Posted on: 02/07/2026
Company Overview :
Trianz Digital Consulting Private Limited is a global digital transformation firm that partners with organizations to execute business strategies through technology-led innovation. The company specializes in cloud, analytics, digital, and infrastructure services, helping enterprises navigate complex digital landscapes. With a focus on delivering measurable business outcomes, Trianz operates across diverse industries, including healthcare, finance, and retail, maintaining a global footprint that supports large-scale digital modernization initiatives for Fortune 500 clients.
Role Overview :
As an Observability Subject Matter Expert, you will serve as the technical authority for designing and implementing comprehensive monitoring and observability frameworks.
You will work closely with cross-functional engineering teams, cloud architects, and client stakeholders to ensure high availability and performance of mission-critical applications.
By transforming raw telemetry data into actionable insights, you will directly influence the stability of complex distributed systems and drive operational excellence across the enterprise.
Key Responsibilities :
- Architect and deploy scalable observability solutions that provide end-to-end visibility into application performance and infrastructure health.
- Design robust monitoring strategies using Prometheus and Grafana to ensure real-time alerting and visualization for complex microservices environments.
- Implement advanced log management and telemetry collection frameworks to streamline incident response and root cause analysis.
- Collaborate with development and DevOps teams to integrate Application Performance Management (APM) tools into the CI/CD pipeline, ensuring performance is a core component of the software development lifecycle.
- Optimize observability costs and performance by refining data collection strategies and reducing noise in monitoring dashboards.
- Mentor junior engineers and provide technical guidance on best practices for system reliability and observability standards.
Required Skillset :
- Possess 8 - 14 years of professional experience in designing and managing enterprise-grade observability and monitoring ecosystems.
- Demonstrate deep technical proficiency in configuring and managing Prometheus, Grafana, and industry-standard Application Performance Management tools.
- Exhibit strong expertise in log management tools and telemetry collection protocols to troubleshoot distributed systems effectively.
- Communicate complex technical concepts clearly to both engineering teams and non-technical stakeholders to drive consensus on architectural decisions.
- Adapt seamlessly to a hybrid work environment across Bangalore, Hyderabad, or Chennai, collaborating effectively with distributed global teams.
- Maintain a proactive approach to problem-solving, with the ability to identify performance bottlenecks before they impact end-user experience.
- Hold a Bachelors or Masters degree in Computer Science, Information Technology, or a related field, backed by a proven track record in high-scale production environments.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1650893