HamburgerMenu
hirist

Staff Engineer - Platform Reliability

Value Lane Consulting
8 - 15 Years
Bangalore

Posted on: 18/09/2026

Job Description

Role Overview :

As a Staff Engineer for Platform Reliability Engineering, you will serve as a technical anchor for our mission-critical infrastructure, ensuring our systems remain resilient, scalable, and performant at massive scale. You will operate at the intersection of software engineering and systems operations, collaborating closely with product engineering teams, security architects, and leadership to define the roadmap for platform stability. By architecting robust backend systems and implementing advanced observability frameworks, you will directly influence the uptime and reliability of services that support millions of users, ultimately driving business growth through operational excellence and technical innovation in our engineering hub.

Key Responsibilities :

- Architect and implement highly available backend systems on AWS to ensure seamless service delivery and fault tolerance for global customers.

- Lead the design and deployment of advanced observability services and monitoring tools to proactively identify bottlenecks and reduce mean time to resolution.

- Drive the evolution of our platform architecture by mentoring senior engineers and establishing best practices for system reliability and performance engineering.

- Optimize cloud infrastructure costs and resource utilization by implementing automated scaling policies and efficient backend architecture patterns.

- Partner with cross-functional teams to conduct deep-dive post-mortems and implement systemic improvements that prevent recurring incidents and enhance overall platform health.

Required Skillset :

- Demonstrated expertise in designing and maintaining complex, distributed backend architectures with a focus on high-concurrency and low-latency requirements.

- Advanced proficiency in Python for building automation tools, custom monitoring agents, and infrastructure-as-code solutions.

- Deep hands-on experience managing and scaling cloud-native environments on AWS, including mastery of VPC networking, IAM, and managed database services.

- Proven ability to implement and manage enterprise-grade observability stacks, translating raw telemetry data into actionable insights for engineering stakeholders.

- Exceptional communication skills, with the ability to articulate complex technical trade-offs to non-technical stakeholders and influence senior leadership decisions.

- Strong collaborative mindset, capable of fostering a culture of reliability and engineering rigor within a high-growth, hybrid work environment.

- A Bachelor's or Master's degree in Computer Science or a related field, complemented by 8 - 15 years of progressive experience in site reliability or platform engineering roles.

Mandatory - Platform Reliability :

- 8+ years of software engineering experience, with substantial time owning production systems and their reliability.

- Demonstrated ownership of SLOs, error budgets and on-call for business-critical systems.

- Proven incident command experience on serious, customer-impacting outages, and a track record of eliminating repeat causes.

- Strong resilience engineering instincts - failure modes, degradation strategies, capacity planning and performance tuning.

- Exceptional debugging ability across application, database, messaging, network and infrastructure layers.

Mandatory - Full Stack Engineering :

- Deep backend engineering expertise in Java, Python, Go or equivalent, writing production-quality code rather than scripting alone.

- Strong, current frontend capability with React/TypeScript or equivalent.

- Proven track record designing distributed systems - APIs, data modelling, event-driven architecture, concurrency and consistency trade-offs.

- Comfortable reading and changing application code across the stack to fix reliability problems at source.

Mandatory - LLM & Agentic AI :

- Hands-on experience building and running LLM-powered or Agentic AI systems in production, including their operational failure modes.

- Practical understanding of prompt and context strategies, RAG, tool use and agent orchestration.

- Experience evaluating and testing AI systems where the output is not deterministic.

- Fluent use of Claude Code, Cursor or equivalent AI development tools in daily engineering work.

Mandatory - Leadership & Ways of Working :

- Track record of technical influence beyond personal output - standards adopted, teams unblocked and engineers levelled up.

- Ability to work independently with ambiguous requirements and bring clarity to others.

- Strong written communication across design documents, RFCs, post-mortems and runbooks.

- Bias towards owning outcomes rather than tickets, and towards fixing causes rather than symptoms.

Desirable :

- IoT, energy, industrial or real-time data experience.

- Kafka, MQTT, time-series databases or stream processing at scale.

- Experience operating AI-first products or Agentic AI systems in production.

- Experience building a platform or reliability function from 0 to 1.

- Security engineering, compliance or disaster recovery and business continuity experience.

- Experience in high-velocity product or startup environments.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...