HamburgerMenu
hirist

AWS DevOps/Site Reliability Engineer

Big Ideas Social Media Recruitment
5 - 7 Years
Gurgaon/Gurugram

Posted on: 10/06/2026

Job Description

Job Summary :

We are looking for a highly skilled AWS DevOps / SRE Engineer with strong expertise in cloud infrastructure, application monitoring, observability, and production support. The ideal candidate should have hands-on experience in AWS environments, CI/CD pipelines, incident management, and end-to-end troubleshooting across application stacks.

The role requires strong analytical and debugging capabilities to identify root causes, improve system reliability, and ensure high availability of business-critical applications.

Key Responsibilities :

- Manage and maintain AWS cloud infrastructure and production environments

- Design, implement, and optimize CI/CD pipelines for automated deployments

- Monitor application performance, reliability, and availability across environments

- Perform end-to-end application tracing, debugging, and troubleshooting

- Handle production incidents, conduct root cause analysis (RCA), and implement preventive measures

- Work closely with development, QA, and infrastructure teams to improve application stability

- Implement observability best practices using logs, metrics, and traces

- Support high-availability systems and ensure platform reliability

- Automate operational tasks and infrastructure provisioning

- Participate in on-call rotations and production support activities when required

- Collaborate with cross-functional teams including backend and frontend teams.

Mandatory Skills :

- Strong hands-on experience with AWS Cloud services

- Expertise in DevOps practices, CI/CD pipelines, and deployment automation

- Exposure to SRE concepts, including monitoring, reliability engineering, and incident management

- Experience with application monitoring and tracing tools such as Dynatrace

- Strong understanding of observability concepts (logs, metrics, traces)

- Ability to perform end-to-end application tracing and troubleshooting

- Strong production support and root cause analysis skills

- Experience with backend technologies such as Java and Python

- Strong debugging and troubleshooting skills across the application stack

- Good understanding of application architecture and distributed systems

- Experience working in Linux/Unix environments

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...