Posted on: 19/08/2026
Company Overview :
WM Global Technology Services India Pvt. Ltd. serves as the core technology hub for Walmart Inc., the world's largest retailer. The organization focuses on building and scaling cutting-edge platforms that power global e-commerce, supply chain logistics, and retail operations. Operating at a massive scale, the company leverages advanced engineering to solve complex problems in distributed systems, cloud infrastructure, and data processing, fostering a culture of innovation and high-performance engineering in Bangalore.
Role Overview :
As a Site Reliability Engineer, you will be responsible for ensuring the availability, latency, performance, and efficiency of our critical retail platforms. You will work closely with cross-functional software engineering and product teams to bridge the gap between development and operations, focusing on automation and proactive system health management. Your contributions will directly impact the reliability of services used by millions of customers, ensuring a seamless shopping experience through robust infrastructure design and rapid incident resolution.
Key Responsibilities :
- Design and implement automated solutions to manage infrastructure at scale, reducing manual toil and improving system consistency across environments.
- Lead troubleshooting efforts for complex production incidents, performing root cause analysis to prevent recurrence and maintain high service availability.
- Manage containerized applications using Kubernetes and Docker to ensure efficient resource utilization and seamless deployment cycles.
- Develop and maintain CI/CD pipelines to accelerate software delivery while ensuring rigorous quality and security standards are met.
- Architect infrastructure as code using Terraform and Ansible to provide scalable, repeatable, and secure cloud environments on AWS.
- Configure and optimize monitoring and logging stacks, including Splunk, to gain deep visibility into system performance and user behavior.
- Author and maintain scripts in Python and Shell to automate routine operational tasks and integrate disparate system components.
Required Skillset :
- Demonstrated expertise in managing large-scale distributed systems within an AWS cloud environment, with 3 to 6 years of relevant professional experience.
- Proficiency in Linux system administration, including performance tuning, security hardening, and deep-dive troubleshooting of kernel and application-level issues.
- Strong programming capabilities in Python and Shell scripting, with a focus on writing clean, maintainable code for automation and tooling.
- Hands-on experience with container orchestration platforms like Kubernetes and containerization technologies such as Docker.
- Proven ability to design and manage infrastructure using Terraform and configuration management tools like Ansible.
- Experience with modern observability and monitoring tools, specifically Splunk, to track system health and identify performance bottlenecks.
- Strong analytical mindset with the ability to communicate complex technical issues clearly to both technical and non-technical stakeholders.
- Ability to thrive in a fast-paced, collaborative environment, demonstrating adaptability and a proactive approach to solving operational challenges.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1664394