HamburgerMenu
hirist

Nisum - Senior Site Reliability Engineer - AWS Infrastructure

Nisum
7 - 10 Years
Hyderabad

Posted on: 15/09/2026

Job Description

Responsibilities :

- Manage and support AWS infrastructure and troubleshoot issues across AWS Linux and Ubuntu environments.

- Deploy and maintain applications using Docker, Kubernetes, and Helm.

- Monitor Kubernetes workloads, pods, services, deployments, configurations, and troubleshoot failures.

- Develop, maintain, and troubleshoot CI/CD pipelines using Jenkins and GitHub Actions, with code maintained in Bitbucket/GitHub/GitLab.

- Automate application deployments, validations, rollbacks, and repetitive operational activities.

- Monitor applications and infrastructure using Splunk, Dynatrace, Grafana, Prometheus, ELK, and CloudWatch.

- Analyze logs, metrics, alerts, and performance data to identify issues proactively.

- Follow Change, Release, Hot Fix, and ECRQ processes according to ITIL/ITSM standards.

- Work with ServiceNow, Remedy, Jira, and Rally for incident, change, problem, and release tracking.

- Troubleshoot MySQL/database-related issues and coordinate with database teams when deeper investigation is required.

- Use Bash and Python to automate health checks, monitoring, log analysis, deployment activities, and operational tasks.

- Troubleshoot Akamai traffic-routing issues and application connectivity problems.

- Support Kafka and Axon API related issues, including message flow, connectivity, and application integration problems.

- Apply Agentic AI where appropriate for log analysis, incident detection, deployment assistance, auto-remediation, and escalation.

Tech Stack :

- Cloud: AWS

- Containers: Docker, Kubernetes, Helm

- Job / Scheduler: Crontab

- CI/CD: Jenkins, Bitbucket, GitHub Actions

- ITIL/ITSM: Incident, Problem, Change, Release, Hot fix, ECRQ Management

- Certificate Renewal: Venafi / CerTIS

- Database: MySQL, SQL Developer

- Operating Systems: AWS Linux, Ubuntu, Putty

- Monitoring: Splunk, Dynatrace, Grafana, Prometheus, ELK, CloudWatch

- Service Management: Remedy, Jira, ServiceNow, Rally

- Version Control: Bitbucket, GitHub, GitLab

- Scripting: Bash Shell, Python

- Traffic Routing: Akamai

- Messaging: Apache Kafka, Axon API

Education:

- Bachelor's degree in Computer Science, Information Systems, Engineering, Computer Applications, or related fields.

Benefits:

- Continuous Learning: Year-round training sessions offered for skill enhancement.

- Certifications: Company-sponsored certifications on an as-needed basis.

- Parental Medical Insurance: Opt-in parental medical insurance provided.

- Activities: Team building activities including hackathons, sports, and festival celebrations.

- Meals: Free daily dinner and subsidized lunch provided.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Posted by

Recruiter

HR at Nisum

Last Active: NA as recruiter has posted this job through third party tool.

Job Views:  
1
Applications:  0
Recruiter Actions:  0

Posted in

DevOps / SRE

Functional Area

Site Reliability Engineering

Job Code

1671599

Loading chat...