HamburgerMenu
hirist

Site Reliability/DevOps Engineer - Kubernetes/OpenShift

True Tech Professionals
7 - 15 Years
Remote

Posted on: 24/09/2026

Job Description

Role : SRE / DevOps Engineer - Linux | Kubernetes | OpenShift

We are looking for a strong SRE / DevOps Engineer with 7+ years of experience who has deep hands-on expertise in Linux infrastructure, Kubernetes/OpenShift administration, automation, troubleshooting, and production operations.

Location : India - Remote / Hybrid (Full time role)

Experience : 7+ Years

Role : SRE / DevOps Engineer

Key Responsibilities & Requirements :

- Linux / RHEL Administration

- Strong hands-on experience with Linux, RHEL/CentOS administration and troubleshooting

1. Strong understanding of Linux internals :

- CPU scheduling

- Memory allocation

- Processes & threads

- Networking

- Filesystems

- Container internals

- User-space to kernel-level troubleshooting

2. Kubernetes & OpenShift :

- Hands-on Kubernetes and OpenShift administration

- Cluster operations, deployments, scaling and upgrades

- Networking, storage and RBAC

- Troubleshooting pods, nodes, services and cluster issues

- Strong understanding of production Kubernetes/OpenShift environments

- Containers

- Strong Docker/Podman experience

- Container image creation, optimization and troubleshooting

- Understanding of container runtime and networking concepts

3. Python / Automation :

- Basic to intermediate Python skills

- Scripting, API integration, debugging and testing

- Understanding of SDLC and code quality

- Ability to explain Python concepts and time/space complexity

- Strong Bash scripting experience

4. Infrastructure Automation :

- Hands-on Ansible and Terraform

- Infrastructure provisioning and configuration management

- Build and deployment automation

5. CI/CD & DevOps :

- Git / GitLab

- GitLab CI / GitHub Actions

- CI/CD pipeline development and maintenance

- Build automation and release management

- Networking

6. Strong fundamentals in :

- HTTP

- DNS

- DHCP

- ARP

- VLAN

- VXLAN

- Container networking

- SRE / Production Operations

- Incident management

- Production troubleshooting

- Root Cause Analysis (RCA)

- SLIs / SLOs

- Reliability engineering

- Performance and availability troubleshooting

- Observability

- Prometheus

- ELK / Sumo Logic

- Logs, metrics and traces

- Monitoring and production issue analysis

7. Infrastructure & Security :

- Cloud and bare-metal infrastructure

- Redfish APIs

- Access control

- Security and compliance

Good to Have :

- ArgoCD / GitOps

- Knative

- Istio

- Advanced OpenShift administration

- Cluster management and upgrades

Ideal Candidate

We are specifically looking for engineers who are strong in Linux troubleshooting and infrastructure internals, rather than candidates focused only on application-level DevOps.

You should be comfortable troubleshooting issues across the full infrastructure stack - from Linux/kernel concepts and networking to containers, Kubernetes/OpenShift and production systems

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...