Posted on: 29/09/2026
Job Description :
Key Responsibilities :
- Administer and support Hadoop ecosystem components (HDFS, YARN) across Production, Test, and Development environments
- Manage and maintain Apache Spark and Airflow infrastructure
- Monitor cluster health, troubleshoot performance issues, and optimize system performance
- Implement and maintain High Availability (HA), Disaster Recovery (DR), backup, and restore strategies
- Manage logging, monitoring, and alerting systems using Prometheus, Grafana, and ELK Stack
- Perform Linux (Ubuntu) system administration, including patching, upgrades, and security hardening
- Develop and maintain automation scripts using Shell and/or Python
- Handle Kerberos authentication setup and maintenance
- Support incident management, root cause analysis (RCA), and ensure SLA adherence
- Manage capacity planning and system scalability
- Follow CI/CD and DevOps best practices for platform improvements and deployments
Required Skills :
- Strong experience with Hadoop ecosystem (HDFS, YARN)
- Handson expertise in Apache Spark and Apache Airflow
- Proficiency in Linux system administration (Ubuntu), including security hardening
- Scripting skills in Shell and Python for automation
- Familiarity with monitoring and observability tools: Prometheus, Grafana, ELK Stack
- Knowledge of High Availability, Disaster Recovery, backup, and restore concepts
- Experience with Kerberos authentication setup and maintenance
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Systems Administration
Job Code
1675489