Posted on: 23/09/2026
The role involves working on Kafka platform administration, production support, performance monitoring, troubleshooting, cloud infrastructure, and automation in a 24/7 operational environment.
Key Responsibilities :
- Design, deploy, configure, and administer Confluent Kafka / Apache Kafka clusters in enterprise environments.
- Manage Kafka brokers, topics, partitions, replication, consumer groups, and access controls.
- Monitor Kafka cluster health, performance, capacity, replication status, and overall platform availability.
- Troubleshoot Kafka-related issues including broker failures, consumer lag, replication issues, partition
imbalance, connectivity problems, and performance bottlenecks.
- Perform Kafka cluster upgrades, configuration changes, patching, scaling, and maintenance activities.
- Configure and manage Kafka security, including authentication, authorization, ACLs, SSL/TLS, and related security mechanisms.
- Work with Kafka producers and consumers to identify and resolve message delivery, throughput, and latency issues.
- Monitor and optimize Kafka performance based on throughput, latency, broker utilization, partition distribution, and consumer lag.
- Support Kafka deployments and infrastructure hosted on AWS.
- Work with AWS services and networking components relevant to Kafka infrastructure and application connectivity.
- Develop and maintain Terraform configurations for provisioning and managing cloud infrastructure.
- Automate repetitive infrastructure and operational activities using Infrastructure as Code practices.
- Implement monitoring, alerting, logging, and operational dashboards for Kafka environments.
- Participate in incident management, root-cause analysis, problem management, and service restoration activities.
- Analyze production incidents and provide permanent solutions to recurring Kafka and infrastructure issues.
- Maintain operational documentation, architecture diagrams, configuration standards, and troubleshooting procedures.
- Collaborate with application, DevOps, cloud, infrastructure, and security teams to support Kafka-based applications.
- Participate in planned maintenance activities and provide production support in a 24/7 rotational shift environment.
- Ensure Kafka environments adhere to availability, reliability, security, and operational standards.
Required Skills and Experience :
- 6+ years of experience in Kafka administration, engineering, platform support, or a closely related role.
- Strong hands-on experience with Confluent Kafka and/or Apache Kafka.
- Good understanding of Kafka architecture, including brokers, topics, partitions, replicas, producers, consumers, and consumer groups.
- Experience in Kafka cluster administration, configuration, monitoring, troubleshooting, and performance tuning.
- Strong understanding of Kafka replication, partition management, leader election, and consumer lag.
- Experience troubleshooting Kafka production environments and handling critical incidents.
- Hands-on experience with AWS cloud infrastructure.
- Good understanding of AWS networking, compute, storage, and security concepts relevant to distributed platforms.
- Strong experience with Terraform and Infrastructure as Code.
- Ability to create, maintain, and troubleshoot Terraform configurations and modules.
- Experience with monitoring and logging tools used for Kafka and cloud infrastructure environments.
- Good understanding of Linux/Unix administration and command-line troubleshooting.
- Strong analytical and problem-solving skills.
- Good communication skills and ability to work effectively with cross-functional teams.
Preferred Skills :
- Experience with Confluent Platform and Confluent ecosystem components.
- Experience with Kafka Connect, Schema Registry, and related Kafka components.
- Knowledge of Kafka security mechanisms such as SSL/TLS, SASL, and ACLs.
- Experience with AWS-based Kafka deployments and integrations.
- Knowledge of CI/CD and DevOps practices.
- Experience with scripting or automation using Shell, Python, or similar technologies.
- Familiarity with containerized or cloud-native environments.
- Experience working in enterprise production support environments.
- Exposure to high-availability and disaster-recovery architectures for distributed systems.
Operational Responsibilities :
- Provide L2/L3 production support for Kafka environments.
- Monitor and respond to alerts within defined operational SLAs.
- Participate in 24/7 rotational shifts and on-call activities as required.
- Perform incident investigation and root-cause analysis.
- Coordinate with relevant teams during major incidents and planned maintenance.
- Execute standard operating procedures for Kafka cluster maintenance and recovery.
- Ensure timely resolution and documentation of production issues.
Education :
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related technical discipline is preferred.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Systems Administration
Job Code
1673944