HamburgerMenu
hirist

Senior Linux Systems Engineer - High Performance Computing

The Judge Group
6 - 8 Years
Chennai

Posted on: 17/08/2026

Job Description

Job Title : Senior Linux Systems Engineer High Performance Computing (HPC)

Location : Chennai - Hybrid

Job Type : Permanent Role

Experience Required : 6+ Years in IT (Minimum 4+ Years in Active Infrastructure Automation or Systems Development)

Notice Period : Immediate to 15 Days Max

Education Required : Bachelor's Degree in IT/Engineering

Job Summary :

We are seeking a highly skilled and motivated Linux Systems Engineer to join our global HPC Supercomputing platform team. In this role, you will focus on the design, optimization, and scaling of core Computer-Aided Engineering (CAE) simulation, graphical rendering, and analytics capabilities used by our global product development groups.

The ideal candidate is a seasoned Level 3 systems engineer who bridges the gap between low-level Linux kernel/network structures and high-concurrency scientific computing.


You will be responsible for profiling and tuning heavy distributed simulations, integrating scalable CAE applications across thousands of compute nodes, building custom command-line interface (CLI) tooling and APIs for engineer consumption, and identifying systemic architecture bottlenecks through advanced telemetry layers.

Key Responsibilities :

- HPC Workload Optimization & Tuning : Install, profile, benchmark, and tune core CAE application workflows and highly distributed workloads to maximize compute utilization across the HPC platform.

- Hardware Evaluation : Test and evaluate demanding scientific workloads across the latest CPU and GPU microarchitectures to drive hardware choices and ensure cost-efficient simulation delivery.

- Tooling & API Development : Design and program custom CLI tools, utilities, and inner APIs (using Python, Go, or Bash) that internal engineering teams consume to streamline access to supercomputing infrastructure.

- Parallel Architecture Integration : Manage, monitor, and scale distributed communication layers and Message Passing Interface standards (IntelMPI, OpenMPI, etc.) across clustered network paths.

- Cluster Management & Batch Scheduling : Configure and manage automated environment state lines and job queues using batch schedulers (such as PBS Pro or Slurm).

- Infrastructure as Code (IaC) & Containerization : Provision, automate, and configure cluster nodes via configuration management utilities (Ansible, Puppet, or Chef) and deploy secure compute containers using Docker, Apptainer (Singularity), or Kubernetes.

- Deep Telemetry & Observability : Architect platform telemetry, cluster dashboards, and resource tracking pipelines using metric collection and visualization engines like Prometheus and Grafana to identify system limits.

- Technical Troubleshooting : Investigate, diagnose, and resolve complex structural failures across Linux kernels, high-speed fabrics, clustered storage volumes, and scientific runtime configurations.

Skills & Qualifications :

Required Technical DNA :

- Operating Systems & Tuning : Strong, production-level engineering and command of Linux operating systems, ideally inside an HPC or high-concurrency server cluster environment.

- Automation & Coding : Advanced coding and systems scripting proficiency in Python, Go, or Bash.

- Parallel Workloads : Solid familiarity with distributed workloads, network topologies, and Message Passing Interface (MPI) frameworks.

- HPC Applications : Hands-on operational experience supporting, profiling, or writing automation for CAE or parallel scientific analysis tools.

- Tenure : 6+ years of total IT experience, with at least 4 years spent in active infrastructure automation, product delivery, or system software development.

Preferred Qualifications (Nice to Have) :

- Queue Managers : Hands-on experience configuring batch engines like PBS Pro or Slurm.

- Configuration Management : Production experience automating bare-metal or cloud platforms using Ansible.

- Technical utility deploying modern telemetry systems (Prometheus, Grafana) and container nodes (Apptainer, Docker).

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...