HamburgerMenu
hirist

Storage & Data Protection Engineer - HPC/AI Infrastructure

AHEAD
5 - 10 Years
Gurgaon/Gurugram

Posted on: 22/09/2026

Job Description

Key Responsibilities :

- Provide enterprise-level operational support to Managed Services customers for incident, problem, and change management activities.

- Administer parallel and distributed filesystems such as Lustre, GPFS, BeeGFS, Ceph, Weka, or Vast.

- Optimize storage performance, throughput, metadata operations, and data locality for AI training and inference.

- Build and maintain automation for storage provisioning, monitoring, alerting, quota management, and lifecycle operations.

- Plan and perform maintenance activities.

- Assess customer environments for performance and design issues and propose resolutions.

- Work across technical teams to troubleshoot complex infrastructure issues.

- Create and maintain detailed documentation.

- Serve as a subject matter expert and escalation point for storage technologies.

- Work with vendors to resolve storage issues.

- Communicate with customers and internal team with transparency.

- Support data movement workflows including ingest, replication, caching, tiering, and archiving.

- Troubleshoot storage, Linux, network, and I/O bottlenecks across storage clusters and fabrics.

- Partner with infrastructure, platform, and research teams to support production AI/HPC workloads.

- Evaluate new storage architectures and technologies for scalability, resilience, and cost efficiency.

- Participate in on-call rotation.

Required Qualifications :

- 5+ years of experience with HPC, AI infrastructure, or large-scale storage engineering.

- Bachelor's degree or equivalent Information Systems or related field.

- Strong experience with Linux systems administration.

- Hands-on experience configuring, managing, and tuning distributed or parallel filesystems.

- Experience tuning storage for performance-sensitive workloads.

- Knowledge of HPC schedulers such as Slurm and/or container platforms such as Kubernetes.

- Familiarity with high-speed interconnects such as InfiniBand or RDMA.

- Ability to troubleshoot complex issues across storage, compute, and networking layers.

- Understanding of data protection mechanisms, including data replication, backup strategies, and disaster recovery in HPC environments.

- Experience with machine learning or data science workflows in HPC environments.

- Managed Services or consulting experience.

- Strong background with customer service.

- High level problem-solving and communication skills.

- Strong oral and written communications skills.

Tech Stack :

- Storage : Vast, Ceph, EMC, Pure, Netapp, Parallel storage systems

- Networking : High speed networking with RDMA and InfiniBand

- Tools : Linux, HPC schedulers (Slurm), Kubernetes, Terraform, Ansible, Helm, GitOps, Prometheus, Grafana, Python, Bash

Preferred Qualifications :

- Experience supporting storage solutions for GPU clusters and AI/ML workflows.

- Familiarity with object storage such as S3, MinIO, or Ceph Object Gateway.

- Experience with multi-petabyte environments, caching architectures, and storage isolation in multi-tenant systems.

- Related Storage certifications are a bonus.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...