Posted on: 22/09/2026
Key Responsibilities :
- Provide enterprise-level operational support to Managed Services customers for incident, problem, and change management activities.
- Administer parallel and distributed filesystems such as Lustre, GPFS, BeeGFS, Ceph, Weka, or Vast.
- Optimize storage performance, throughput, metadata operations, and data locality for AI training and inference.
- Build and maintain automation for storage provisioning, monitoring, alerting, quota management, and lifecycle operations.
- Plan and perform maintenance activities.
- Assess customer environments for performance and design issues and propose resolutions.
- Work across technical teams to troubleshoot complex infrastructure issues.
- Create and maintain detailed documentation.
- Serve as a subject matter expert and escalation point for storage technologies.
- Work with vendors to resolve storage issues.
- Communicate with customers and internal team with transparency.
- Support data movement workflows including ingest, replication, caching, tiering, and archiving.
- Troubleshoot storage, Linux, network, and I/O bottlenecks across storage clusters and fabrics.
- Partner with infrastructure, platform, and research teams to support production AI/HPC workloads.
- Evaluate new storage architectures and technologies for scalability, resilience, and cost efficiency.
- Participate in on-call rotation.
Required Qualifications :
- 5+ years of experience with HPC, AI infrastructure, or large-scale storage engineering.
- Bachelor's degree or equivalent Information Systems or related field.
- Strong experience with Linux systems administration.
- Hands-on experience configuring, managing, and tuning distributed or parallel filesystems.
- Experience tuning storage for performance-sensitive workloads.
- Knowledge of HPC schedulers such as Slurm and/or container platforms such as Kubernetes.
- Familiarity with high-speed interconnects such as InfiniBand or RDMA.
- Ability to troubleshoot complex issues across storage, compute, and networking layers.
- Understanding of data protection mechanisms, including data replication, backup strategies, and disaster recovery in HPC environments.
- Experience with machine learning or data science workflows in HPC environments.
- Managed Services or consulting experience.
- Strong background with customer service.
- High level problem-solving and communication skills.
- Strong oral and written communications skills.
Tech Stack :
- Storage : Vast, Ceph, EMC, Pure, Netapp, Parallel storage systems
- Networking : High speed networking with RDMA and InfiniBand
- Tools : Linux, HPC schedulers (Slurm), Kubernetes, Terraform, Ansible, Helm, GitOps, Prometheus, Grafana, Python, Bash
Preferred Qualifications :
- Experience supporting storage solutions for GPU clusters and AI/ML workflows.
- Familiarity with object storage such as S3, MinIO, or Ceph Object Gateway.
- Experience with multi-petabyte environments, caching architectures, and storage isolation in multi-tenant systems.
- Related Storage certifications are a bonus.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
IT Infrastructure Services
Job Code
1673364