HamburgerMenu
hirist

DDN - L4 Staff Engineer - File System

DDN
8 - 12 Years
Pune

Posted on: 25/09/2026

Job Description

Job Description :

As a Staff Engineer - L4, you'll be an escalation point for the most complex and critical issues affecting enterprise and hyperscale environments.

This hands-on role is ideal for a deep technical expert who thrives under pressure and has a passion for solving distributed system challenges at scale.

This role is part of the Infinia Core engineering team.

You'll collaborate with Engineering, Product Management, and Field teams to drive root cause resolutions, define architectural best practices, and continuously improve product resiliency.

Leveraging AI tools and automation, you'll reduce time-to-resolution, streamline diagnostics, and elevate the support experience for strategic customers.

Key Responsibilities :

1. Technical Expertise & Escalation Leadership :

- Own critical customer case escalations end-to-end, including deep root cause analysis and mitigation strategies.

- Act as one of the technical escalation points for Infinia incidents - especially in production-impacting scenarios.

- Lead war rooms, live incident bridges, and cross-functional response efforts with other engineering, QA, and Field teams.

- Utilize AI-powered debugging, log analysis, and system pattern recognition tools to accelerate resolution.

2. Product Knowledge & Value Creation :

- Become a subject-matter expert on Infinia internals: metadata handling, storage fabric interfaces, performance tuning, AI integration, etc.

- Reproduce complex customer issues and propose product improvements or workarounds.

- Author and maintain detailed runbooks, performance tuning guides, and RCA documentation.

- Feed real-world support insights back into the development cycle to improve reliability and diagnostics.

3. Customer Engagement & Business Enablement :

- Partner with Field CTOs, Solutions Architects, and Sales Engineers to ensure customer success.

- Translate technical issues into executive-ready summaries and business impact statements.

- Participate in post-mortems and executive briefings for strategic accounts.

- Drive adoption of observability, automation, and self-healing support mechanisms using AI/ML tools.

- Delivering training to customer support and field engineering.

Required Qualifications :

- 8+ years in enterprise storage, distributed systems, or cloud infrastructure support/engineering.

- Deep understanding of file systems (S3, POSIX, NFS), storage performance, and Linux kernel internals.

- Scripting and Coding using Python, Go, C++.

- Proven debugging skills at system/protocol/app levels (e.g., strace, tcpdump, perf).

- Hands-on experience with troubleshooting on Linux.

- Exposure to RDMA, NVMe-oF, or high-performance networking stacks.

- Exceptional communication and executive reporting skills.

- Experience using AI tools (e.g., log pattern analysis, LLM-based summarisation, automated RCA tooling) to accelerate diagnostics and reduce MTTR.

Preferred Qualifications :

- Experience with DDN, VAST, Weka, or similar scale-out file systems.

- Strong scripting/coding ability in Python, Bash, or Go.

- Familiarity with observability platforms: Prometheus, Grafana, ELK, OpenTelemetry.

- Knowledge of replication, consistency models, and data integrity mechanisms.

- Exposure to Sovereign AI, LLM model training environments, or autonomous system data architectures.

- This position requires participation in an on-call rotation to provide after-hours support as needed.

Success Metrics - First 30 Days :

1. Technical Ramp-Up :

- Complete Infinia training, labs, and architecture deep dives.

- Stand up a fully functioning Infinia test system.

- Shadow at least 5 complex escalations and participate in 2 customer calls.

2. Operational Integration :

- Lead one live incident response and deliver a full RCA within 48 hours.

- Propose 3+ enhancements to internal tools, AI/automation usage, or documentation.

- Establish key partnerships with Engineering and Field teams.

3. Strategic Insight :

- Deliver a written 30-day reflection with gaps and high-impact recommendations.

- Begin identifying patterns where AI or automation can reduce MTTR or improve proactive detection.

Success Metrics - Beyond 30 Days :

- MTTR on high-severity cases is consistently below internal SLAs.

- Volume and quality of resolved L4 escalations.

- Strategic tooling or automation contributions adopted across the support org.

- Executive-ready RCAs that inform product improvement.

- High-impact engagements with strategic accounts (prevention, performance tuning, etc.).

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...