- Own the platform infrastructure strategy for the AI capability programme across all six pods, covering CI/CD, environment provisioning, secrets management, observability, and reliability engineering.
- Design and maintain multi-environment infrastructure (development, staging, production) with appropriate isolation, promotion gates, and security controls for classified, internal, and public workloads.
- Define infrastructure-as-code standards (Terraform, Ansible, or equivalent) and drive their adoption across pods.
- Establish the observability stack metrics, logs, tracing, alerting covering both AI-model and application layers, in coordination with the MLOps Lead.
- Own incident response protocols and on-call rotation across the programme; conduct post-incident reviews and drive corrective actions.
- Define availability and reliability SLOs for each pod's production services; track error budgets and coordinate with Product Managers on trade-offs.
- Oversee infrastructure security hardening network policies, WAF operations, DDoS protection, VAPT remediation at infrastructure layer in coordination with the Security & Compliance Engineer.
- Manage relationships with NIC, IndiaAI Compute, MeghRaj, and other government infrastructure providers; coordinate capacity planning and reserved-versus-on-demand mixing.
- Ensure Data Residency, Data Storage, and Data Lifecycle Management compliance per the LoE (data within India, AES-256 at rest, TLS 1.3 in transit, sanitisation at project closure).
- Mentor DevOps/SRE Engineers and coordinate with MLOps Lead on shared platform responsibilities.