Posted on: 06/10/2026
Key Responsibilities :
- Design and implement CI/CD pipelines for models, prompts, agents, and supporting infrastructure across development, test, and production environments.
- Build and maintain deployment automation for AI workloads, including versioning and rollback mechanisms, environment promotion workflows, and runtime safeguards.
- Set up and operate observability for AI applications and agents : tracing, monitoring, alerting, token-consumption analysis, latency tracking, and incident diagnostics.
- Implement evaluation pipelines and acceptance gates for quality, groundedness, task adherence, safety, and agent-specific behaviour.
- Drive prompt lifecycle management, RAG optimisation, and semantic retrieval tuning, and integrate vector- or search-based knowledge components where needed.
- Partner with security, engineering, and data teams to build identity, secrets management, compliance controls, and cost optimisation into the operating model.
- Contribute to platform automation, runbooks, and on-call readiness, and lead post-incident reviews that improve production AI services.
Qualifications :
- DevOps and platform engineering : Experience operating production cloud workloads with CI/CD, monitoring, and infrastructure automation.
- MLOps / LLMOps / AgentOps : Hands-on practice in deployment, monitoring, retraining or re-evaluation, and controlled release management for AI systems.
- Observability : Strong grasp of logs, metrics, traces, runtime telemetry, and production diagnostics for AI workloads.
- Generative AI engineering : Familiarity with retrieval-augmented systems, prompt engineering, tool-calling flows, and agent behaviour debugging.
- Operational excellence : A reliability-first mindset with strong attention to security, incident response, and cost-performance trade-offs.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
ML / DL Engineering
Job Code
1676982