Posted on: 03/09/2026
About the Role :
We are looking for an experienced Applied AI Site Reliability Engineer III to join a high-performing Product Engineering team responsible for building and operating reliable, scalable, cloud-native platforms and AI-enabled products.
The ideal candidate will combine strong SRE/Production Engineering, Cloud, Observability, Performance Engineering, Kubernetes and Infrastructure-as-Code expertise with hands-on experience operating AI/ML and agentic workloads in production.
Key Responsibilities :
- Own reliability, performance, availability and cost outcomes through SLIs, SLOs, SLAs and error budgets.
- Design and operate production-grade observability across metrics, logs and distributed tracing.
- Build dashboards, actionable alerts, runbooks and automated reliability checks.
- Drive production readiness, release gating and operational acceptance of applications and platforms.
- Manage cloud-native infrastructure across AWS, Azure or GCP.
- Develop and maintain infrastructure using Terraform, Kubernetes, Docker and CI/CD.
- Perform performance, load, capacity and resilience testing.
- Implement chaos engineering and failure-testing practices to improve system resilience.
- Participate in incident management, root-cause analysis and blameless postmortems.
- Identify and eliminate operational toil through automation.
- Support and operate AI/ML and agentic workloads in production.
- Address AI-specific reliability challenges including model drift, train/serve skew, output variance, latency and token/GPU cost anomalies.
- Work with engineering, security, risk, architecture and data teams to establish production standards and controls.
- Contribute to MLOps/LLMOps, AI control planes, model/LLM gateways and guardrails from an operability and performance perspective.
- Drive cloud and AI cost optimization / FinOps initiatives.
- Continuously improve reliability, scalability, security and operational efficiency.
Required Skills & Experience :
- 5+ years of software engineering / SRE / production engineering experience.
- 3+ years working with large-scale, distributed, cloud-native production systems.
- Strong experience with one or more : Python / Go / Bash / Java / C#/.NET, SQL / NoSQL, Kubernetes & Docker, Terraform, CI/CD, GitHub / Azure DevOps.
- Hands-on experience with AWS, Azure or GCP.
- Strong understanding of : SLI / SLO / SLA, Error budgets, Incident management, On-call operations, Production readiness, Capacity planning, Autoscaling, Environment integrity & configuration drift.
- Experience with observability tools such as OpenTelemetry, Prometheus, Grafana, Datadog, Dynatrace, CloudWatch, Azure Monitor, Google Cloud Operations or Splunk.
- Experience with load/performance testing using tools such as JMeter, LoadRunner or k6.
- Exposure to chaos engineering tools such as AWS Fault Injection Simulator or Azure Chaos Studio.
- Experience operating AI/ML or GenAI workloads in production.
- Understanding of MLOps / LLMOps, AI reliability and agentic systems.
- Experience with AI platforms such as Azure OpenAI, AWS Bedrock or Vertex AI is highly desirable.
- Familiarity with MLflow, LangSmith, LangFuse or equivalent AI/agent orchestration platforms.
- Understanding of DevSecOps, SRE, Lean/XP methodologies and secure production practices.
- Knowledge of RBAC, secrets management, least privilege and deployment controls.
- Strong software engineering fundamentals including OOP/OOD, data structures, algorithms, system design and code instrumentation.
What We're Looking For :
- Strong hands-on engineering mindset.
- Excellent troubleshooting and problem-solving skills.
- Ability to work across engineering, security and architecture teams.
- Strong understanding of production systems and operational excellence.
- Data-driven approach to reliability and performance improvements.
- Excellent communication and stakeholder-management skills.
- Willingness to continuously learn emerging cloud and AI technologies.
Travel : Up to 10% may be required.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1668388