HamburgerMenu
hirist

Applied AI Site Reliability Engineer III

Clarus- Impact Network
5 - 10 Years
Multiple Locations

Posted on: 03/09/2026

Job Description

About the Role :

We are looking for an experienced Applied AI Site Reliability Engineer III to join a high-performing Product Engineering team responsible for building and operating reliable, scalable, cloud-native platforms and AI-enabled products.

The ideal candidate will combine strong SRE/Production Engineering, Cloud, Observability, Performance Engineering, Kubernetes and Infrastructure-as-Code expertise with hands-on experience operating AI/ML and agentic workloads in production.

Key Responsibilities :

- Own reliability, performance, availability and cost outcomes through SLIs, SLOs, SLAs and error budgets.

- Design and operate production-grade observability across metrics, logs and distributed tracing.

- Build dashboards, actionable alerts, runbooks and automated reliability checks.

- Drive production readiness, release gating and operational acceptance of applications and platforms.

- Manage cloud-native infrastructure across AWS, Azure or GCP.

- Develop and maintain infrastructure using Terraform, Kubernetes, Docker and CI/CD.

- Perform performance, load, capacity and resilience testing.

- Implement chaos engineering and failure-testing practices to improve system resilience.

- Participate in incident management, root-cause analysis and blameless postmortems.

- Identify and eliminate operational toil through automation.

- Support and operate AI/ML and agentic workloads in production.

- Address AI-specific reliability challenges including model drift, train/serve skew, output variance, latency and token/GPU cost anomalies.

- Work with engineering, security, risk, architecture and data teams to establish production standards and controls.

- Contribute to MLOps/LLMOps, AI control planes, model/LLM gateways and guardrails from an operability and performance perspective.

- Drive cloud and AI cost optimization / FinOps initiatives.

- Continuously improve reliability, scalability, security and operational efficiency.

Required Skills & Experience :

- 5+ years of software engineering / SRE / production engineering experience.

- 3+ years working with large-scale, distributed, cloud-native production systems.

- Strong experience with one or more : Python / Go / Bash / Java / C#/.NET, SQL / NoSQL, Kubernetes & Docker, Terraform, CI/CD, GitHub / Azure DevOps.

- Hands-on experience with AWS, Azure or GCP.

- Strong understanding of : SLI / SLO / SLA, Error budgets, Incident management, On-call operations, Production readiness, Capacity planning, Autoscaling, Environment integrity & configuration drift.

- Experience with observability tools such as OpenTelemetry, Prometheus, Grafana, Datadog, Dynatrace, CloudWatch, Azure Monitor, Google Cloud Operations or Splunk.

- Experience with load/performance testing using tools such as JMeter, LoadRunner or k6.

- Exposure to chaos engineering tools such as AWS Fault Injection Simulator or Azure Chaos Studio.

- Experience operating AI/ML or GenAI workloads in production.

- Understanding of MLOps / LLMOps, AI reliability and agentic systems.

- Experience with AI platforms such as Azure OpenAI, AWS Bedrock or Vertex AI is highly desirable.

- Familiarity with MLflow, LangSmith, LangFuse or equivalent AI/agent orchestration platforms.

- Understanding of DevSecOps, SRE, Lean/XP methodologies and secure production practices.

- Knowledge of RBAC, secrets management, least privilege and deployment controls.

- Strong software engineering fundamentals including OOP/OOD, data structures, algorithms, system design and code instrumentation.

What We're Looking For :

- Strong hands-on engineering mindset.

- Excellent troubleshooting and problem-solving skills.

- Ability to work across engineering, security and architecture teams.

- Strong understanding of production systems and operational excellence.

- Data-driven approach to reliability and performance improvements.

- Excellent communication and stakeholder-management skills.

- Willingness to continuously learn emerging cloud and AI technologies.

Travel : Up to 10% may be required.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...