HamburgerMenu
hirist

Site Reliability Engineer - Artificial Intelligence & Machine Learning Plaform

SysMind
8 - 10 Years
Multiple Locations

Posted on: 06/08/2026

Job Description

About the Role :

We are looking for an experienced AI Site Reliability Engineer (AI SRE) to drive the reliability, scalability, and operational excellence of enterprise AI and Machine Learning platforms.


This role combines expertise in Site Reliability Engineering (SRE), MLOps, Production Support, and AI Platform Engineering to ensure highly available, secure, and resilient AI-powered applications.

Note: Candidates with Investment Banking domain experience and a strong background in Application Production Support are highly preferred.

Key Responsibilities :

1. AI Platform Reliability :

- Ensure high availability, scalability, and reliability of AI/ML platforms and production environments.

- Build resilient systems capable of supporting enterprise-scale AI workloads.

- Improve operational stability through proactive monitoring and automation.

2. MLOps & AI Operations :

Drive end-to-end MLOps lifecycle, including :

1. Model experimentation

2. Training

3. Validation

4. Deployment

5. Monitoring

6. Retirement

- Maintain model versioning, reproducibility, governance, and audit readiness.

- Collaborate closely with Data Science teams to productionize machine learning models efficiently.

3. Model Monitoring & Governance :

- Monitor model performance to identify:

1. Model drift

2. Data drift

3. Bias

4. Accuracy degradation

5. Data quality issues

- Implement governance standards and ensure compliance across AI solutions.

4. Production Support & Site Reliability :

- Provide L2/L3 production support for enterprise Java-based applications.

- Investigate and resolve production incidents with minimal business impact.

- Perform root cause analysis (RCA), defect resolution, and continuous service improvements.

- Debug application failures and optimize system performance.

5. Automation & Platform Engineering :

- Develop scripts and automation to reduce manual operational effort.

- Improve deployment, monitoring, and recovery processes through automation.

- Support platform reliability initiatives across cloud environments.

6. Infrastructure & Operations :

- Administer and troubleshoot Linux and Windows environments.

- Support enterprise applications running on cloud infrastructure.

- Work with observability tools to monitor application health and infrastructure performance.

Required Skills :

- 8 to 10 years of experience in Site Reliability Engineering (SRE), Production Support, or Platform Engineering.

- Strong experience supporting Java-based enterprise applications.

- Good knowledge of SQL, PostgreSQL, and PL/SQL.

- Hands-on experience with Python scripting and automation.

- Experience troubleshooting applications on Linux and Windows platforms.

- Strong debugging, defect analysis, and production incident management skills.

- Knowledge of cloud platforms, preferably Google Cloud Platform (GCP).

- Experience implementing automation for operational processes.

- Strong analytical, troubleshooting, and problem-solving skills.

- Excellent communication and stakeholder management abilities.

Preferred Skills:

- Prior experience in the Investment Banking or Financial Services domain.

- Knowledge of MLOps frameworks and AI platform operations.

- Experience with cloud observability and monitoring solutions.

- Familiarity with Autosys job scheduling.

- Exposure to AI governance, compliance, and model lifecycle management.

Key Technologies :

- Java, Python, PostgreSQL, PL/SQL, Google Cloud Platform (GCP), Linux, Windows.

Good to Have :

- Autosys, Cloud Observability Tools, MLOps Platforms.

Why Join Us :

- Work on enterprise-scale AI and Machine Learning platforms.

- Build reliable, resilient, and cloud-native AI infrastructure.

- Collaborate with AI Engineers, Data Scientists, Platform Engineers, and DevOps teams.

- Drive automation, observability, and operational excellence across mission-critical systems.

- Be part of innovative AI transformation programs in a fast-paced enterprise environment.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...