Posted on: 12/09/2026
About the Role :
We're looking for an MLOps Engineer to support and evolve a production machine-learning platform that produces daily risk scores used by operational leaders to intervene before incidents occur. You'll own the reliability, deployment, and monitoring of a production ML pipeline that spans AWS (S3, Glue, SageMaker, Step Functions), and a fully automated Terraform + GitHub Actions delivery pipeline across multi environments.
This is a hands-on infrastructure-and-operations role. The model itself is trained offline; your focus is keeping the scoring pipeline dependable, observable, secure, and easy to change.
What You'll Do :
- Operate and improve the end-to-end inference pipeline (SAP HANA - AWS Glue - AWS SageMaker - AWS Step Functions - SAP HANA), keeping scheduled daily runs reliable.
- Own the infrastructure-as-code: maintain Terraform modules and configuration, manage remote state in Terraform Cloud, and promote changes through dev - qa - prd.
- Maintain and extend the CI/CD workflows in GitHub Actions.
- Keep the pipeline observable: structured JSON logging, CloudWatch Logs Insights queries, and Step Functions execution monitoring.
- Support model monitoring (drift detection) and data-quality checks; tune alert thresholds and triage failure alerts.
- Troubleshoot production incidents using established runbooks (database connectivity, SageMaker job failures, Glue write-back, scheduler issues) and drive root-cause fixes.
- Enforce security and compliance controls: encryption, least-privilege IAM, secrets handling, and no-PII-in-logs discipline.
- Collaborate with data scientists to deploy retrained models through the SageMaker Model Registry and cross-account promotion process.
Required Qualifications :
- 3+ years in MLOps supporting production systems.
- Strong AWS experience, especially the data/ML services.
- Production Terraform experience with a remote backend and multi-environment promotion.
- Solid Python for data processing and operational scripting.
- Experience operating CI/CD pipelines and diagnosing production failures from logs and metrics.
Technical Skillset Breakdown :
The percentages reflect the relative weight of each area for day-to-day success in this role.
1. AWS Cloud & ML Services: ~35% :
The core of the platform runs on AWS.
- AWS Glue - Spark/PySpark ETL jobs, Glue connections (VPC/JDBC), Data Catalog, crawlers, job bookmarks.
- Amazon SageMaker - Processing Jobs, Model Registry, cross-account model packages, execution roles.
- AWS Step Functions - orchestration, Amazon States Language (ASL), retry/error handling.
- Amazon EventBridge Scheduler - cron scheduling, timezone handling.
- Amazon S3 - bucket policies, versioning, server-side encryption, lifecycle.
- Amazon CloudWatch - Logs, Logs Insights, monitoring.
- AWS Lake Formation - fine-grained catalog access control.
- VPC / Security Groups / Prefix Lists - network connectivity to on-prem.
2. Infrastructure as Code (Terraform): ~20% :
Nearly all infrastructure is managed as code.
- Terraform module authoring and reuse.
- Terraform Cloud (remote state, workspaces, locking).
- Multi-environment configuration (.tfvars, per-env config files).
- Provider configuration (AWS, AWSCC).
- Managing drift, plan/apply safety, state troubleshooting.
3. CI/CD & Automation: ~15% :
Delivery is fully automated.
- GitHub Actions: workflow authoring, reusable composite actions, environment approvals.
- Sequential environment promotion (dev - qa - prd).
- Git and branch/PR workflows, GitHub CLI / AWS Kiro.
- Make for build automation.
4. Python & ML Tooling: ~10% :
Supporting the model code and tests.
- Python 3.8+ for inference/ETL scripts and operational tooling.
- pandas, pyarrow, scipy for data processing.
- boto3 for AWS automation.
- SHAP (model explainability), imbalanced-learn (SMOTE) - familiarity to support the model.
- Scikit-learn model artifacts (Random Forest, MinMaxScaler) and pickle handling.
- pytest + coverage (50% gate), Black, flake8, isort.
5. Docker & Containerization: ~10% :
Containers underpin local testing and reproducible builds.
- Building and running Docker images for development and CI.
- Containerized test execution (e.g., PySpark unit tests that don't run natively on Windows).
- Managing dependencies and reproducible environments across local and pipeline runs.
- Familiarity with container-based execution in AWS (Glue/SageMaker managed containers).
6. Observability, Reliability & Security: ~10% :
Keeping it production-grade.
- Structured JSON logging and a typed error hierarchy (config/data/model/inference/downstream).
- Model drift and data-quality monitoring; alert threshold tuning.
- Incident triage using runbooks; alerting via SNS/SES dispatcher integration.
- Security controls: encryption in transit/at rest, least-privilege IAM, no PII in logs.
- Data-quality validation.
Working Environment :
- This is an offshore role, you will be working from your base location. No travel required.
- Ability to work US Pacific Standard Time Zone (i.e. 8 pm IST to 5 am IST).
Did you find something suspicious?
Posted by
Recruiter
Last Active: NA as recruiter has posted this job through third party tool.
Posted in
DevOps / SRE
Functional Area
ML / DL Engineering
Job Code
1671015