HamburgerMenu
hirist

Guddge - MLOps Engineer

Guddge Tech
5 - 10 Years
Others

Posted on: 12/09/2026

Job Description

About the Role :

We're looking for an MLOps Engineer to support and evolve a production machine-learning platform that produces daily risk scores used by operational leaders to intervene before incidents occur. You'll own the reliability, deployment, and monitoring of a production ML pipeline that spans AWS (S3, Glue, SageMaker, Step Functions), and a fully automated Terraform + GitHub Actions delivery pipeline across multi environments.

This is a hands-on infrastructure-and-operations role. The model itself is trained offline; your focus is keeping the scoring pipeline dependable, observable, secure, and easy to change.

What You'll Do :

- Operate and improve the end-to-end inference pipeline (SAP HANA - AWS Glue - AWS SageMaker - AWS Step Functions - SAP HANA), keeping scheduled daily runs reliable.

- Own the infrastructure-as-code: maintain Terraform modules and configuration, manage remote state in Terraform Cloud, and promote changes through dev - qa - prd.

- Maintain and extend the CI/CD workflows in GitHub Actions.

- Keep the pipeline observable: structured JSON logging, CloudWatch Logs Insights queries, and Step Functions execution monitoring.

- Support model monitoring (drift detection) and data-quality checks; tune alert thresholds and triage failure alerts.

- Troubleshoot production incidents using established runbooks (database connectivity, SageMaker job failures, Glue write-back, scheduler issues) and drive root-cause fixes.

- Enforce security and compliance controls: encryption, least-privilege IAM, secrets handling, and no-PII-in-logs discipline.

- Collaborate with data scientists to deploy retrained models through the SageMaker Model Registry and cross-account promotion process.

Required Qualifications :

- 3+ years in MLOps supporting production systems.

- Strong AWS experience, especially the data/ML services.

- Production Terraform experience with a remote backend and multi-environment promotion.

- Solid Python for data processing and operational scripting.

- Experience operating CI/CD pipelines and diagnosing production failures from logs and metrics.

Technical Skillset Breakdown :

The percentages reflect the relative weight of each area for day-to-day success in this role.

1. AWS Cloud & ML Services: ~35% :

The core of the platform runs on AWS.

- AWS Glue - Spark/PySpark ETL jobs, Glue connections (VPC/JDBC), Data Catalog, crawlers, job bookmarks.

- Amazon SageMaker - Processing Jobs, Model Registry, cross-account model packages, execution roles.

- AWS Step Functions - orchestration, Amazon States Language (ASL), retry/error handling.

- Amazon EventBridge Scheduler - cron scheduling, timezone handling.

- Amazon S3 - bucket policies, versioning, server-side encryption, lifecycle.

- Amazon CloudWatch - Logs, Logs Insights, monitoring.

- AWS Lake Formation - fine-grained catalog access control.

- VPC / Security Groups / Prefix Lists - network connectivity to on-prem.

2. Infrastructure as Code (Terraform): ~20% :

Nearly all infrastructure is managed as code.

- Terraform module authoring and reuse.

- Terraform Cloud (remote state, workspaces, locking).

- Multi-environment configuration (.tfvars, per-env config files).

- Provider configuration (AWS, AWSCC).

- Managing drift, plan/apply safety, state troubleshooting.

3. CI/CD & Automation: ~15% :


Delivery is fully automated.

- GitHub Actions: workflow authoring, reusable composite actions, environment approvals.

- Sequential environment promotion (dev - qa - prd).

- Git and branch/PR workflows, GitHub CLI / AWS Kiro.

- Make for build automation.

4. Python & ML Tooling: ~10% :

Supporting the model code and tests.

- Python 3.8+ for inference/ETL scripts and operational tooling.

- pandas, pyarrow, scipy for data processing.

- boto3 for AWS automation.

- SHAP (model explainability), imbalanced-learn (SMOTE) - familiarity to support the model.

- Scikit-learn model artifacts (Random Forest, MinMaxScaler) and pickle handling.

- pytest + coverage (50% gate), Black, flake8, isort.

5. Docker & Containerization: ~10% :

Containers underpin local testing and reproducible builds.

- Building and running Docker images for development and CI.

- Containerized test execution (e.g., PySpark unit tests that don't run natively on Windows).

- Managing dependencies and reproducible environments across local and pipeline runs.

- Familiarity with container-based execution in AWS (Glue/SageMaker managed containers).

6. Observability, Reliability & Security: ~10% :

Keeping it production-grade.

- Structured JSON logging and a typed error hierarchy (config/data/model/inference/downstream).

- Model drift and data-quality monitoring; alert threshold tuning.

- Incident triage using runbooks; alerting via SNS/SES dispatcher integration.

- Security controls: encryption in transit/at rest, least-privilege IAM, no PII in logs.

- Data-quality validation.

Working Environment :

- This is an offshore role, you will be working from your base location. No travel required.

- Ability to work US Pacific Standard Time Zone (i.e. 8 pm IST to 5 am IST).

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...