Posted on: 17/08/2026
Role Overview :
We are seeking an AWS AI/ML & GenAI Platform Engineer to design, deploy, and operate scalable, secure, production-grade AI/ML and Generative AI platforms on AWS. The role focuses on MLOps, LLMOps, DevOps, AI observability, and platform engineering supporting lending, risk, and customer-facing solutions.
Key Responsibilities :
- Operate and monitor production ML, LLM, RAG, and agentic AI workloads on AWS.
- Build end-to-end MLOps/LLMOps pipelines using SageMaker, Model Registry, CI/CD, and GenAI evaluation frameworks.
- Deploy and manage Amazon Bedrock, LLMs, RAG pipelines, vector databases, and AI agents.
- Design cloud infrastructure using Terraform/CloudFormation, Docker, Kubernetes, and Amazon EKS.
- Build CI/CD pipelines using AWS CodePipeline, CodeBuild, GitHub Actions, and/or Jenkins.
- Develop data and orchestration pipelines using AWS Glue, Lambda, Step Functions, S3, Redshift, and DynamoDB.
- Implement monitoring and observability using CloudWatch, Prometheus, Grafana, and ELK.
- Ensure AI platforms meet security, governance, Responsible AI, compliance, SLA/SLO, and reliability requirements.
- Develop reusable MLOps/LLMOps frameworks and collaborate with Data Scientists, AI Architects, and Platform Engineers.
Required Skills & Experience :
- Strong hands-on experience with AWS SageMaker and end-to-end ML lifecycle management.
- Experience with Amazon Bedrock, LLMs, Foundation Models, Generative AI, RAG, and AI/agentic workflows.
- Experience with MLOps/LLMOps, including ML pipelines, Model Registry, CI/CD, experiment tracking, prompt lifecycle management, and GenAI evaluation.
- Strong knowledge of AWS services including S3, Lambda, API Gateway, OpenSearch, Glue, Step Functions, Redshift, and DynamoDB.
- Hands-on experience with Docker, Kubernetes/EKS, Terraform, and/or CloudFormation.
- Experience with AWS CodePipeline, CodeBuild, GitHub Actions, and/or Jenkins.
- Strong Python and Bash/shell scripting skills; FastAPI preferred.
- Experience building and integrating REST APIs and enterprise AI platforms.
- Experience with CloudWatch, Prometheus, Grafana, and ELK for monitoring and observability.
- Understanding of AI/ML observability, model monitoring, logging, alerting, SLAs/SLOs, and incident management.
- Strong understanding of AWS IAM, KMS, Secrets Manager, security, data governance, privacy, and compliance.
- Knowledge of Responsible AI, model governance, risk management, and AI security.
- Experience designing scalable, highly available, resilient, secure, and cost-optimized AWS platforms.
- Experience collaborating with AI Architects, Data Scientists, ML Engineers, and Platform/DevOps teams.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
ML / DL Engineering
Job Code
1663642