Posted on: 26/08/2026
Position Overview:
The AI Observability Engineer will be instrumental in implementation of scalable, cloud-native solutions to meet the growing needs of our Data & Development team. The successful candidate will demonstrate the ability to abstract complexity and create reusable, scalable patterns that accelerate development. The AI Observability Engineer will build and maintain a robust framework to ensure the reliability and maintainability of DPR Construction's complex AI systems.
Responsibilities:
- Standardize observability practices across AI/ML and other development teams including logging, metrics, tracing, and model performance monitoring, ingesting data from multiple platforms.
- Lead hands-on implementation of automation-first DevOps and MLOps practices, enabling infrastructure-as-code and consistent, repeatable environment provisioning.
- Design and manage intelligent DataOps pipelines with automated data quality monitoring and anomaly detection.
- Deploy, maintain and monitor containerized ML workloads.
- Extend existing CI/CD pipelines to support automated infrastructure changes and ML workflows.
- Implement AI-driven data validation, schema and concept drift detection and metadata management.
- Establish governance frameworks for AI systems, including bias detection, explainability, and auditability.
- Extend existing Azure RBAC strategy by automating role and permission management to reduce manual intervention.
- Develop automated test suites for model performance, regression, edge cases and bias validation.
- Monitor model KPIs (accuracy, precision, recall, latency, calibration).
- Ensure reproducability of experiments and production models.
- Act as a technical point of contact for DevOps and MLOps practices, developing reusable patterns, documentation, and proof-of-concepts to drive adoption.
Qualifications:
- Bachelors degree in computer science, Data Science, Information Systems, or a related field.
- 5+ years of experience in DevOps, MLOps, Data Engineering, Software Engineering or Site Reliability Engineering.
- Strong understanding of cloud infrastructure and experience working with at least one major cloud provider, preferably Azure.
- Proficiency in at least one objected-oriented programming language, preferably python with hands-on experience in ml frameworks like TensorFlow, PyTorch or Scikit-learn.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
ML / DL Engineering
Job Code
1666253