HamburgerMenu
hirist

Job Description

Position Overview:

The AI Observability Engineer will be instrumental in implementation of scalable, cloud-native solutions to meet the growing needs of our Data & Development team. The successful candidate will demonstrate the ability to abstract complexity and create reusable, scalable patterns that accelerate development. The AI Observability Engineer will build and maintain a robust framework to ensure the reliability and maintainability of DPR Construction's complex AI systems.

Responsibilities:

- Standardize observability practices across AI/ML and other development teams including logging, metrics, tracing, and model performance monitoring, ingesting data from multiple platforms.

- Lead hands-on implementation of automation-first DevOps and MLOps practices, enabling infrastructure-as-code and consistent, repeatable environment provisioning.

- Design and manage intelligent DataOps pipelines with automated data quality monitoring and anomaly detection.

- Deploy, maintain and monitor containerized ML workloads.

- Extend existing CI/CD pipelines to support automated infrastructure changes and ML workflows.

- Implement AI-driven data validation, schema and concept drift detection and metadata management.

- Establish governance frameworks for AI systems, including bias detection, explainability, and auditability.

- Extend existing Azure RBAC strategy by automating role and permission management to reduce manual intervention.

- Develop automated test suites for model performance, regression, edge cases and bias validation.

- Monitor model KPIs (accuracy, precision, recall, latency, calibration).

- Ensure reproducability of experiments and production models.

- Act as a technical point of contact for DevOps and MLOps practices, developing reusable patterns, documentation, and proof-of-concepts to drive adoption.

Qualifications:

- Bachelors degree in computer science, Data Science, Information Systems, or a related field.

- 5+ years of experience in DevOps, MLOps, Data Engineering, Software Engineering or Site Reliability Engineering.

- Strong understanding of cloud infrastructure and experience working with at least one major cloud provider, preferably Azure.

- Proficiency in at least one objected-oriented programming language, preferably python with hands-on experience in ml frameworks like TensorFlow, PyTorch or Scikit-learn.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...