HamburgerMenu
hirist

Observability Architect - AIOps & Data Science

INFINITES HR SERVICES
9 - 14 Years
rupee27-35 LPA
Hyderabad

Posted on: 05/10/2026

Job Description

Position Overview :

We are seeking an experienced Observability Architect to join our platform engineering and reliability team focused on developing intelligent systems for AIOps and Data Science. In addition to building machine learning models for log and telemetry data, this role will contribute to product strategy, collaborate closely with engineering teams, and help shape the architecture of our observability and operational intelligence platform.

Key Responsibilities :

- Design and build machine learning models for error detection, anomaly detection, and root cause analysis from system logs and metrics.

- Develop log correlation algorithms identifying relationships between disparate log entries across distributed systems.

- Build and maintain knowledge graphs representing system dependencies, service relationships, and failure patterns.

- Create automated incident analysis pipelines ingesting raw logs, correlating events, and suggesting root causes in real time.

- Collaborate with platform engineers to design scalable, resilient telemetry ingestion, aggregation, and analytics pipelines supporting real-time operational intelligence.

- Contribute to defining platform standards, technology selections, and engineering practices aligned with the long-term product vision.

- Partner with SRE, engineering, and product teams to translate enterprise observability challenges into scalable, AI-powered platform solutions.

- Participate in design reviews and help optimize platform reliability, performance, security, automation, and scalability.

- Mentor and guide junior data scientists and cross-functional teams in applied ML and observability techniques.

- Document model methodologies and assumptions and create dashboards for model performance and stakeholder visibility.

- Stay abreast of emerging observability technologies, AI workflows, and telemetry-driven operational intelligence trends.

- Support the development of backend services or APIs in languages such as Python, Java, or .NET to integrate ML components with platform services.

Required Skills & Qualifications :

- Bachelor's degree in Data Science, Computer Science, Statistics, or a related field, or equivalent experience.

- 3+ years of professional experience in data science, machine learning engineering, or observability analytics.

- Strong programming skills in Python (preferred), with optional experience in Java, Go, or .NET for production code integration.

- Experience developing ML models using frameworks such as TensorFlow, PyTorch, or scikit-learn.

- Expertise in NLP techniques for log analysis and time-series anomaly detection.

- Familiarity with distributed computing environments such as Spark, Flink, or Kafka.

- Knowledge of knowledge graph technologies and graph neural networks, such as Neo4j, PyG, or DGL.

- Experience with telemetry and observability platforms such as ELK, Datadog, Splunk, Prometheus, or similar tools.

- Understanding of observability concepts, including metrics, logs, traces, dashboards, alerting, SLOs/SLIs, incident management, and operational analytics.

- Exposure to AI-powered operational intelligence, including agentic AI workflows, LLM-based assistants, graph-based dependency mapping, automated incident detection and triage, root cause analysis, and remediation recommendations.

- Strong problem-solving skills and the ability to communicate complex models to technical and non-technical stakeholders.

Preferred Qualifications :

- Master's degree in Machine Learning, Computer Science, or a related discipline.

- Experience in AIOps, Site Reliability Engineering, or IT Operations domains.

- Published research or contributions to open-source ML or observability projects.

- Knowledge of causal inference techniques for root cause analysis.

- Familiarity with containerization technologies, including Docker and Kubernetes, and CI/CD pipelines.

- Experience with incident management systems and on-call tooling.

- Background working with microservices architectures or cloud platforms; Azure experience is preferred.

- Awareness of SRE practices, self-healing automation, capacity prediction, and operational decision intelligence.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...