Posted on: 24/09/2026
Job Description:
About the Role:
We are looking for a Senior Data Engineer to design, build, and operate enterprise-scale data and AI platforms on Microsoft Azure.
You will own the lakehouse and ETL/ELT layer end to end using Microsoft Fabric, Azure Databricks, and Azure Data Factory.
You will also package and ship data services, pipelines, and AI workloads as production-grade Docker containers.
The role suits someone who combines strong data engineering fundamentals with a DevOps mindset and is comfortable extending into Generative AI and RAG solutions.
Key Responsibilities:
- Design and deliver scalable, metadata-driven ETL/ELT pipelines using Azure Data Factory, Fabric Data Factory, Databricks Workflows, and Lakeflow Jobs.
- Build and maintain lakehouse platforms using Medallion Architecture (Bronze, Silver, Gold) on OneLake, ADLS Gen2, and Delta Lake.
- Develop high-quality PySpark, Spark SQL, Python, and T-SQL transformations covering incremental loads, CDC, data-quality rules, and business-rule validation.
- Optimize Spark and Delta Lake workloads using AQE, partitioning, join strategies, Z-ORDER, Liquid Clustering, data skipping, and file compaction.
- Containerize data applications, ingestion services, APIs, and pipeline components using Docker.
- Own the full container lifecycle: build, tag, scan, publish, deploy, and monitor.
- Author optimized, secure multi-stage Dockerfiles and docker-compose setups for local development, testing, and integration.
- Manage container images in Azure Container Registry (ACR) and deploy to Azure Kubernetes Service (AKS) or Azure Container Apps.
- Build CI/CD pipelines in Azure DevOps for automated container builds, testing, vulnerability scanning, and release.
- Use Fabric Deployment Pipelines and Databricks Asset Bundles for platform deployments.
- Implement data governance and security using Unity Catalog, Microsoft Entra ID, RBAC, Managed Identity, and Azure Key Vault.
- Build RAG pipelines and AI Agent workflows (document parsing, chunking, embeddings, vector search, model serving), and deploy agent and retrieval services as containerized microservices.
- Enable self-service analytics through SQL Endpoints, Power BI datasets, and natural-language interfaces such as Databricks Genie.
- Own production support: SLA monitoring, incident troubleshooting, root-cause analysis, and performance tuning.
- Mentor junior engineers, conduct code reviews, and contribute to engineering standards and best practices.
Required Skills & Experience:
Data Engineering Core:
- 8+ years in data engineering, ETL/ELT, and data warehousing.
- Strong hands-on experience with Azure Databricks (Delta Lake, Auto Loader, Unity Catalog, Lakeflow/DLT, Databricks Workflows).
- Strong hands-on experience with Microsoft Fabric (OneLake, Lakehouse, Data Warehouse, Fabric Data Factory, Notebooks, Dataflows Gen2, SQL Endpoint).
- Expert in Azure Data Factory, Azure Synapse Analytics, and ADLS Gen2.
- Advanced proficiency in Python, PySpark, Spark SQL, and SQL/T-SQL, including stored procedures, query tuning, and performance optimization.
- Proven experience with Spark and Delta Lake performance tuning at scale.
- Solid grounding in dimensional modelling and data-quality frameworks.
Docker & Containerization (Strong, Hands-on):
- 3+ years of production experience with Docker, covering image design, layering, caching, and size and security optimization.
- Proficient in writing multi-stage Dockerfiles, docker-compose, and container networking, volumes, and environment/secret management.
- Experience containerizing Python and PySpark applications, data ingestion services, REST/FastAPI services, and AI/RAG microservices.
- Hands-on experience with Azure Container Registry (ACR), plus deployment on AKS or Azure Container Apps.
- Working knowledge of Kubernetes fundamentals (pods, deployments, services, config maps, secrets, scaling).
- Container security practices: base-image hardening, non-root users, image scanning (e.g., Trivy or Microsoft Defender for Cloud), and dependency management.
- Ability to integrate container builds and deployments into CI/CD pipelines.
- Experience with Databricks custom containers or containerized job runtimes is a strong plus.
DevOps & Governance:
- Azure DevOps, Git, and CI/CD pipelines for data and container workloads.
- Infrastructure-as-code exposure (Terraform or Bicep), Databricks Asset Bundles, and Fabric Deployment Pipelines.
- Unity Catalog, Entra ID, RBAC, Managed Identity, and Key Vault.
Good to Have:
- Generative AI and Agentic AI: RAG, Azure OpenAI / Azure AI Foundry, Azure AI Search, Databricks Mosaic AI Vector Search and Model Serving, LangChain/LangGraph, MLflow, and Model Context Protocol (MCP).
- Power BI and dashboard development.
- Experience in BFSI or financial-services data platforms.
- Exposure to SSIS/SSRS and legacy-to-cloud migration.
- Observability tooling for containers and pipelines (Azure Monitor, Log Analytics, Prometheus/Grafana).
Certifications (Preferred):
- Microsoft Certified: Fabric Data Engineer Associate (DP-700)
- Any of: Azure Data Engineer Associate (DP-203), Databricks Data Engineer Professional, Azure Developer (AZ-204), Certified Kubernetes Application Developer (CKAD)
Education:
Bachelor's degree in Engineering, Computer Science, or a related field.
Soft Skills:
- Strong ownership and end-to-end accountability for production platforms
- Clear communication with business and technical stakeholders
- Structured problem-solving and troubleshooting
- Mentoring and collaborative working style
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1674194