HamburgerMenu
hirist

Netscribes - Lead Data Engineer - Databricks & Cloud

NETSCRIBES DATA INSIGHTS
7 - 15 Years
Bangalore

Posted on: 02/06/2026

Job Description

About Netscribes :

Netscribes is a global leader in data, insights, and digital solutions, helping the worlds largest organizations accelerate growth and innovation. As a growth catalyst, we empower sales, marketing, product development, and strategy through a unique blend of domain expertise and technological capabilities. Our end-to-end solutions span data engineering, advanced analytics, AI, and intelligent automation - built to scale and adapt to dynamic business environments. We partner with clients across the implementation journey, aligning with their ecosystem to deliver actionable intelligence, operational efficiency, and competitive advantage.

LoB Overview :

Netscribes delivers integrated solutions across content, process, and technology to help global businesses scale efficiently and stay competitive in fast-moving markets. Our services are structured across three key areas:

Content Solutions :

We specialize in building, managing, and optimizing large-scale content ecosystems for global brands. Our expertise spans product data enrichment, taxonomy development, catalog and digital shelf management, content operations, and digital asset optimization. These solutions help improve product discoverability, streamline content workflows, and enhance customer experience across digital platforms.

Process Solutions :

Our research and insights-driven services support core business processes across marketing, product, customer experience, and operations. This includes market and competitive research, campaign analysis, customer intelligence, and content performance tracking. Through Information Management Services (IMS) and managed service models, we embed scalable, high-quality support directly into client workflows, ensuring agility, consistency, and speed.

Technology Solutions :

Complementing our content and process offerings, we bring advanced technology capabilities through data engineering, data analytics, and AI. From building unified data platforms and real-time dashboards to deploying intelligent automation and machine learning models, we help organizations modernize operations and unlock data-driven decisions at scale.

Role : Lead Data Engineer - Databricks & Cloud

Overview :

Location : Bengaluru, India (Work From Office) Shift IST - 9:30 AM to 6:30 PM


Experience : 7+ Years


Employment Type : Full-Time

About the Role :

We are looking for a seasoned Lead Data Engineer with 7+ years of deep, hands-on expertise across the full data engineering spectrum. The ideal candidate will bring strong command over Databricks, PySpark, Spark architecture, SQL, and Python, combined with the ability to architect and deliver end-to-end ETL/ELT solutions at enterprise scale.


This role requires a practitioner who has led data migration programmes, governed data assets across the Databricks platform - from ingestion through data quality, validation, semantic layers, and Medallion Architecture and who can confidently manage stakeholder relationships and mentor junior engineers. If you have built production-grade lakehouses, driven migrations from legacy stacks to Databricks, and can bridge the gap between engineering rigour and business outcomes, we want to hear from you.

Mandatory Skills :

- Python / PySpark advanced proficiency in distributed data processing, performance optimisation, and reusable framework development

- SQL / Spark SQL from fundamental queries to complex multi-level transformations, window functions, and query plan optimisation

- Databricks Spark Architecture : deep understanding of DAGs, shuffle optimisation, partitioning strategies, caching, and cluster tuning

- Databricks Platform : Unity Catalog, Delta Lake, Delta Live Tables, Auto Loader, Workflows, Model Serving, and Databricks Asset Bundles

- Medallion Architecture : Bronze / Silver / Gold layer design, data flow patterns, and governance at each layer

- ETL / ELT Solution Architecture - end-to-end pipeline design, incremental loads, CDC, SCD strategies, and error handling frameworks

- Data Modelling : dimensional modelling, data vault, star/snowflake schema, and semantic layer design

- Data Migration Leadership : planning and executing large-scale migrations from legacy platforms to Databricks / cloud lakehouse

- Data Governance on Databricks - Unity Catalog policies, data lineage, tagging, access controls, and audit trails

- Data Quality & Validation : building automated DQ frameworks, expectation checks, reconciliation, and alerting pipelines

- Semantic Layer Development : building logical data models and metrics layers consumable by BI and analytics tools

- Azure Data Engineering : Azure Data Factory, ADLS Gen2, Event Hubs, and Azure Monitor integration

- Team Leadership : mentoring engineers, conducting code reviews, and driving engineering best practices across a migration programme

- Stakeholder Management : gathering requirements, presenting architectural decisions, and communicating project status to technical and business audiences

Good to Have Skills :

- Azure Synapse Analytics, Microsoft Fabric

- AWS Data Engineering: Glue, S3, Redshift, or Lake Formation

- dbt (Data Build Tool) for transformation layer management

- Apache Airflow, Prefect, or Dagster for pipeline orchestration

- Streaming / real-time pipelines: Spark Structured Streaming, Kafka, or Azure Event Hubs

- BI tools : Power BI, Tableau, Qlik, or Looker - especially semantic model integration

- CI/CD for data pipelines : GitHub Actions, Azure DevOps, or Databricks Asset Bundles

- Certifications : Databricks Data Engineer Professional, DP-203 Azure Data Engineer, or equivalent

Key Responsibilities :

Spark & Databricks Engineering :

- Design, develop, and optimise large-scale PySpark and Spark SQL workloads on Databricks for batch and streaming use cases

- Build and manage Delta Lake tables with proper partitioning, Z-ordering, vacuuming, and schema evolution strategies

- Implement Auto Loader and Delta Live Tables for reliable, scalable data ingestion and transformation pipelines

- Configure and tune Databricks clusters - instance types, autoscaling, Photon acceleration, and spot strategies - to optimise cost and performance

- Manage Databricks Workflows and orchestrate multi-task jobs with dependency management, retry logic, and alerting

ETL / ELT Architecture & Solution Delivery :

- Architect end-to-end ETL/ELT solutions covering ingestion, transformation, data quality, and consumption layers

- Design Medallion Architecture (Bronze - Silver - Gold) with clear data contracts, SLAs, and governance policies at each layer

- Develop reusable, modular, and well-documented code frameworks in Python, PySpark, and SQL

- Implement CDC (Change Data Capture), incremental load patterns, and SCD handling for accurate and efficient data movement

- Integrate Azure Data Factory pipelines with Databricks for orchestrated, event-driven data workflows

Data Modelling & Semantic Layer :

- Design logical and physical data models using dimensional, data vault, and OBT (One Big Table) approaches based on use case requirements

- Build semantic layers and metrics definitions on top of Gold layer datasets, consumable by Power BI, Tableau, or other BI tools

- Collaborate with BI and analytics teams to ensure data models are optimised for query performance and self-service reporting

Data Governance, Quality & Validation :

- Govern data assets using Databricks Unity Catalog - define and enforce data access policies, lineage tracking, tagging, and audit logging

- Build automated data quality frameworks with expectation checks, threshold alerting, and reconciliation reporting

- Implement data validation pipelines to ensure completeness, accuracy, consistency, and timeliness across all data products

- Define and enforce data contracts between upstream producers and downstream consumers

Data Migration Programme Leadership :

- Lead end-to-end migration of data workloads from legacy platforms (on-premises warehouses, legacy ETL tools) to Databricks Lakehouse

- Define migration strategies, cutover plans, rollback procedures, and data validation gates

- Mentor and guide junior and mid-level engineers through complex migration tasks, conducting regular code reviews and technical sessions

- Track migration progress, manage risks and blockers, and provide transparent status updates to project and business stakeholders

Stakeholder Management & Collaboration :

- Engage with business stakeholders, data owners, and product teams to gather requirements and translate them into technical solutions

- Articulate architectural decisions, trade-offs, and data engineering concepts clearly to both technical and non-technical audiences

- Provide regular project status updates, proactively identify risks, and propose mitigation strategies

- Champion data engineering best practices, coding standards, and documentation within the team

Technical Expertise Area :

- Technologies / Skills : Databricks Delta Lake, Delta Live Tables, Unity Catalog, Auto Loader, Workflows, Photon, Asset Bundles

- Languages : Python, PySpark, SQL, Spark SQL

- Spark Architecture : DAG optimisation, partitioning, shuffle tuning, caching, cluster configuration

- ETL / ELT : Medallion Architecture, CDC, SCD, incremental loads, error handling frameworks

- Data Modelling : Dimensional Modelling, Data Vault, Star/Snowflake Schema, Semantic Layer

- Data Governance : Unity Catalog, lineage, tagging, access controls, data contracts, DQ frameworks

- Azure DE : Azure Data Factory, ADLS Gen2, Event Hubs, Azure Monitor

- Leadership : Team mentoring, code reviews, migration planning, stakeholder management

- Good to Have : dbt, Airflow, Synapse, AWS Glue, Power BI, Tableau, CI/CD for data pipelines

Other Requirements :

- 7+ years of experience in data engineering roles with progressive responsibility

- Demonstrated experience delivering end-to-end Databricks Lakehouse solutions in production environments

- Proven track record of leading data migration programmes from legacy platforms to cloud lakehouse architectures

- Experience mentoring engineers and driving technical excellence across a team

- Strong communication skills - written and verbal; ability to present complex technical concepts clearly to business stakeholders

- Strong problem-solving skills with a structured, detail-oriented approach to root cause analysis and solution design

- Ability to manage multiple workstreams, prioritise effectively, and deliver under deadlines

- Must be available to work from our Bengaluru office - Monday to Friday, 9:30 AM to 6:30 PM IST

Employment Type : Permanent

Work Mode : Hybrid

Working Days :
5 days (Sat-Sun week off)

Shift : Fixed Shift (9 hours)

Shift Timings : General Shift

Location : Bengaluru

Education : Graduates only

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...