Posted on: 11/09/2026
Data Engineer (AWS & Databricks)
Experience : 3 - 5 Years
Role Summary :
We are looking for a Data Engineer with 3 - 5 years of experience specializing in Databricks and AWS to design, build, and optimize scalable data pipelines for analytical and BI workloads.
The ideal candidate brings hands-on expertise in PySpark, Delta Lake, and cloud-native integration between AWS and Databricks.
Key Responsibilities :
- Develop and maintain batch and streaming ETL/ELT data pipelines using Databricks and PySpark on AWS.
- Implement Lakehouse architectures using Delta Lake on Amazon S3 following the Medallion architecture (Bronze/Silver/Gold).
- Orchestrate data workflows using Databricks Workflows, AWS Step Functions, or Apache Airflow (MWAA).
- Write, tune, and optimize complex SQL queries for Amazon Redshift, Athena, and Databricks SQL.
- Configure secure data ingestion, storage, and cataloging leveraging AWS IAM, S3, and Databricks Unity Catalog.
- Build automated data validation checks and unit tests to ensure high data quality and pipeline reliability.
- Implement CI/CD deployment workflows using Git, Databricks Asset Bundles (DABs), or AWS CI/CD tools.
- Collaborate with data analysts, data scientists, and downstream stakeholders to deliver performant data models.
Technical Skills Required :
- Big Data & Compute: Databricks, Apache Spark, PySpark, Delta Lake
- AWS Services: Amazon S3, Redshift, Athena, IAM, CloudWatch
- Programming & Querying: Python, PySpark, Advanced SQL
- Orchestration: Databricks Workflows, Apache Airflow, or AWS Step Functions
- Streaming & Messaging: Apache Kafka, AWS Kinesis, or Spark Structured Streaming
- Data Governance & Formats: Unity Catalog, Parquet, Delta format
- Version Control & DevOps: Git, GitHub/GitLab, Docker basics
Good to Have :
- Databricks Certified Data Engineer Associate/Professional certification.
- Familiarity with AWS Glue, EMR, or DynamoDB.
- Experience with dbt (data build tool) for SQL transformations.
- Exposure to Generative AI integration, vector databases, or AWS Bedrock.
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1670801