Posted on: 02/06/2026
About the Role :
We are seeking a highly skilled Senior Data Engineer to join our team. This role focuses on designing, building, and maintaining scalable data pipelines, and machine learning solutions using Databricks and cloud technologies. You will lead efforts in orchestrating data workflows, optimizing cloud infrastructure, and developing production-grade data pipelines. A critical aspect of the role is ensuring strong data governance, security, and compliance, while seamlessly integrating cloud-native services.
As a Senior Data Engineer, you will collaborate with data scientists, product teams, and external engineering partners to define best practices in data engineering, infrastructure automation, and large-scale data processing. This position provides opportunities for leadership, creativity, and significant impact through data-driven product decisions and robust data solutions.
Key Responsibilities :
Data Engineering & Pipeline Development :
- Design, build, and optimize ETL/ELT pipelines for batch and real-time data processing.
- Implement data ingestion frameworks for multiple sources, including APIs, streaming platforms (Kafka, Kinesis), and third-party datasets.
- Develop high-performance, distributed data pipelines using PySpark, Delta Lake, and SQL.
- Perform schema design, normalization/denormalization, and data modeling for analytical and operational data stores.
- Implement data quality, lineage, and auditing mechanisms to ensure reliable and compliant datasets.
Databricks Platform Administration :
- Coordinate and manage Databricks workspaces, clusters, and workflows.
- Configure role-based access control (RBAC), manage user groups, entitlements, and integrate with identity providers like Okta.
- Optimize cluster configurations, auto-scaling, and cost management.
- Monitor performance, debug Spark job failures, and troubleshoot performance bottlenecks in notebooks and SQL queries.
- Maintain high availability and disaster recovery plans for critical data pipelines.
Cloud Platform Experience :
- Implement secure cloud environments using role-based access, encryption, and compliance frameworks.
- Integrate Databricks and other data engineering tools with cloud-native services to streamline data workflows.
- Design scalable, cost-effective data storage and processing architectures for structured and unstructured data.
Security, Governance & Compliance :
- Implement data security policies including encryption, masking, and tokenization.
- Ensure compliance with GDPR, HIPAA, SOC2, and other regulatory standards.
- Develop and maintain data cataloging, metadata management, and governance standards.
Automation, Orchestration & DevOps :
- Automate data pipelines, tasks, and monitoring using REST APIs, Python, Bash, or Terraform.
- Implement workflow orchestration with Apache Airflow, Prefect, or Databricks Workflows.
- Build CI/CD pipelines for data and ML code using Jenkins, Azure DevOps, or GitHub Actions.
- Monitor and alert on pipeline failures, data anomalies, and system health.
Machine Learning & Data Science Support :
- Collaborate with data scientists to build, deploy, and maintain ML models at scale.
- Integrate ML workflows into production data pipelines.
- Implement monitoring and retraining strategies for model performance and business relevance.
- Support experimentation, feature engineering, and scalable model training pipelines.
Collaboration :
- Establish and promote best practices for data engineering, pipeline development, and cloud infrastructure.
- Collaborate with product teams and external engineering partners to deliver high-impact solutions.
- Document technical decisions, architecture diagrams, and runbooks for operational excellence.
Must-Have Skills :
- 6+ years of professional experience in data engineering, or machine learning.
- Strong proficiency in Python (PySpark, Pandas) and SQL for large-scale data processing.
- Experience designing and maintaining production-grade ETL/ELT pipelines for batch and streaming data.
- Deep understanding of cloud platforms (AWS, Azure), infrastructure automation, and CI/CD for data workloads.
- Knowledge of containerization (Docker, Kubernetes) and orchestration for data and ML workloads.
- Hands-on experience with data cataloging, lineage tracking, and metadata management tools.
- Advanced skills in Spark performance tuning, Delta Lake optimization, and cloud cost management.
- Solid understanding of data governance, compliance, and security best practices.
- Strong communication and stakeholder management skills.
Preferred Qualifications :
- Familiarity with generative AI models and applying them to product-specific use cases.
- Expertise in anomaly detection systems, vector search, and embedding-based retrieval for unstructured data.
- Experience with streaming data platforms like Kafka, Kinesis, or Event Hubs.
- Experience in Databricks administration, Spark, and distributed data processing.
- Experience building and supporting ML pipelines and working closely with ML team
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1640952