Posted on: 09/09/2026
About the Role:
We are seeking a Lead Product AI Data Engineer to take ownership of the development of data engineering and AI capabilities within our healthcare-focused product ecosystem.
The role is highly hands-on and suited for an engineer who can independently design and develop scalable data pipelines while contributing to AI-enabled features, machine learning workflows, and intelligent data applications.
You will collaborate with product, analytics, data science, and engineering teams to convert complex datasets into reliable product capabilities.
Key Responsibilities:
- Develop and maintain scalable data pipelines supporting AI-enabled healthcare and medical device products.
- Build data ingestion, transformation, processing, and integration workflows using Python and PySpark.
- Develop ETL/ELT pipelines and workflow orchestration using Apache Airflow.
- Implement data processing and storage solutions using Databricks, Snowflake, and Delta Lake.
- Write optimized SQL queries and develop data models for analytical and product workloads.
- Create reusable data pipelines for machine learning and AI use cases.
- Work with Data Scientists to prepare training datasets, implement feature engineering workflows, and integrate ML models into production applications.
- Support the implementation of Generative AI features using LLMs, RAG, embeddings, vector databases, and AI APIs.
- Contribute to the development of AI agents and intelligent workflows for data-driven product use cases.
- Build data preparation and retrieval workflows for enterprise AI applications.
- Implement data validation, quality checks, monitoring, and pipeline observability.
- Troubleshoot pipeline failures, data issues, performance bottlenecks, and production incidents.
- Deploy and operate data workloads on AWS or Azure cloud environments.
- Participate in technical design and architecture discussions and contribute practical implementation recommendations.
- Conduct code reviews and promote clean, maintainable, and scalable engineering practices.
- Work closely with Product Managers, Data Scientists, Analysts, and other engineers to understand product requirements and deliver solutions.
Technical Skills:
- Strong hands-on programming experience in Python.
- Good experience with PySpark and distributed data processing.
- Hands-on experience with Snowflake, Databricks, and Delta Lake.
- Experience building and managing Airflow workflows.
- Strong SQL skills and experience with databases such as PostgreSQL, Oracle, Snowflake, or Databricks.
- Working knowledge of AWS or Azure cloud platforms.
- Understanding of data engineering concepts including ETL/ELT, data modeling, data quality, pipeline optimization, and distributed processing.
- Exposure to Generative AI, LLMs, RAG, vector databases, embeddings, or AI agents.
- Understanding of machine learning data pipelines and basic MLOps concepts.
- Exposure to specification-driven development approaches such as SpecKit or OpenSpec is an advantage.
Preferred Domain Experience:
- Experience with Healthcare, Life Sciences, Pharmaceutical, Biotechnology, or Medical Device data is desirable.
- Awareness of data privacy, security, governance, and regulatory considerations is an advantage.
Qualifications:
- Bachelor's degree or equivalent qualification in Computer Science, Software Engineering, Data Engineering, Artificial Intelligence, or a related discipline.
- 5 - 7 years of experience in data engineering, software engineering, AI engineering, or related technical roles.
- Strong problem-solving skills with the ability to independently own technical deliverables.
- Good communication and collaboration skills with cross-functional teams.
Success in this Role:
- Success will be measured by the quality and reliability of data pipelines, successful delivery of AI-enabled product features, pipeline performance, production stability, engineering efficiency, and the ability to translate product requirements into scalable technical solutions.
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1669859