Job Description :
Key Responsibilities :
Solutioning & Architecture :
- Translate client requirements into Databricks-based solution designs data pipelines, Lakehouse layouts, and serving layers.
- Recommend the right Databricks components (Delta Live Tables, Workflows, Unity Catalog, Databricks SQL, Genie) for a given use case, based on data volume, latency, and governance needs.
- Participate in pre-sales and proposal discussions, contributing effort estimates and technical approach for Databricks-based engagements.
- Review architecture decisions with senior architects and flag risks or better alternatives early.
Hands-On Build & Delivery :
- Build and maintain end-to-end pipelines : ingestion (Auto Loader, DLT), transformation (dbt or native PySpark/SQL), and serving (Unity Catalog, Databricks SQL).
- Work directly with pharma commercial datasets IQVIA, Symphony, CRM, Hub/SP, claims modeling them into clean, governed Delta Lake structures.
- Develop and maintain reusable components : notebooks, job templates, SQL libraries, and data quality checks.
- Configure and tune Genie Spaces and AI/BI dashboards for client-facing analytics use cases.
- Own workspace-level hygiene: cluster policies, job scheduling, cost tracking, and basic performance tuning.
Platform Currency & Best Practices :
- Stay closely tracked with new Databricks releases and features (e.g. Lakehouse/RT, Genie enhancements, Metric Views) and assess their relevance to pharma use cases.
- Bring new capabilities into existing client engagements where they create real value, not just for novelty.
- Contribute to and maintain DataZymes' internal Databricks standards, templates, and knowledge base.
- Support the certification and upskilling of junior engineers and analysts on the team.
Client & Team Collaboration :
- Act as the day-to-day Databricks technical point of contact on assigned client engagements.
- Explain technical trade-offs in plain terms to non-technical stakeholders when needed.
- Collaborate with analytics, forecasting, and delivery teams to make sure the platform serves the actual business question, not just the data movement.
Practice Building :
- Help establish the Databricks practice at DataZymes codifying reusable design patterns, reference architectures, and coding standards as the team's project count grows.
- Design and build solution accelerators for common pharma use cases (prescription analytics, patient cohort analysis, omnichannel attribution) that can be reused and adapted across clients.
- Maintain the internal Databricks knowledge base templates, checklists, and lessons learned from delivery.
- Support partnership conversations with Databricks by contributing technical input solution briefs, architecture references, and demo material that the practice lead and account teams can take into partner and client discussions.
- Help identify gaps in team capability and contribute to certification and enablement plans for engineers joining the practice.
Ideal Candidate :
- Strong hands-on Databricks Architect/Databricks Engineer Profile with end-to-end build-and-solution capability and Databricks professional certification.
Mandatory Experience :
- Must have 5+ years in Data engineering/Data architect roles, with at least recent 3 years of hands-on Databricks experience.
- Must be able to design a solution end-to-end and then build it themselves.
- Current-role project work must clearly align with the JD - the resume must describe the current project (what it is, the candidate's own scope, and the Databricks components used), not just list skills.
- Generic or JD-mirrored bullets without a concrete project will not be considered.
- Must have experience translating client/business requirements into Databricks solution designs - data pipelines, Lakehouse layouts, and serving layers.
- Must have built and maintained end-to-end pipelines - ingestion (Auto Loader, DLT), transformation (PySpark/SQL or dbt), and serving (Unity Catalog, Databricks SQL).
- Mandatory (Certification) : Must hold at least one active Databricks Professional-level certification (Data Engineer Professional preferred).
Mandatory (Tech skill 1) :
- Must have solid working knowledge across the Databricks stack - Delta Lake, Delta Live Tables, Unity Catalog, Auto Loader, Databricks SQL, Workflows, and cluster/job configuration.
- Must have strong SQL and PySpark skills, able to read and reason about existing pipelines quickly.
- Mandatory (Communication) : Must be able to act as the client-facing Databricks technical point of contact and explain technical trade-offs in plain terms to non-technical stakeholders.
- Mandatory (Company) : Must come from an IT services/consulting background with direct delivery on client engagements, US or global clients preferred.
- Mandatory (Stability) : Must show stable tenure : 2+ years average per employer, and no unexplained career gaps.
- Mandatory (Note 1) : Must be currently working hands-on on Databricks in their present role, not on an adjacent platform with past Databricks experience.
- Mandatory (Note 2) : CTC is inclusive of 20% variable.
- Mandatory (Note 3) : Role is Hybrid, WFH flexibility as well upto 6 days a month.
- Preferred (Domain) : Pharma or life sciences background.