Posted on: 04/09/2026
The Role :
We are looking for a Data Scientist who wants to work on some of the hardest and most interesting problems in healthcare AI.
You will work with large-scale, longitudinal EHR data to develop models across areas such as :
- Clinical prediction and risk stratification
- Patient journey and disease progression modelling
- Treatment and medication intelligence
- Clinical outcome prediction
- Doctor and patient behaviour modelling
- Healthcare recommendation systems
- Medical NLP and clinical language understanding
- Data quality and clinical entity modelling
- Features and datasets for foundation and generative AI models
You will work closely with ML engineers, clinicians, product managers and data engineers to take ideas from research to production.
What You Will Do :
1. Build healthcare models :
- Design, train and evaluate statistical and machine learning models using real-world clinical data.
2. Work with longitudinal EHR data :
- Understand how diagnoses, symptoms, medications, investigations and clinical encounters evolve over time and translate these patterns into meaningful features and models.
3. Solve ambiguous problems :
- Turn open-ended healthcare questions into measurable modelling problems. Decide what data is required, how the target should be defined and how success should be evaluated.
4. Develop clinical intelligence :
- Identify patterns in patient journeys that can help predict risks, outcomes, adherence, treatment response and disease progression.
5. Work with unstructured clinical data :
- Build pipelines and models using clinical notes, prescriptions, medical terminology and other unstructured healthcare information.
6. Experiment rapidly :
- Develop hypotheses, build prototypes, run experiments and iterate quickly. We value strong problem-solving and scientific thinking over simply applying standard modelling techniques.
7. Build production-ready solutions :
- Work with engineering teams to take models from notebooks into scalable production systems and continuously monitor their performance.
8. Establish rigorous evaluation :
- Healthcare models need more than good offline metrics. You will design evaluation frameworks that consider clinical relevance, bias, calibration, robustness and real-world performance.
What We're Looking For :
- 2 - 6 years of experience in Data Science, Machine Learning, Applied Statistics or a related field
- Strong understanding of machine learning and statistical modelling
- Strong Python and SQL skills
- Experience with libraries such as Pandas, NumPy, Scikit-learn, PyTorch or equivalent
- Strong understanding of model evaluation, experimentation and statistical reasoning
- Experience working with large and messy real-world datasets
- Ability to translate business or domain problems into well-defined modelling problems
- Strong analytical and problem-solving skills
- Ability to communicate technical findings clearly to non-technical stakeholders
Strong Plus :
Experience in one or more of the following will be highly valuable :
- Healthcare, EHR or claims data
- Clinical NLP / medical language models
- Time-series or longitudinal modelling
- Survival analysis
- Causal inference
- Recommendation systems
- Risk prediction
- Generative AI / LLMs
- Medical terminology and ontology systems
- Large-scale data processing using Spark or similar technologies
- Building and deploying ML models in production
What Makes This Role Different :
- Real clinical data. You will work with longitudinal healthcare data generated through actual clinical workflows, not synthetic datasets or toy problems.
- Hard problems. Healthcare data is noisy, incomplete, heterogeneous and highly contextual. Building useful models requires going beyond standard ML approaches.
- Large opportunity. We are building a foundation for healthcare AI across millions of patient journeys and millions of clinical interactions.
- Research meets product. You will have the freedom to explore new modelling approaches while also seeing your work deployed into products used by doctors.
- Clinical impact. The ultimate measure of our models is not just AUC or F1. It is whether they can improve clinical decision-making and healthcare outcomes.
Our Ideal Candidate :
You are someone who gets excited when you see a messy dataset and immediately start asking :
- What is the underlying pattern?
- What could we predict from this?
- Is there a better way to represent this clinical journey?
- Can we build a model that actually generalises to the real world?
You enjoy going deep into a problem, challenging assumptions and building things that have never existed before.
If you want to work at the intersection of EHR data, machine learning, clinical intelligence and generative AI, this is the opportunity.
Did you find something suspicious?