HamburgerMenu
hirist

Job Description

Applied Data Scientist

Full-Time | 3 to 6 Years Experience | Hybrid

Role Overview :

We are looking for a passionate and highly skilled Applied Data Scientist with a strong foundation in machine learning, statistical modelling, and real-world problem solving. The ideal candidate will have 36 years of hands-on experience in building, fine-tuning, and deploying end-to-end data science solutions, including experience with Large Language Models (LLMs) and Generative AI techniques. You will be responsible for translating complex business challenges into scalable, production-grade AI/ML solutions that power our next-generation products and services.

Key Responsibilities :

Model Development & Research :

- Design and develop machine learning and deep learning models for NLP tasks including text classification, summarisation, sentiment analysis, and question answering.

- Fine-tune and adapt pre-trained LLMs (e.g., GPT, BERT, T5) to domain-specific use cases with a focus on performance and reliability.

Data Handling & Feature Engineering :

- Work with large-scale structured and unstructured datasets to build robust data pipelines, conduct exploratory data analysis, and engineer high-quality features.

- Ensure rigorous data preprocessing, augmentation, and validation to maximize model accuracy and generalization.

Algorithm Design & Optimisation :

- Collaborate with cross-functional teams to design and implement scalable algorithms optimised for production environments.

- Benchmark and continuously improve model performance using appropriate evaluation metrics and experimentation frameworks.

MLOps & Production Deployment :

- Own the end-to-end model lifecycle from experimentation and versioning to deployment, monitoring, and retraining.

- Build and maintain CI/CD pipelines for ML models using MLOps best practices and tooling (e.g., MLflow, Kubeflow, Airflow, SageMaker Pipelines).

- Deploy and scale machine learning models on cloud infrastructure (AWS, GCP, Azure), optimising for latency, throughput, and cost efficiency.

- Set up model monitoring, data drift detection, and automated alerting to ensure reliability in production.

Cross-Team Collaboration :

- Partner closely with product, engineering, and business teams to align AI/ML solutions with organisational goals.

- Translate ambiguous business problems into well-scoped data science problem statements with clear success criteria.

Research & Innovation :

- Stay current with state-of-the-art advancements in machine learning, NLP, and MLOps. Evaluate and adopt relevant new techniques into the team's workflow.

- Contribute to internal knowledge sharing and, where applicable, to external publications, technical blogs, or patents.

Mentoring & Knowledge Sharing :

- Mentor junior data scientists and ML engineers, providing guidance on modelling approaches, code quality, and production readiness.

Qualifications :

Education :

- Bachelor's or Master's degree from a premier institution in Computer Science, Statistics, Mathematics, Data Science, or a related quantitative field.

Experience :

- 3 to 6 years of professional experience in applied data science, machine learning engineering, or NLP.

- Proven track record of deploying ML models to production and managing their lifecycle end-to-end.

- Hands-on experience with training, fine-tuning, and deploying Large Language Models (LLMs) such as GPT, BERT, T5, and similar architectures.

- Strong understanding of transformer architectures, attention mechanisms, embeddings, and neural network fundamentals.

Technical Skills :

- Core Languages : Python

- ML Frameworks : TensorFlow, PyTorch, Hugging Face Transformers, Scikit-learn

- MLOps Tooling : MLflow, Kubeflow, Apache Airflow, SageMaker Pipelines, DVC model versioning, pipeline orchestration, CI/CD for ML

- Cloud Platforms : AWS, GCP, Azure model deployment, scaling, and cloud-native ML services

- NLP Techniques : Text generation, embeddings, sentiment analysis, summarisation, semantic search

- Production Observability : Model monitoring, data drift detection, A/B testing, performance benchmarking

- Data Engineering : SQL, Spark / PySpark for large-scale data processing

Nice to Have :

- Experience with vector databases (e.g., Pinecone, Weaviate, ChromaDB) for semantic search and RAG architectures.

- Familiarity with prompt engineering, RLHF, and LLM evaluation frameworks.

- Contributions to open-source ML projects or published research.

- Experience in a regulated industry (BFSI, healthcare) with awareness of responsible AI and model governance practices.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...