HamburgerMenu
hirist

AI/ML Engineer

Jash Data Sciences
2 - 6 Years
Maharashtra

Posted on: 07/07/2026

Job Description

We are a fast-growing startup based in Pune, India, specializing in cutting-edge Data Science and Data Engineering solutions. Our team of dedicated professionals is committed to solving complex data challenges for companies worldwide.

Our Culture:

We foster a vibrant startup culture that values:

- Intellectual curiosity

- Continuous learning

- Positive work environment

- Collaborative problem-solving

Role Overview:

We are seeking a highly technical and proactive AI-ML Engineer to join our dynamic team. The ideal candidate will possess a strong blend of software engineering rigor and deep technical expertise in modern AI/ML architectures, infrastructure optimization, and generative AI systems. This role demands critical thinking, production-grade coding, and the ability to deploy, optimize, and scale robust AI models to solve complex, real-world problems. You should bring a deep curiosity for understanding model internals, and the adaptability required to navigate and implement rapidly evolving, state-of-the-art technologies.

Key Responsibilities:

- Deliver end-to-end AI engineering projects by designing, building, and deploying production-grade Machine Learning, Deep Learning, and Generative AI applications.

- Develop high-quality software solutions in Python, collaborating with cross-functional engineering teams to integrate AI models into existing application codebases.

- Implement advanced training strategies including mixed-precision training (FP16/BF16), gradient accumulation, and distributed training (Data/Model/Pipeline parallel) while profiling and maximizing GPU utilization.

- Optimize, serialize (ONNX, TorchScript), and deploy models via high-throughput REST APIs (FastAPI) while managing latency vs. throughput trade-offs.

- Implement robust MLOps workflows using DVC, Docker, and cloud platforms for model versioning, pipeline automation, and production monitoring.

- Architect and optimize high-performance Retrieval-Augmented Generation (RAG) systems using hybrid search, semantic chunking, and cross-encoder reranking.

- Design autonomous, multi-agent systems and orchestrated workflows utilizing function calling, the ReAct pattern, and advanced memory management architectures.

- Implement state-of-the-art serving techniques (vLLM, speculative decoding, prompt caching) and quantization (INT8/INT4 via GPTQ/AWQ) to maximize inference throughput.

- Move workflows from notebooks to production-grade pipelines, writing clean code and implementing unit/integration tests for ML (pytest, Great Expectations).

- Actively diagnose production anomalies, including data/concept distribution shifts, silent model failures, and training bottlenecks (underfitting/overfitting).

- Evaluate and benchmark LLM outputs using appropriate metrics and testing frameworks.

- Design high-throughput data pipelines and optimized SQL/NoSQL queries for large-scale data processing and model feature injection.

- Practice active listening to understand project requirements and team inputs.

- Collaborate with stakeholders to translate complex business requirements into scalable AI/ML solutions and communicate technical trade-offs clearly.

- Demonstrate strong technical communication skills, a high degree of ownership, and an action-biased approach to debugging and solving ambiguous engineering problems.

- Apply responsible AI principles, jailbreak awareness, and output validation guardrails to ensure ethical and safe model development.

- Plan strategically and multitask efficiently to meet project deadlines.

Required Skills:

Core Programming & ML:

- Strong Python programming skills with hands-on project experience, Git proficiency, and a basic understanding of CUDA.

- Expertise in Deep Learning architectures (Transformers, CNNs, RNNs) alongside a strong theoretical understanding of foundational ML algorithms (GBMs, Random Forests).

- Deep proficiency in PyTorch or TensorFlow, with a strong emphasis on custom layer implementation and neural network training loops.

- Hands-on experience with modern training paradigms, including Self-supervised Learning, Contrastive Learning, and advanced Transfer Learning.

- Experience with NLP, Computer Vision, or Time Series Analysis.

- Proven experience writing clean, production-grade code, utilizing pytest and Great Expectations for data and model validation.

Generative AI & LLMs:

- Hands-on experience with commercial LLM APIs (OpenAI, Anthropic, Groq) and hosting/deploying open-source foundational models (Llama, Mistral).

- Proficiency with modern orchestration frameworks (LangChain, LlamaIndex) and building stateful multi-agent architectures using LangGraph or DSPy.

- Experience with vector databases (Pinecone, Weaviate, Chroma, pgvector), advanced chunking (fixed, semantic, recursive), and hybrid search implementation.

- Deep technical understanding of Parameter-Efficient Fine-Tuning (PEFT) mechanicsspecifically LoRA/QLoRA low-rank decompositionalongside instruction tuning and model alignment methodologies (DPO, RLHF).

- Mastery of advanced prompt engineering, including structured output forcing, Chain-of-Thought, system prompt design, and building resilient multi-agent coordination systems.

- Hands-on experience with inference optimization and high-throughput serving, including quantization (GPTQ, AWQ), speculative decoding, prompt caching, and vLLM acceleration.

MLOps & Deployment:

- Experience with MLOps practices, logging, and model registries (MLflow, Weights & Biases, DVC) along with model serving via Triton Inference Server or FastAPI.

- Production experience deploying, scaling, and monitoring models natively on cloud AI platforms, specifically AWS SageMaker or GCP Vertex AI.

- Experience building CI/CD pipelines for ML applications, with an emphasis on data and model versioning/lineage using DVC and Git.

Data Engineering & Databases:

- Solid understanding of SQL, including advanced concepts like windowing functions and query optimization.

- Experience building and orchestrating data pipelines using Airflow or Prefect to feed specialized SQL and NoSQL vector/feature databases.

Soft Skills & Professional Attributes:

- Strong critical thinking and problem-solving skills.

- Excellent written and verbal communication abilities.

- High degree of flexibility and adaptability to stay ahead of the rapid velocity of the open-source AI ecosystem (Hugging Face, vLLM, etc.).

- Practical understanding of AI safety, compliance, guardrails (e.g., NeMo Guardrails or Llama Guard), and responsible AI practices.

Nice-to-Have:

- Contributions to open-source AI/ML repositories or GenAI orchestration frameworks.

- Experience with real-time streaming data processing.

- Active participation in competitive ML spaces (Kaggle) or track record of reviewing/reproducing state-of-the-art AI research papers.

- Published research papers or conference presentations.

- Experience building and scaling Graph Databases or Knowledge Graphs for advanced RAG.

- Experience with multimodal architectures (Vision-Language models, Audio processing).

Qualifications:

- AI-ML Engineer: 2-5 years of hands-on experience engineering machine learning systems, optimizing infrastructure, and implementing LLM/GenAI workflows in production.

- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Data Science, Statistics, or a related highly quantitative field.

- Demonstrated commitment to continuous learning through contributing to open-source, certifications, or self-study (especially in deep learning internals, MLOps, and modern GenAI frameworks).

What We Offer:

- Competitive salary commensurate with experience.

- Opportunity to work on diverse, cutting-edge AI/ML projects.

- Collaborative and innovation-driven work environment.

- Rapid growth and continuous learning opportunities.

- Exposure to latest AI technologies and industry best practices.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...