Posted on: 07/07/2026
We are a fast-growing startup based in Pune, India, specializing in cutting-edge Data Science and Data Engineering solutions. Our team of dedicated professionals is committed to solving complex data challenges for companies worldwide.
Our Culture:
We foster a vibrant startup culture that values:
- Intellectual curiosity
- Continuous learning
- Positive work environment
- Collaborative problem-solving
Role Overview:
We are seeking a highly technical and proactive AI-ML Engineer to join our dynamic team. The ideal candidate will possess a strong blend of software engineering rigor and deep technical expertise in modern AI/ML architectures, infrastructure optimization, and generative AI systems. This role demands critical thinking, production-grade coding, and the ability to deploy, optimize, and scale robust AI models to solve complex, real-world problems. You should bring a deep curiosity for understanding model internals, and the adaptability required to navigate and implement rapidly evolving, state-of-the-art technologies.
Key Responsibilities:
- Deliver end-to-end AI engineering projects by designing, building, and deploying production-grade Machine Learning, Deep Learning, and Generative AI applications.
- Develop high-quality software solutions in Python, collaborating with cross-functional engineering teams to integrate AI models into existing application codebases.
- Implement advanced training strategies including mixed-precision training (FP16/BF16), gradient accumulation, and distributed training (Data/Model/Pipeline parallel) while profiling and maximizing GPU utilization.
- Optimize, serialize (ONNX, TorchScript), and deploy models via high-throughput REST APIs (FastAPI) while managing latency vs. throughput trade-offs.
- Implement robust MLOps workflows using DVC, Docker, and cloud platforms for model versioning, pipeline automation, and production monitoring.
- Architect and optimize high-performance Retrieval-Augmented Generation (RAG) systems using hybrid search, semantic chunking, and cross-encoder reranking.
- Design autonomous, multi-agent systems and orchestrated workflows utilizing function calling, the ReAct pattern, and advanced memory management architectures.
- Implement state-of-the-art serving techniques (vLLM, speculative decoding, prompt caching) and quantization (INT8/INT4 via GPTQ/AWQ) to maximize inference throughput.
- Move workflows from notebooks to production-grade pipelines, writing clean code and implementing unit/integration tests for ML (pytest, Great Expectations).
- Actively diagnose production anomalies, including data/concept distribution shifts, silent model failures, and training bottlenecks (underfitting/overfitting).
- Evaluate and benchmark LLM outputs using appropriate metrics and testing frameworks.
- Design high-throughput data pipelines and optimized SQL/NoSQL queries for large-scale data processing and model feature injection.
- Practice active listening to understand project requirements and team inputs.
- Collaborate with stakeholders to translate complex business requirements into scalable AI/ML solutions and communicate technical trade-offs clearly.
- Demonstrate strong technical communication skills, a high degree of ownership, and an action-biased approach to debugging and solving ambiguous engineering problems.
- Apply responsible AI principles, jailbreak awareness, and output validation guardrails to ensure ethical and safe model development.
- Plan strategically and multitask efficiently to meet project deadlines.
Required Skills:
Core Programming & ML:
- Strong Python programming skills with hands-on project experience, Git proficiency, and a basic understanding of CUDA.
- Expertise in Deep Learning architectures (Transformers, CNNs, RNNs) alongside a strong theoretical understanding of foundational ML algorithms (GBMs, Random Forests).
- Deep proficiency in PyTorch or TensorFlow, with a strong emphasis on custom layer implementation and neural network training loops.
- Hands-on experience with modern training paradigms, including Self-supervised Learning, Contrastive Learning, and advanced Transfer Learning.
- Experience with NLP, Computer Vision, or Time Series Analysis.
- Proven experience writing clean, production-grade code, utilizing pytest and Great Expectations for data and model validation.
Generative AI & LLMs:
- Hands-on experience with commercial LLM APIs (OpenAI, Anthropic, Groq) and hosting/deploying open-source foundational models (Llama, Mistral).
- Proficiency with modern orchestration frameworks (LangChain, LlamaIndex) and building stateful multi-agent architectures using LangGraph or DSPy.
- Experience with vector databases (Pinecone, Weaviate, Chroma, pgvector), advanced chunking (fixed, semantic, recursive), and hybrid search implementation.
- Deep technical understanding of Parameter-Efficient Fine-Tuning (PEFT) mechanicsspecifically LoRA/QLoRA low-rank decompositionalongside instruction tuning and model alignment methodologies (DPO, RLHF).
- Mastery of advanced prompt engineering, including structured output forcing, Chain-of-Thought, system prompt design, and building resilient multi-agent coordination systems.
- Hands-on experience with inference optimization and high-throughput serving, including quantization (GPTQ, AWQ), speculative decoding, prompt caching, and vLLM acceleration.
MLOps & Deployment:
- Experience with MLOps practices, logging, and model registries (MLflow, Weights & Biases, DVC) along with model serving via Triton Inference Server or FastAPI.
- Production experience deploying, scaling, and monitoring models natively on cloud AI platforms, specifically AWS SageMaker or GCP Vertex AI.
- Experience building CI/CD pipelines for ML applications, with an emphasis on data and model versioning/lineage using DVC and Git.
Data Engineering & Databases:
- Solid understanding of SQL, including advanced concepts like windowing functions and query optimization.
- Experience building and orchestrating data pipelines using Airflow or Prefect to feed specialized SQL and NoSQL vector/feature databases.
Soft Skills & Professional Attributes:
- Strong critical thinking and problem-solving skills.
- Excellent written and verbal communication abilities.
- High degree of flexibility and adaptability to stay ahead of the rapid velocity of the open-source AI ecosystem (Hugging Face, vLLM, etc.).
- Practical understanding of AI safety, compliance, guardrails (e.g., NeMo Guardrails or Llama Guard), and responsible AI practices.
Nice-to-Have:
- Contributions to open-source AI/ML repositories or GenAI orchestration frameworks.
- Experience with real-time streaming data processing.
- Active participation in competitive ML spaces (Kaggle) or track record of reviewing/reproducing state-of-the-art AI research papers.
- Published research papers or conference presentations.
- Experience building and scaling Graph Databases or Knowledge Graphs for advanced RAG.
- Experience with multimodal architectures (Vision-Language models, Audio processing).
Qualifications:
- AI-ML Engineer: 2-5 years of hands-on experience engineering machine learning systems, optimizing infrastructure, and implementing LLM/GenAI workflows in production.
- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Data Science, Statistics, or a related highly quantitative field.
- Demonstrated commitment to continuous learning through contributing to open-source, certifications, or self-study (especially in deep learning internals, MLOps, and modern GenAI frameworks).
What We Offer:
- Competitive salary commensurate with experience.
- Opportunity to work on diverse, cutting-edge AI/ML projects.
- Collaborative and innovation-driven work environment.
- Rapid growth and continuous learning opportunities.
- Exposure to latest AI technologies and industry best practices.
Did you find something suspicious?