Description :
Notice period : 15 to 30days
Location : Bangalore
Exp : 6 to 9 years
Key Responsibilities :
- Design, build, and deploy AI/ML solutions using Large Language Models (LLMs) and Generative AI technologies.
- Fine-tune open-source and proprietary foundation models for domain-specific use cases.
- Work on model optimization techniques including quantization, pruning, distillation, and efficient inference.
- Develop Retrieval-Augmented Generation (RAG) pipelines using embeddings and vector databases.
- Optimize models for edge/on-device deployment on low-resource hardware.
- Collaborate with product, platform, and application teams to integrate AI capabilities into products.
- Evaluate model performance using appropriate benchmarks and metrics.
- Stay updated with the latest advancements in AI, GenAI, multimodal AI, and edge AI ecosystems.
- Mentor junior engineers and contribute to technical design discussions.
Essential Skills :
- Strong programming expertise in Python.
- Hands-on experience with :
1. LLM fine-tuning
2. Transformer architectures
3. PyTorch / TensorFlow
4. Embedding models and semantic search
5. Vector databases (FAISS, ChromaDB, Pinecone)
- Experience in model quantization and optimization techniques (INT8, 4-bit, GGUF, ONNX, TensorRT, llama.cpp, etc.
- Good understanding of RAG architectures and prompt engineering.
- Strong debugging, analytical, and problem-solving skills