Posted on: 08/10/2026
Position : AI/GenAI Engineer (LLM Integration Specialist)
Experience : 2 - 4 years
CTC : 12 to 16 LPA
Location : Hyderabad
About the Role :
We're building a high-performance chat application and looking for an AI/GenAI Engineer to lead the integration and optimization of Large Language Models (LLMs). You'll be responsible for connecting LLM APIs, implementing domain-specific fine-tuning strategies, prompt engineering, and ensuring optimal performance for production use.
Key Responsibilities :
LLM Integration & Architecture :
- Integrate multiple LLM APIs (OpenAI, Anthropic Claude, Google Gemini, or open-source models)
- Design and implement robust API wrapper services with retry logic, fallback mechanisms, and error handling
- Implement streaming responses for real-time chat experience
- Build rate limiting and quota management systems
- Handle token counting, context window management, and cost optimization
Domain Customization & Fine-tuning :
- Develop domain-specific prompt engineering strategies
- Implement RAG (Retrieval Augmented Generation) pipelines using vector databases
- Fine-tune or adapt models for specific use cases using techniques like LoRA, prompt tuning
- Create and maintain knowledge bases for domain-specific responses
Performance & Optimization :
- Optimize API response times and reduce latency
- Implement caching strategies for common queries
- Monitor and optimize token usage to control costs
Safety & Quality :
- Implement content moderation and safety filters
- Build guardrails to prevent prompt injection and jailbreaking
- Develop evaluation frameworks to measure response quality
Technical Stack :
- Languages : Python (primary), JavaScript/TypeScript (basic understanding)
- LLM APIs : OpenAI, Anthropic, Google Gemini, Cohere
- Frameworks : LangChain, LlamaIndex, FastAPI
- Vector DBs : Pinecone, Weaviate, Qdrant, or ChromaDB
- Infrastructure : Docker, Redis, PostgreSQL, Message Queues
- Cloud : AWS/GCP/Azure
Infrastructure :
- Design scalable architecture for handling concurrent LLM requests
- Implement queue systems for managing high-volume API calls
- Set up monitoring and logging for LLM interactions
- Work with DevOps to deploy models (if self-hosted)
Required Skills & Experience :
Must Have :
- Experience working with LLMs and GenAI technologies
- Strong experience with OpenAI API, Anthropic Claude, or similar LLM APIs
- Proficiency in Python (FastAPI, LangChain, LlamaIndex preferred)
- Strong understanding of prompt engineering techniques and best practices
- Experience with vector databases (Pinecone, Weaviate, Qdrant, ChromaDB)
- Knowledge of RAG (Retrieval Augmented Generation) implementation
- Understanding of transformer architecture and attention mechanisms
- Experience with API integration, webhooks, and streaming responses
- Strong problem-solving skills and ability to debug complex AI systems
- Experience with LangChain, LlamaIndex, or similar LLM frameworks
- Knowledge of fine-tuning techniques (LoRA, QLoRA, PEFT)
- Experience with embedding models and semantic search
- Familiarity with HuggingFace Transformers library
- Experience deploying models using vLLM, TGI (Text Generation Inference)
- Knowledge of function calling/tool use with LLMs
- Experience with model evaluation metrics (BLEU, ROUGE, BERTScore)
- Understanding of token economics and cost optimisation
- Experience with open-source models (Llama, Mistral, Falcon)
- Knowledge of model quantization and optimization techniques
- Experience with multi-modal models (vision, audio)
- Familiarity with MLOps practices and experiment tracking (Weights & Biases, MLflow)
- Experience with AWS SageMaker, Google Vertex AI, or Azure ML
- Understanding of chain-of-thought prompting, ReAct, agents
- Experience building chatbots or conversational AI systems
- Publications or contributions to AI/ML community
Did you find something suspicious?