Posted on: 16/04/2026
Key Responsibilities
- Lead end-to-end SLM lifecycle (data ingestion - dataset creation - fine-tuning - evaluation - deployment)
- Architect scalable ETL pipelines for large-scale unstructured data
- Design and build training datasets (instruction tuning, classification, entity extraction)
- Implement fine-tuning techniques (LoRA, QLoRA, PEFT)
- Define evaluation frameworks (accuracy, F1, domain-specific KPIs)
- Optimize models using quantization (INT4/INT8), inference tuning
- Drive system design for SLM-powered solutions (APIs, pipelines, infra)
- Mentor engineers and guide best practices
Must-Have Skills :
- Strong Python (data processing, ML pipelines)
- Hands-on experience in unstructured text processing (NLP)
- Experience designing ETL pipelines (batch/streaming)
- Proven experience in SLM/LLM fine-tuning
- Understanding of GPU memory management (VRAM), model loading, batching, and throughput optimization
- Experience with CUDA-enabled environments / GPU clusters (NVIDIA GPUs such as A100, L40S, T4, etc.)
- Familiarity with deployment on GPU-enabled platforms (Kubernetes / cloud GPU instances / on-prem setups)
- Experience with dataset creation (JSONL, instruction tuning)
- Familiarity with Hugging Face, PEFT, Unsloth or similar
- Experience with vector DBs (pgvector / milvus etc)
- Experience building APIs (FastAPI/Flask)
Good to Have :
- Experience in RAG pipelines
- Knowledge of model optimization (quantization, pruning)
- Exposure to Ollama / on-device inference / GGUF models
- Experience with evaluation frameworks (RAGAS / OpenAI evals)
- Domain experience (telecom, logs, enterprise ops)
- Exposure to cloud platforms (Azure / AWS / GCP)
Did you find something suspicious?