HamburgerMenu
hirist

Machine Learning Engineer - Speech Recognition

HuntingCube Recruitment Solution
3 - 7 Years
Remote

Posted on: 14/08/2026

Job Description

About the Role :

We are looking for a Machine Learning Engineer to build and deploy production-grade AI systems powered by Large Language Models (LLMs). In this role, you will work on retrieval-augmented generation (RAG), search, multilingual NLP, document understanding, evaluation frameworks, and AI-powered workflows. You will collaborate closely with research, engineering, and product teams to deliver reliable, scalable AI solutions used in real-world environments.

Key Responsibilities :

- Design, build, and deploy production-ready LLM and NLP applications.

- Develop retrieval and search pipelines using RAG, embeddings, vector databases, and reranking techniques.

- Build AI agents and workflow automation for complex business use cases.

- Fine-tune and adapt open-source LLMs using techniques such as LoRA, QLoRA, or PEFT.

- Develop multilingual NLP solutions including summarization, translation, information extraction, and question answering.

- Design evaluation frameworks to measure model quality, hallucinations, factuality, latency, and overall user experience.

- Build feedback pipelines that capture production failures and improve model performance.

- Work closely with data, engineering, and product teams to translate business requirements into scalable ML solutions.

- Deploy and maintain ML models in production while ensuring scalability, monitoring, and reliability.

- Continuously improve system performance, inference efficiency, and operational costs.

Required Skills :

- 3-6 years of experience building production Machine Learning or NLP systems.

- Strong programming skills in Python.

- Hands-on experience with PyTorch and Hugging Face Transformers.

- Experience working with Large Language Models (LLMs).

- Strong understanding of Retrieval-Augmented Generation (RAG) architectures.

- Experience with vector search technologies such as FAISS, Pinecone, Milvus, Chroma, or similar.

- Experience with prompt engineering, LLM evaluation, and model fine-tuning.

- Familiarity with LangChain, LangGraph, LlamaIndex, or similar orchestration frameworks.

- Experience building REST APIs using FastAPI or similar frameworks.

- Experience deploying ML solutions using Docker and cloud platforms (AWS, Azure, or GCP).

- Good understanding of software engineering best practices, version control, testing, and CI/CD.

Preferred Skills :

- Experience with AI agents or multi-agent systems.

- Experience with multilingual NLP applications.

- Experience with recommendation systems or semantic search.

- Experience with vLLM, Text Generation Inference (TGI), or LLM serving frameworks.

- Familiarity with ML evaluation tools such as Arize Phoenix, LangSmith, or similar.

- Experience working with enterprise AI applications in domains such as legal, healthcare, finance, or public sector.

- Publications or open-source contributions in AI/ML are a plus.

Nice to Have :

- Experience with distributed training, DeepSpeed, or FSDP.

- Experience optimizing inference latency and deployment costs.

- Knowledge of model monitoring, observability, and experimentation frameworks.

- Familiarity with cloud-native architectures and microservices.

What You'll Work On :

- Production LLM applications

- Retrieval-Augmented Generation (RAG)

- AI Agents & Workflow Automation

- Search & Recommendation Systems

- Multilingual NLP

- Evaluation & Observability

- Document Intelligence

- Prompt Engineering & Fine-tuning

- Cloud-native ML Deployment

If you're passionate about building reliable AI systems that solve real-world problems at scale, we'd love to hear from you.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...