HamburgerMenu
hirist

Lead Applied AI/LLM Engineer

AKS INTELLIGENT SYSTEMS LLP
3 - 7 Years
rupee27-32 LPA
Gurgaon/Gurugram

Posted on: 18/08/2026

Job Description

Lead Applied AI / LLM Engineer

PERMANENT FULL-TIME OPPORTUNITY

RAG, Agents & Evaluation | Remote-first within India | Reports to the CTO

Experience : 3 - 7 years

Joining : ASAP; within 30 days preferred

Location : Anywhere in India; Gurugram/NCR attendance as required


Work model : Flexible hours; core overlap 11:00 AM - 4:00 PM IST

The mandate :

AKS is building a permanent technical core of engineers who own outcomes, not isolated tickets. You will take end-to-end responsibility for applied AI systems across discovery, architecture, implementation, evaluation, deployment, UAT, production support, documentation and handover. You will coordinate directly with the CTO, collaborate with the full-stack lead and, where appropriate, guide interns or junior engineers while retaining accountability for the final result.

This is primarily an applied LLM and application-engineering role. Model training and research are part of the toolkit, but the central objective is to deliver reliable systems that work with real client data, constraints and acceptance criteria. A typical portfolio may include one or two active PoC/UAT builds and several systems in hypercare or maintenance.

What makes this role different :

AI coding agents and approved model/tool subscriptions are provided and encouraged. You remain accountable for architecture, correctness, security, tests and every accepted change. You must understand and be able to defend the system without relying on the tool that generated a suggestion.

What you will own :

- Convert ambiguous business and client problems into written technical outcomes, acceptance criteria and delivery plans.

- Design, build and maintain production-oriented RAG, agentic and document-intelligence systems with explicit boundaries between deterministic logic, conventional ML and generative components.

- Own ingestion, retrieval, embeddings, vector search, prompting, structured outputs, tool use, orchestration, guardrails and human-review paths.

- Build defensible evaluations for retrieval quality, groundedness, citations, entity isolation, refusal behaviour, latency, cost and failure cases; turn production failures into regression tests.

- Fine-tune models using LoRA or related parameter-efficient methods and deploy reproducible inference-serving workflows.

- Develop production-quality Python services and APIs; integrate AI components with FastAPI applications, PostgreSQL/vector stores and frontend or workflow systems.

- Deploy and troubleshoot services on cloud and client-controlled Linux/GPU environments using Docker, CI/CD, logging and monitoring.

- Reuse proven components across applications without creating hidden coupling; improve existing repositories rather than rebuilding everything by default.

- Maintain clear Git history, pull requests, tests, runbooks, architecture decisions, model/evaluation records and knowledge-transfer material.

- Communicate progress, dependencies, risks and production incidents clearly to the CTO, delivery stakeholders and clients where required.

- Review work and nurture interns or junior engineers when assigned; ownership cannot be delegated away.

Required capabilities :

- 3 - 7 years of relevant software, ML or applied-AI experience, including personally shipped systems.

- Strong Python engineering and hands-on PyTorch proficiency.

- Strong foundations in statistics, statistical machine learning, deep learning, experimental design and practical model evaluation.

- Production experience with LLM applications, RAG or agentic workflows; ability to diagnose retrieval, grounding, hallucination, identity and tool-use failures.

- Hands-on LoRA fine-tuning and inference serving, including data preparation, baselines, leakage control, versioning, evaluation, capacity and rollback considerations.

- Working knowledge of embeddings, vector databases, information retrieval, structured extraction and document-processing pipelines.

- Strong Git and GitHub proficiency: branches, pull requests, reviewable commits, code review and repository hygiene.

- Strong Linux, Docker, REST API, testing and debugging skills; working knowledge of CI/CD, monitoring and at least one cloud platform.

- Ability to reason about security, privacy, client-data boundaries, model/provider selection, cost, latency and production fallback behaviour.

- Clear written and verbal communication; direct client discovery or review experience is highly valued.

- Bachelors degree in Computer Science, AI/ML, Data Science, Statistics, Mathematics, Engineering or another relevant quantitative discipline.

Preferred experience :

- Quantization, GPU-memory/performance optimization, vLLM or comparable model-serving frameworks.

- PostgreSQL with pgvector or other vector stores; Pydantic/FastAPI; evaluation and observability tooling such as MLflow, Langfuse, Phoenix or OpenTelemetry.

- Enterprise document intelligence, legal/research assistants, RFQ/quotation automation, voice/calling automation or other high-trust workflows.

- Local/on-premise model deployment, distributed inference, Kubernetes or infrastructure-as-code.

- Conventional ML problems such as classification, forecasting or ranking where labels, calibration and business-valid metrics matter.

- Informal technical leadership, code review and guidance of interns/junior engineers.

Success in the first 90 days :

- Establish a verified architecture, data-flow, evaluation and deployment map for the assigned portfolio.

- Reproduce existing baselines and known failures; resolve at least one bounded maintenance issue through a reviewed deployment.

- Put representative evaluation/regression checks into the delivery workflow and ship one measurable reliability improvement.

- Demonstrate a reproducible LoRA training and serving workflow in staging.

- Lead one active PoC/UAT workstream while maintaining transparent status and risk control across 4-5 systems.

- Deliver or materially advance one system to a written acceptance milestone and leave reusable, documented components behind.

Work model and support expectations :

- Permanent, full-time and remote-first within India. Attendance at the Gurugram office or client meetings may be required with reasonable notice.

- Eight working hours on each scheduled workday, excluding breaks. The normal reference window is 9:30 AM - 6:30 PM IST with a one-hour meal break; start/end times may flex while preserving mandatory availability from 11:00 AM - 4:00 PM IST.

- Regular Monday - Friday schedule plus two pre-notified working Saturdays per calendar month. The Saturday schedule is communicated before the month begins.

- Routine overnight support is not expected. Exceptional production incidents may occasionally require availability outside normal hours, handled under company policy and applicable requirements.

- Client travel is currently infrequent and primarily within NCR. Approved business travel outside NCR is reimbursed; other required travel arrangements are confirmed in advance.

Compensation and support :

- Annual fixed gross salary : INR 24,00,000 - 28,00,000.

- Performance-linked variable pay : up to INR 3,00,000 per annum, assessed and payable quarterly against written, measurable outcomes. Variable pay is not guaranteed and may be prorated for eligible service.

- Company-funded access to approved coding agents, model APIs and role-relevant software subscriptions.

- Approved learning resources or certifications may be company-funded where relevant to delivery.

- A company laptop/workstation may be provided based on role and information-security requirements.

- Paid leave, holidays and statutory benefits in accordance with company policy and applicable law.

Selection process :

1. Structured initial evidence and logistics screen with HR.

2. Technical architecture and ownership interview with the CTO.

3. Role-relevant synthetic work sample; approved coding agents may be used with disclosure and verification.

4. Final alignment and, where appropriate, consent-based reference or document verification.

The job is for:

May work from home
info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Posted by

Recruiter

HR at AKS INTELLIGENT SYSTEMS LLP

Last Active: NA as recruiter has posted this job through third party tool.

Job Views:  
52
Applications:  27
Recruiter Actions:  0

Posted in

AI/ML

Functional Area

ML / DL Engineering

Job Code

1664136

Loading chat...