Posted on: 16/07/2026
Key Responsibilities :
- Optimize deep learning models for deployment using quantization and graph optimizations.
- Validate model accuracy after optimization to ensure production quality.
- Benchmark inference latency, throughput, memory footprint, and power consumption.
- Build and maintain benchmarking frameworks for consistent performance evaluation.
- Create deployment playbooks and technical documentation for every model release.
- Partner with product and engineering teams during integration to resolve performance and accuracy issues.
- Profile inference pipelines and identify CPU, GPU, DSP, or NPU bottlenecks.
- Debug ONNX export issues, unsupported operators, dynamic shapes, and runtime compatibility.
- Improve deployment workflows across multiple inference runtimes and hardware accelerators.
Required Qualifications :
- 3+ years of experience in Machine Learning Systems, Performance Engineering, or AI Infrastructure.
- Strong proficiency in Python and PyTorch.
- Hands-on experience exporting production models to ONNX.
- Deep understanding of :
1. Dynamic Shapes
2. Control Flow
3. Custom Operators
4. ONNX Graph Optimizations
- Experience implementing production-grade model quantization.
- Expertise with at least two of the following inference runtimes :
1. ONNX Runtime
2. TensorRT
3. CoreML
4. OpenVINO
5. Qualcomm QNN
6. LiteRT (TensorFlow Lite Runtime)
Did you find something suspicious?