HamburgerMenu
hirist

Job Description

Key Responsibilities :

- Optimize deep learning models for deployment using quantization and graph optimizations.

- Validate model accuracy after optimization to ensure production quality.

- Benchmark inference latency, throughput, memory footprint, and power consumption.

- Build and maintain benchmarking frameworks for consistent performance evaluation.

- Create deployment playbooks and technical documentation for every model release.

- Partner with product and engineering teams during integration to resolve performance and accuracy issues.

- Profile inference pipelines and identify CPU, GPU, DSP, or NPU bottlenecks.

- Debug ONNX export issues, unsupported operators, dynamic shapes, and runtime compatibility.

- Improve deployment workflows across multiple inference runtimes and hardware accelerators.

Required Qualifications :

- 3+ years of experience in Machine Learning Systems, Performance Engineering, or AI Infrastructure.

- Strong proficiency in Python and PyTorch.

- Hands-on experience exporting production models to ONNX.

- Deep understanding of :

1. Dynamic Shapes

2. Control Flow

3. Custom Operators

4. ONNX Graph Optimizations

- Experience implementing production-grade model quantization.

- Expertise with at least two of the following inference runtimes :

1. ONNX Runtime

2. TensorRT

3. CoreML

4. OpenVINO

5. Qualcomm QNN

6. LiteRT (TensorFlow Lite Runtime)

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...