HamburgerMenu
hirist

Software Stack Validation & Performance Engineer - AI Hardware

Square Root Consulting
8 - 15 Years
Bangalore

Posted on: 21/09/2026

Job Description

AI Software Stack Validation & Performance Engineer

Location : Bangalore

About the Role :

We are building a next-generation AI software stack optimized for high-performance execution on custom AI hardware. We are looking for a strong AI/ML Systems Engineer who can use and validate the platform end-to-end - from ML frameworks and model serving to compilers, runtimes, and low-level kernels.

This is an engineering-focused role, not traditional QA. You will develop and run real-world AI workloads, benchmark performance, debug complex issues across the software stack, and work closely with compiler, runtime, framework, and kernel teams to improve the platform.

Key Responsibilities :

- Develop, port, run, and validate real-world AI/ML workloads end-to-end.

- Build and benchmark workloads to identify functional issues, performance bottlenecks, regressions, and usability gaps.

- Develop automated validation and regression test suites using Python, C++, PyTest, and related tools.

- Work across PyTorch, model serving, compilers, runtimes, and AI kernels.

- Write, modify, optimize, and benchmark Triton/CUDA kernels.

- Debug issues across multiple layers, from Python/framework behavior to compiled kernels and runtime execution.

- Perform performance profiling and benchmarking using relevant ML performance tools, including MLPerf.

- Collaborate with compiler, runtime, framework, and kernel engineers to triage and resolve issues.

- Build scalable automation and CI-integrated validation infrastructure.

- Contribute to release quality, regression tracking, and performance analysis.

Required Qualifications :

- 8+ years of overall industry experience, with 3+ years in AI/ML systems, software validation, or performance engineering.

- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field.

- Strong programming skills in Python and C++.

- Hands-on experience with PyTorch and ML model development/debugging.

- Experience with LLMs, inference, model serving, or vLLM.

- Experience writing or modifying GPU/AI accelerator kernels using Triton, CUDA, or similar technologies.

- Strong experience with PyTest or similar automated testing frameworks.

- Understanding of compilers, runtimes, operators, kernel execution, and software stacks.

- Strong analytical and debugging skills with the ability to trace issues across different layers of a complex system.

Good to Have :

- Experience enabling new ML models/architectures on GPUs or AI accelerators.

- Experience with ML workload profiling and performance optimization.

- Ability to understand compiler-generated code and IR-level representations.

- Experience building CI/CD and automated validation infrastructure.

- Exposure to GPU, custom accelerator, or heterogeneous computing platforms.

- Experience with containers and orchestration for ML serving.

- Experience with software release processes and quality metrics.

Key Skills :

- Python, C++, PyTorch, LLM, vLLM, Triton, CUDA, AI/ML Systems, GPU, AI Accelerators, Model Serving, MLPerf, Performance Engineering, Compilers, Runtime, PyTest, CI/CD, Kernel Development, ML Infrastructure

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...