Posted on: 23/09/2026
Position : AI Software QA Engineer
Additional details :
We are building a next-generation software stack designed for high-performance execution on custom hardware. Our mission is to deliver industry-leading low-latency systems by developing an optimized, modular, and deeply integrated platform - from workload ingestion to hardware-level execution.
We are looking for strong engineers who will serve as real user of our AI software stack. You will exercise the full stack end-to-end - from ML frameworks and serving platforms through kernel authoring and compilation - validating functionality, performance, and usability from the perspective of ML engineers and infrastructure teams. Your feedback will directly shape the quality and developer experience of the platform. This is an engineering position.
Key Responsibilities :
- Validate the AI software stack end to end by developing, porting, and running representative ML workloads that reflect real world usage.
- Write, optimize, and benchmark AI workloads to identify functional gaps, performance bottlenecks, regressions, and usability issues.
- Design, develop, and maintain a comprehensive benchmarking test suite spanning model serving, ML frameworks, compilers, and kernel layers.
- Reproduce, triage, and deeply analyze issues across the stack - from Python level framework behavior to compiled kernel correctness and runtime performance.
- Perform performance benchmarking and profiling (e.g., MLPerf) to track scalability, regressions, and optimization impact across releases.
- Collaborate closely with compiler, runtime, framework, and kernel teams to provide actionable feedback and drive timely issue resolution.
- Build scalable validation platforms, automation frameworks, and CI integrated infrastructure to enable continuous and reliable quality assurance.
Required Qualifications :
- 8+ years of relevant industry experience (3+ Years AI/ML systems), software validation, performance engineering, or related domains.
- Bachelor's degree or higher in Computer Science, Electrical Engineering, or a related field.
- Strong programming skills in Python and C++, with experience developing, debugging, and maintaining production quality systems.
- Hands on experience with ML frameworks such as PyTorch, including model authoring and debugging.
- Experience with model serving platforms and inference workflows.
- Experience writing or modifying GPU kernels using Triton, CUDA, or similar kernel authoring technologies.
- Proficiency with PyTest or similar testing frameworks for building automated validation and regression test suites.
- Comfort working across multiple layers of a complex software stack, from Python frameworks to compilers, runtimes, and kernels.
- Strong analytical and systematic debugging skills, with the ability to isolate issues across framework, compiler, and runtime boundaries.
Strong Advantage :
- Experience enabling new model architectures or workloads on GPU or AI accelerator platforms.
- Hands on experience with performance profiling and benchmarking tools for ML workloads.
- Ability to read and reason about compiler generated code, including IR level representations.
- Experience designing or maintaining CI/CD pipelines and automated test infrastructure for ML systems.
- Exposure to GPU, custom accelerator, or heterogeneous compute ecosystems.
- Familiarity with container based deployment and orchestration for ML serving.
- Familiarity with release process management, defining and reporting quality matrix / release.
- Python programming.
- Large Language Models (LLMs), vLLM, and associated concepts (based on relevant experience).
- Operators, Fusion, and related technologies/platforms.
Did you find something suspicious?
Posted by
Posted in
Quality Assurance
Functional Area
QA & Testing
Job Code
1673813