Posted on: 23/07/2026
Company Overview :
Scaling Theory is an emerging leader in the artificial intelligence infrastructure space, focused on building the foundational layers that power next-generation generative models. By optimizing the intersection of hardware efficiency and algorithmic performance, the company enables enterprises to train and deploy massive models with unprecedented speed and cost-effectiveness. Operating at the cutting edge of AI research and systems engineering, Scaling Theory provides the computational backbone for high-stakes machine learning applications across global industries.
Role Overview :
As an ML/AI Engineer focused on Foundation Model Engineering, you will be at the core of our efforts to push the boundaries of large-scale model training. You will work closely with research scientists and infrastructure engineers to design, implement, and optimize training pipelines that handle billions of parameters.
Your day-to-day will involve navigating the complexities of distributed computing, ensuring that our models achieve peak performance on GPU clusters. By bridging the gap between theoretical model architecture and practical hardware execution, you will directly influence the efficiency and capability of the models that define our product roadmap.
Key Responsibilities :
- Architect and maintain distributed training pipelines to ensure high throughput and fault tolerance during the training of large-scale foundation models.
- Optimize deep learning kernels and memory management strategies to maximize GPU utilization and reduce training latency.
- Implement advanced parallelization techniques, including tensor, pipeline, and data parallelism, to scale model training across multi-node clusters.
- Collaborate with the research team to integrate and test novel transformer architectures, ensuring they are optimized for production-grade hardware.
- Develop robust monitoring and diagnostic tools to identify bottlenecks in distributed training jobs, ensuring consistent model convergence.
- Contribute to the development of internal libraries and frameworks that streamline the deployment of AI models in resource-constrained environments.
Required Skillset :
- Demonstrated expertise in designing and training deep learning models, with a specific focus on transformer-based architectures.
- Proven ability to manage distributed training workloads using frameworks such as DeepSpeed, PyTorch Distributed, or similar technologies.
- Strong proficiency in GPU programming and optimization, with a deep understanding of how to extract maximum performance from modern hardware accelerators.
- Advanced command of Linux environments, including system-level debugging, performance profiling, and shell scripting for automation.
- Experience in handling large-scale datasets using PySpark or similar distributed data processing frameworks to prepare high-quality training corpora.
- Ability to communicate complex technical concepts clearly to cross-functional teams, fostering a collaborative environment that values rigorous engineering standards.
- A background in Computer Science or a related quantitative field, with a track record of solving complex algorithmic problems in high-pressure environments.
- Flexibility to work on-site in Bangalore, collaborating closely with the engineering team to iterate rapidly on hardware-software co-design challenges.
- Candidates should possess between 0 - 7 years of relevant experience in machine learning engineering or high-performance computing.
Did you find something suspicious?