Posted on: 09/06/2026
Designation: Associate Director of Engineering
About Cloudkeeper :
CloudKeeper is a cloud cost optimization partner that combines the power of group buying & commitments management, expert cloud consulting & support, and an enhanced visibility & analytics platform to reduce cloud cost & help businesses maximize the value from AWS, Microsoft Azure, & Google Cloud.
A certified AWS Premier Partner, Azure Technology Consulting Partner, Google Cloud Partner, and FinOps Foundation Premier Member, CloudKeeper has helped 350+ global companies save an average of 20% on their cloud bills, modernize their cloud set-up and maximize value all while maintaining flexibility and avoiding any long-term commitments or cost.
CloudKeeper hived off from TO THE NEW, a digital technology service company with 2500+ employees and an 8-time GPTW winner.
About the Role :
We are building FinOps for AI a new product category that applies cloud FinOps discipline to AI infrastructure spend. AI workloads are the fastest-growing cloud cost category, and CloudKeeper is making a strategic bet on this market: helping enterprises govern and optimize the cost of GPU instances, LLM API calls, ML platform services, and AI dev tools across AWS, Azure, and GCP.
We are seeking an Associate Director of Engineering to build and lead CloudKeeper's FinOps for AI engineering practice. You will own the end-to-end engineering charter across three pillars visibility (Lens AI), usage optimization (Tuner AI), and rate optimization (Commit AI) extending CloudKeeper's existing Lens, Tuner, and Commit platforms to cover AI workloads.
This is a high-visibility leadership role with executive sponsorship, scaling to a team of 20+ engineers. You will define architecture, hire and lead Tech Leads / Staff Engineers across each optimization vertical, and drive technical strategy at the intersection of AI infrastructure, cloud cost, and large-scale data engineering. You will be the engineering depth-leader on AI workloads inside the team we expect you to go deep on GPU infrastructure, ML systems, and LLM application patterns, not just manage from above.
Responsibilities :
- Develop intelligent and scalable engineering solutions from scratch
- Partner with Product to shape product vision, roadmap, and goals
- Own high-level and low-level design for new products and features, alongside a team of strong developers
- Be responsible for server-side component design, detailed technical design, development, testing, deployment, and maintenance
- Build products using a modern Java tech stack (Spring Boot, Hibernate, REST), MySQL, MongoDB, Snowflake, and React
- Lead a team of 20+ engineers sprint planning, code reviews, mentorship, hiring, and unblocking
- Be a hands-on coder write production-grade code, set the quality bar, and lead by example
- Review business requirements and ensure delivery within timelines, with thorough testing and minimal defects
Must Have :
- 12+ years in product engineering with 2- 4+ years leading engineering teams of 5- 15 people building and shipping SaaS products (leadership experience in services, consulting, or systems integration teams does not apply)
- Track record shipping data-intensive SaaS, cloud infrastructure, or platform products in production at scale not just prototypes or internal tools
- Cloud experience is mandatory hands-on with at least one major cloud (AWS, Azure, or GCP); strong understanding of cloud cost levers and infra economics
- Ability to define technical architecture, work as a hands-on coder, and maintain coding standards and team policies
- Experience managing a team of at least 3 engineers
- Strong experience with OOAD frameworks Spring, Hibernate, REST
- Solid grasp of design patterns, system design, and distributed systems
- Experience with TDD, Continuous Integration, and build tools (Maven, Jenkins, Gradle)
- Strong experience with MySQL (or Oracle / equivalent RDBMS) schema design, query optimization, indexing
- Working knowledge of at least one cloud platform AWS preferred; GCP or Azure also acceptable
- Frontend exposure with React + TypeScript this role also mentors a small frontend team and reviews frontend work. Deep React expertise is not required, but the ability to guide engineers, review code, and ramp up quickly is essential
- Awareness of integration patterns queuing (RabbitMQ / Kafka), caching (Redis), messaging
- Strong interpersonal skills with the ability to work effectively across team boundaries
- Understanding of latest technologies and tools
Good to have :
- LLM optimization semantic caching, model routing, prompt optimization, token reduction, quantization, distillation
- Performance engineering : profiling, low-latency serving, throughput tuning, GPU utilization optimization at scale
- Familiarity with cloud billing data (AWS CUR, Azure billing exports, GCP billing) and AI provider billing APIs
- Understanding of FinOps principles and cloud cost optimization patterns
- Background in building data-intensive SaaS products, cloud platforms, or analytics products
- Experience with Kubernetes GPU node pools (EKS, AKS, GKE) and GPU-aware scheduling
- Experience with data pipelines and warehouses at scale Kafka, Spark, Snowflake, BigQuery
- Familiarity with backend frameworks: Spring Boot 2.7.x/3.2.x, Spring Cloud
- Exposure to AI dev tools and developer productivity (Copilot, Cursor, Claude Code)
Did you find something suspicious?
Posted by
Posted in
Backend Development
Functional Area
Engineering Management
Job Code
1642826