Posted on: 05/08/2026
Your Opportunity :
As a Senior Software Engineer within the Container Fabric (CF) organization, you will be a key driver in evolving New Relics global internal platform. We are looking for an operations-heavy engineer with 5 to 8 years of relevant experience who can leverage open-source and custom tooling to orchestrate and maintain large-scale Kubernetes environments. You will play a "Captain" role leading critical deliverables and mentoring junior engineers while maintaining the reliability of our global fleet.
What You'll Do :
1. Architectural Leadership :
Drive the design and implementation of internal tools, specifically focusing on Kubernetes Operators and Controllers to automate resource management.
2. Platform Orchestration :
Lead complex, large-scale infrastructure shifts.
3. Operational Excellence :
Take ownership of incident response, author comprehensive retrospectives, and implement systemic hardening to prevent recurrence using advanced overcommit strategies.
This Role Requires :
Experience :
- 5 to 8 years in a DevOps, Site Reliability, or Infrastructure Engineering role.
- Kubernetes Mastery: Deep internals knowledge of Kubernetes and hands-on experience writing custom operators.
- Tooling Proficiency: Strong experience building production-grade tools and services, specifically for infrastructure automation.
- Operations-Heavy Mindset: A proven track record of Day 1/Day 2 operations for a large-scale Kubernetes fleet, handling high-severity incidents, and improving SLA compliance through automation.
- Cloud Infrastructure: Multi-cloud experience, including familiarity with cloud-native tools and managing OS migrations.
- Mentorship: Ability to lead projects as a "Captain," providing technical direction and unblocking team members across different time zones.
A Huge Plus (Bonus Points) :
We don't expect you to know everything, but we would love it if you have experience with :
- Golang: A basic ability to read, navigate, and understand Golang code.
- Advanced Autoscaling: Experience implementing advanced autoscaling strategies and efficiency optimizations using tools like Karpenter.
- Infrastructure as Code & Control Planes: Hands-on experience with Terraform or Crossplane and Cluster API.
- GitOps & CI/CD: Experience deploying and managing platform infrastructure via Argo CD.
- Cloud Providers: Strong working knowledge of managing resources within AWS and Azure environments.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1660759