Posted on: 10/07/2026
Job Description:
Role & responsibilities:
Platform Engineering and Automation:
- Design, develop, and maintain high-quality software solutions, automation tools, and platform capabilities that improve operational efficiency and developer productivity.
- Write clean, maintainable, scalable, and well-tested code following engineering best practices.
- Develop integrations, APIs, automation workflows, and self-service platform components to support infrastructure and application lifecycle management.
- Participate in code reviews and contribute to continuous improvement of engineering standards.
Cloud Infrastructure and Operations:
- Design, deploy, and maintain cloud infrastructure on AWS using Terraform.
- Manage and optimize compute, networking, storage, and database services.
- Ensure high availability, scalability, security, and cost optimization of cloud resources.
- Implement infrastructure best practices and governance standards.
Container Platform and Reliability:
- Manage and maintain Kubernetes (EKS) clusters across multiple environments.
- Design platform standards for containerized applications.
- Implement cluster upgrades, scaling strategies, security controls, and disaster recovery procedures.
- Troubleshoot and resolve platform-related incidents and performance issues.
CI/CD and Deployment Automation:
- Design and maintain CI/CD pipelines using GitHub Actions and related tooling.
- Implement automated build, test, deployment, and rollback processes.
- Support GitOps practices using ArgoCD.
- Improve deployment reliability and reduce lead time for changes.
Developer Experience and Self-Service Platform:
- Contribute to Internal Developer Platform initiatives.
- Develop and maintain Backstage integrations, templates, and service catalog capabilities.
- Improve developer self-service capabilities and platform adoption.
- Create standardized onboarding and deployment workflows.
Observability and Monitoring:
- Implement and maintain monitoring, logging, tracing, and alerting solutions.
- Integrate observability platforms such as Datadog.
- Define SLOs, SLIs, dashboards, and operational metrics.
Collaboration and Technical Leadership:
- Partner with software engineering teams to improve application reliability.
- Mentor engineers on cloud-native and platform engineering best practices.
- Lead technical discussions, architecture reviews, and platform initiatives.
- Drive continuous improvement through automation and operational excellence.
Documentation and Operational Readiness:
- Create, maintain, and continuously improve comprehensive technical documentation for software applications, cloud infrastructure, deployment processes, platform services, and operational procedures.
- Develop architecture diagrams, runbooks, knowledge base articles, and onboarding materials to support engineering teams.
- Ensure documentation remains accurate, up-to-date, and aligned with organizational standards and compliance requirements.
Required Skills:
- 8+ years of relevant experience in platform engineering, site reliability engineering, cloud infrastructure, or DevOps.
- Proficiency in one or more programming languages such as Java, Python, Go or any other.
- Hands-on experience enabling application development teams through platform, automation, or DevOps solutions.
- Strong experience building and operating automated cloud platforms using Infrastructure as Code.
- Strong experience with AWS and infrastructure automation tools such as Terraform, GitHub Actions, and the AWS SDK.
- Experience with CI/CD pipelines, infrastructure deployment automation, and observability or monitoring tools.
- Experience with Kubernetes, Docker, and deploying containerized or microservices-based applications.
- Strong problem-solving skills with the ability to identify manual inefficiencies and automate them at scale.
- Experience with incident response, root cause analysis, and production operations support.
- Strong communication and collaboration skills, with the ability to work effectively across engineering and operations teams.
Preferred candidate profile:
- Experience contributing to a self-service platform or Internal Developer Platform (IDP).
- Experience with Backstage, ArgoCD, Datadog, or similar platform and observability tools.
- Experience mentoring engineers and leading technical initiatives across teams.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1652936