Posted on: 17/09/2026
Site Reliability Engineer
Hitachi Solutions Innovation Team :
This is a full-time individual contributor role in our product organization for an experienced Site Reliability Engineer focused on cloud operations, deployment reliability, infrastructure automation, observability, and production support for customer-facing SaaS and data platform environments. This is not a lead role. The position is designed for an engineer who can execute hands-on SRE responsibilities across Azure infrastructure, Kubernetes-based deployments, CI/CD pipelines, Terraform and other Infrastructure as Code tools, observability, incident response, release support, and customer escalation triage.
Individuals in this role will support the design, engineering, deployment, troubleshooting, and administration of Azure-based product environments, including Azure Kubernetes Service, Azure PaaS and IaaS services, Azure Key Vault, Azure Storage Accounts, Azure Data Factory, Azure Databricks, GitHub Actions, Azure DevOps pipelines, Terraform, Helm, GitOps deployment patterns, and customer tenant configuration. The role will work closely with application engineering, QA, delivery, security, cloud engineering, and operations teams to maintain stable, repeatable, and well-documented deployment and support processes across development, QA, staging, and production environments.
The ideal candidate will bring practical hands-on experience similar to the current SRE skillsets, Tier 1 and Tier 2 SRE support, Terraform, AKS, GitHub, Azure DevOps, Key Vault, RBAC, customer deployment troubleshooting, README and runbook documentation, CI/CD secret management, environment access setup, operational monitoring, release automation, and escalation support for production deployments.
Responsibilities :
- Support Reliability : Support availability, latency, performance, reliability, efficiency, observability, incident response, emergency response, capacity planning, SLOs, SLIs, error budgets, and operational dashboards for production and pre-production environments.
- Troubleshoot Issues : Analyze, troubleshoot, and resolve operational issues that impact deployment reliability, customer environments, platform stability, service uptime, and defined SLOs.
- Maintain Stability : Maintain site stability, performance, reliability, and uptime for Azure-hosted SaaS and data platform production environments.
- Operate Production Environments : Support 24x7 highly available production environments for SaaS or cloud service provider platforms.
- Support Releases : Support automated release, hotfix, and customer deployment processes across multiple environments.
- Improve CI/CD : Execute and improve CI/CD pipelines using Azure DevOps, GitHub Actions, YAML build pipelines, Git operations, branching strategies, pull requests, release pipelines, and deployment approvals.
- Automate Infrastructure : Implement, maintain, and troubleshoot infrastructure as code using Terraform, Bicep, ARM templates, Helm, Kubernetes manifests, and related technologies.
- Manage Kubernetes : Support AKS, Kubernetes, Helm, containerized workloads, Kubernetes service accounts, GitOps deployment patterns, Flux, Argo CD, and environment configuration.
- Support Identity & Access : Assist with Key Vault, RBAC, service principal, identity, permissions, access policy, Secrets Officer role, CI/CD secrets, and customer tenant configuration issues that affect deployment success.
- Manage Secrets : Support CI/CD secret management and automation patterns that reduce manual updates to pipeline variables, deployment secrets, and environment-specific configuration.
- Restore Services : Perform application-specific production support, incident management, change management, problem management, service restoration, and root cause analysis.
- Lead Incident Response : Contribute to incident root cause analysis, service restoration, coordinated outage response, and incident command activities during outage events when needed.
- Triage Escalations : Triage incoming SRE support, Web Support escalation requests, and customer deployment issues, routing to applicable internal teams when needed and resolving issues directly within scope.
- Validate Readiness : Support release readiness by validating deployment steps, addressing configuration issues, reviewing runbooks, and coordinating with QA, engineering, delivery, security, and operations teams.
- Document Operations : Develop, maintain, and improve technical documentation, including design specifications, user guides, README files, runbooks, deployment instructions, troubleshooting procedures, support guides, and best practice guidelines.
- Improve Readiness : Improve operational readiness by identifying repeatable failure patterns and recommending automation, documentation, configuration updates, or process improvements.
- Advance Observability : Develop and improve multi-environment observability patterns based on existing systems, including logs, metrics, traces, dashboards, and capacity trend analysis.
- Optimize Performance : Actively look for opportunities to improve system availability and performance by applying learnings from monitoring, observability, incident reviews, and production support.
- Identify Improvements : Identify reliability, performance, availability, and operational improvements for product architecture using a data-driven approach.
- Collaborate Cross-Functionally : Collaborate with product managers, software engineers, QA engineers, designers, delivery teams, security teams, and cloud engineering teams to support efficient project execution.
- Participate in Agile : Participate in Agile ceremonies, including sprint planning, stand-up meetings, retrospectives, backlog review, and cross-pod coordination discussions.
- Support Stakeholders : Collaborate with customers and internal stakeholders to understand needs, gather feedback, provide technical support, and guide successful deployment outcomes.
- Continue Learning : Stay current with Azure, Kubernetes, Terraform, Databricks deployment patterns, CI/CD, observability, cloud networking, security, and cloud operations best practices.
Qualifications & Skills :
- SRE Experience : Hands-on experience as a Site Reliability Engineer, DevOps Engineer, Cloud Operations Engineer, Platform Engineer, or similar role supporting SaaS or cloud-based production environments.
- Production Support : Strong background supporting highly available 24x7 production environments for SaaS, product, or cloud service provider platforms.
- Azure Infrastructure : Strong experience with Azure infrastructure, including Azure Kubernetes Service, Azure Key Vault, Azure Data Factory, Azure Databricks, Azure Storage Accounts, service principals, RBAC, networking concepts, Azure PaaS resources, and Azure IaaS resources.
- Terraform Automation : Practical experience with Terraform-based infrastructure automation, including troubleshooting state, resource drift, deployment failures, environment configuration issues, and infrastructure provisioning.
- Infrastructure as Code : Experience with infrastructure as code tools and templates such as Terraform, Bicep, ARM, Helm, or similar technologies.
- CI/CD Automation : Experience with CI/CD and release automation using Azure DevOps, GitHub Actions, YAML build pipelines, Git operations, pull requests, branching strategies, release pipelines, and deployment approvals.
- Kubernetes Operations : Working knowledge of Kubernetes, Helm, containerized architectures, AKS operations, Kubernetes service accounts, container deployment patterns, and troubleshooting cluster-related deployment issues.
- GitOps Practices : Experience with GitOps deployment approaches using tools such as Flux, Argo CD, or similar technologies.
- Release Support : Experience supporting customer deployments, hotfix rollouts, environment promotions, QA, staging, production upgrades, and release coordination.
- Observability Tools : Experience with observability, monitoring, logging, tracing, dashboards, and APM tools such as Datadog, Application Insights, Azure Monitor, Prometheus, Grafana, Splunk, Elastic, Sentry, Dynatrace, New Relic, Nagios, Zabbix, or similar platforms.
- Observability Planning : Experience implementing observability plans across logs, metrics, traces, alerts, dashboards, and operational review processes.
- Troubleshooting : Ability to troubleshoot complex deployment, permissions, networking, secrets, RBAC, identity, configuration, and environment issues across unfamiliar technical environments.
- Data Platform Awareness : Familiarity with Azure Databricks, Data Factory, data platform deployment patterns, Data Lake ACLs, Synapse, Spark data components, Unity Catalog, or similar data ecosystem components are preferred.
- Scripting & Automation : Working knowledge of scripting and automation using Bash, Python, PowerShell, Go, or similar tools.
- Source Control : Excellent command of source control using Git, including branching strategies, policies, pull requests, merge processes, and release tagging.
- Systems Knowledge : Strong understanding of Linux, Windows, software development, systems, networking, cloud concepts, and operational support models.
- Cloud Networking : Familiarity with cloud networking and security elements such as VNets, peering, firewalls, private endpoints, private DNS zones, NAT, service endpoints, and related access controls.
- Environment Readiness : Experience with release automation, system administration, configuration management, deployment validation, and environment readiness.
- Technical Documentation : Ability to create and maintain clear technical documentation, runbooks, deployment guides, support procedures, README files, and best practice documentation.
- Communication & Collaboration : Strong analytical, troubleshooting, communication, and collaboration skills with the ability to work across engineering, QA, delivery, security, operations, and customer-facing teams.
- Independent Execution : Ability to independently execute assigned work while knowing when to escalate architectural, customer-impacting, security-sensitive, or production-risk decisions.
- Deployment Flexibility : Comfortable supporting after-hours or time-zone-aligned deployment windows when required for customer or production release needs.
Additional Skills (Great to have) :
- Product Deployment : Experience supporting data products, SaaS, or Azure Marketplace-style product deployments.
- Databricks Automation : Experience with Databricks deployment automation, configuration promotion, notebooks, secrets, data workflows, and customer environment promotion.
- Microsoft Fabric Automation and Data Operations
- Dynamics 365 experience
- GitOps Orchestration : Familiarity with GitOps practices, Flux, Argo CD, Helm charts, Terraform controllers, and Kubernetes-based deployment orchestration.
- Azure Identity & Security : Experience with service principal permissions, Azure identity and security patterns as it relates to Entra ID and Security Token Service.
- Customer Support : Experience with customer-facing technical support, deployment escalation triage, Web Support escalation routing, and internal SRE support ticket processes.
- Cloud Operations : Familiarity with Docker operations, REST APIs, SaaS deployment APIs, Azure networking, private endpoints, private DNS zones, firewalls, NAT, and cloud security practices.
- MLOps Exposure : Exposure to MLflow, MLOps pipeline technologies, AI platform deployment pipelines, or data science platform operational support is a plus.
- Database Deployment : Experience with database deployment pipelines such as DACPACs or similar technologies is a plus.
- Enterprise Authentication : Experience with SSO, federated security, and enterprise authentication patterns is a plus.
- Testing Frameworks : Familiarity with unit testing and mocking frameworks such as unittest, MSTest, NUnit, RhinoMocks, Moq, NSubstitute, or similar tools is a plus.
- CI/CD Tools : Prior experience with Jenkins or similar CI/CD technology is acceptable.
- Certifications : Optional certifications include :
1. Microsoft Certified : Azure Solutions Architect Expert
2. Microsoft Certified : DevOps Engineer Expert
3. Certified Kubernetes Application Developer
4. Certified Kubernetes Administrator.
Practices, Principles, and Techniques :
- Site Reliability Engineering
- Instrumentation strategy and observability engineering
- Incident management, emergency response, root cause analysis, and service restoration
- SLOs, SLIs, error budgets, operational dashboards, and reliability reporting
- CI/CD and release automation
- Infrastructure as Code with Terraform, Bicep, ARM, Helm, and related tools
- Kubernetes, AKS, containerized workloads, and GitOps
- Azure identity, RBAC, Key Vault, service principals, and secret management
- Monitoring, observability, logging, tracing, alerting, and capacity analysis
- Release communication, deployment coordination, and cross-team collaboration
- Security, compliance, and production access governance
- Test Driven Development concepts as they relate to CI/CD and DevOps
- Deployment documentation, runbooks, and operational knowledge management
- Agile delivery, sprint execution, and cross-functional collaboration
- Automation-first mindset to reduce toil and improve deployment consistency
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1672264