Posted on: 30/09/2026
Role : Infrastructure Operations & Cloud Monitoring Engineer
Experience : 8 - 12 Years
Location : Bangalore (Onsite)
Key Responsibilities :
1. Infrastructure Monitoring & Maintenance :
- Continuously monitor server and infrastructure performance, availability, and security.
- Perform regular VM server updates and apply security patches.
- Proactively review server logs and troubleshoot issues before they escalate.
- Maintain documentation of server configurations, update schedules, and recovery procedures.
- Oversee backup and disaster recovery planning and execution.
2. Incident & Outage Response :
- Lead response efforts during outages or performance degradation.
- Coordinate with internal teams and vendors to resolve issues quickly and effectively.
- Create and manage service desk tickets with third-party vendors.
- Track resolution timelines and escalate as needed.
- Prepare RCA (root cause analysis) document for high severity incidents and share with the stakeholders.
- Maintain documentation of recurring issues and suggest possible resolutions.
3. Performance & KPI Tracking :
- Monitor and report on key performance indicators (KPIs) related to server/database efficiency and uptime.
- Recommend improvements based on trend analysis and performance data.
- Prepare regular reports for leadership on system health, incidents, and vendor performance.
- Communicate technical issues and resolutions clearly to non-technical stakeholders.
4. Security & IT Compliance :
- Serve as the primary point of contact between operations, IT, and security teams.
- Ensure infrastructure aligns with corporate security policies and compliance standards.
- Collaborate with DevOps, software engineering, and QA teams to align infrastructure with product needs.
- Manage access controls and permissions for internal team members across infrastructure systems.
- Support annual security audits by preparing documentation, responding to audit requests, and implementing required changes.
- Ensure infrastructure complies with internal policies and external regulations (e.g., SOC 2, ISO 27001).
5. Financial Oversight :
- Manage payments and contracts for service providers.
- Track service costs and ensure budget adherence.
6. Automation & Tooling:
- Identify opportunities to automate monitoring, reporting, and patching workflows.
- Evaluate and implement tools for infrastructure observability (e.g., Datadog, Splunk, Prometheus).
Preferred Skills :
- Strong understanding of the Azure cloud environment.
- Infrastructure Monitoring Tools: Experience with observability platforms like Datadog, Splunk, Prometheus, Grafana, or similar.
- Virtualization & Server Management: Strong knowledge of VMware, Hyper-V, or other virtualization technologies.
- Operating Systems: Proficiency in Windows Server and Linux administration.
- Networking Fundamentals: Understanding of TCP/IP, DNS, firewalls, and load balancing.
- Backup & Disaster Recovery: Hands-on experience with backup solutions and DR planning.
- Security & Compliance: Familiarity with SOC 2, ISO 27001, and corporate security standards.
- Automation & Scripting: Ability to automate tasks using PowerShell, Python, or Bash.
- Incident Management: Skilled in ITIL processes, RCA preparation, and escalation workflows.
Additional Competencies :
- Ability to co-ordinate among global user groups and internally within the solution center.
- Ability to multi-task and prioritize.
- Willing to work flexible hours as the job demands (collaborate with global colleagues, when necessary).
- Communication: Ability to explain technical issues clearly to non-technical stakeholders.
- Collaboration: Comfortable working with DevOps, QA, and security teams.
- Vendor Management: Skilled in coordinating with third-party providers and managing SLAs.
- Documentation: Strong attention to detail for maintaining accurate infrastructure documentation.
- Ready to learn new skills on the job.
- Excellent troubleshooting skills.
- Action-oriented, delivers on commitments.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
IT Infrastructure Services
Job Code
1675863