Posted on: 05/05/2026
Job Description :
We are seeking a highly skilled candidate with 10+ years of experience on on-premises and cloud support or consulting with customer facing roles in systems architecture, administration, operations, software support, IT consulting, for enterprise and/or government customer base, and advanced troubleshooting..
The role requires expertise in Windows Failover Clustering, Hyper-V,Virtualization, Windows Platform, Performance Troubleshooting and Azure.
Understanding of Microsofts Modern IT Architecture and Strategy and Proven track record in successfully planning, deploying, operating, and optimizing on-premises and cloud environments.
Relevant Windows Hybrid and Azure certifications would be preferred.
The ideal candidate will be responsible for managing and optimizing enterprise infrastructure platforms, ensuring high availability, resilience, security, and compliance, and delivering robust solutions to meet business and regulatory requirements backed with customer-facing technical leadership.
Primary Skills Required :
- Windows Server (2012 R2 - 2025 advanced to expert level))
- Windows Performance & Advanced Troubleshooting
- Windows Failover Clustering
- Hyper-V
- Virtualization (migration from VMWare to the following, Hyper-V, Azure VMWare Solutions, Azure Stack HCI, Azure VMs)
- Azure Infrastructure (Azure Arc, Azure Monitor, Azure Backup, Azure SiteRecovery, Defender for Cloud, Azure Local / hybrid infrastructure scenarios)
- Windows Defender (incl. Defender for Endpoint on servers)
- Systems Center Virtual Machine Manager
- Windows Firewall
- Windows Server Security (Defender, Firewall, Device Guard, Credential Guard,and BitLocker)
Key Responsibilities :
1. Windows Failover Cluster & Virtualization Management :
i. Creation, decommissioning, migration, and troubleshooting of Windows Failover Clusters.
ii. Design and support of Hyper-V clusters, Cluster Shared Volumes (CSV),. quorum models, and live migration.
iii. Planning and execution of failover cluster migrations with minimal downtime.
iv. Design and validation of disaster recovery solutions using Hyper-V Replica,including cluster-to-cluster and S2D-backed scenarios.
2. Operating System & Platform Troubleshooting :
i. Addressing issues related to OS in-place upgrades and platform lifecycle management.
ii. Troubleshooting high CPU, memory, disk latency, storage, and networking performance issues.
iii. Deep-dive analysis of unexpected restarts, VM crashes, cluster instability, and file-system anomalies.
iv. Proactively improving server reliability and performance through design and configuration recommendations.
3. Root Cause Analysis (RCA) & Critical Incident Handling :
i. End-to-end ownership of RCA activities, including log collection, dump analysis coordination, hypothesis validation, and corrective actions.
ii. Providing clear, evidence-based RCA reports for reactive and critical incidents(CritSits).
iii. Working closely with internal engineering teams, vendors, and customers during high-severity incidents.
4. Security, Vulnerability & Compliance Mitigations :
i. Implementing strategies to mitigate vulnerabilities and meet security and audit compliance requirements, especially in regulated environments.
ii. Configuring and managing :
a. Windows Defender (including performance-safe exclusions for Hyper-V)
b. Windows Defender (including performance-safe exclusions for Hyper-V)
c. Windows Firewall
d. Device Guard and Credential Guard
e. BitLocker, including cluster-aware and CSV scenarios
iii. Supporting regular audits, compliance reviews, and security assessments with audit-safe documentation and explanations.
5. Storage & High Availability :
i. Design, deployment, and support of Storage Spaces Direct (S2D) clusters,including resiliency selection (mirror / dual parity), capacity planning, and failure handling.
ii. Advising on storage architecture, firmware alignment, and operational best practices for clustered environments.
6. Workshops, Enablement & Knowledge Sharing :
i. Conducting customer enablement activities and internal workshops on :
a. Hyper-V and Failover Clustering best practices
b. Common failure patterns and troubleshooting
c. Performance tuning and observability
ii. Sharing knowledge on new technologies and adoptions like Windows Server2025, security advancements, and Windows Admin Center (WAC).
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Technical / Solution Architect
Job Code
1633534