Position Summary :
We are looking for a Site Reliability Engineer (SRE) / Production Support Engineer with hands-on experience supporting Trading, Capital Markets, Brokerage, Investment Banking, OMS/RMS, Exchange, or FinTech platforms.
The ideal candidate should have strong expertise in Linux, SQL, and Python/Bash scripting, with experience in production monitoring, incident management, RCA, automation, and ensuring high availability of mission-critical trading applications. Candidates with exposure to the trading lifecycle (Order Management, Trade Execution, Settlement, and Post-Trade Support) will be preferred.
Experience & Required Skill Sets :
- Ensure 24-7 uptime and stability of production systems.
- Investigate and troubleshoot production issues.
- Collaborate with developers to optimize system performance.
- Participate in on-call rotation to provide 24/7 support for critical systems.
- Work on automation and enhancements to reduce manual processes / intervention.
- Relevant 5+ years of experience in SRE / Production/Product Support role, with a track record of implementing SRE practices.
- Basic understanding of cloud solutions provided by providers such as AWS or Azure.
- Basic-Intermediate knowledge of Scripting in either of Bash/Python/PowerShell.
- Good presentation, communication and interpersonal skills with the ability to collaborate effectively with cross-functional teams and stakeholders across different countries and cultures.
- Good problem solving and troubleshooting skills.
- Continuous learning mindset and willingness to adapt to new technologies and industry trends.
- Good Understanding of Operating System Commands (Linux), SQL (Ability to write, analyze queries and deduce / build important information per requirement).
- In-depth knowledge of Trading Life Cycle : The candidate should possess a comprehensive understanding of trading life cycle, including order management, trade execution, settlement and post-trade processes. Familiarity with various financial products like Equities, Derivatives, Currencies, Commodities, FX is a plus.
- Incident and Problem Management Expertise : The candidate must demonstrate strong problem-solving skills and the ability to manage incidents frequently and efficiently within a fast paced trading environment. This includes identifying, analyzing and resolving issues related to trading systems and processes as well as collaborating with cross-functional teams to implement long-term solutions and improve operational efficiency.
- Good Understanding of Tools :
1. Orchestration : Autosys / Airflow or Cron
2. Monitoring & Logging : PagerDuty, Prometheus & Grafana or Datadog, Splunk
3. Project Management / ITSM : Service Now (Basic ability to navigate / create change tickets / incidents), Jira (Basic ability to create Jira Tickets, ability to filter your work)
Education :
- Bachelors degree or masters in computer science, Engineering, Software Engineering or a relevant field.