Join the Platform & Production Reliability team at CXM and play a vital role in ensuring the performance, availability, and resilience of our mission-critical trading systems. As a Site Reliability Engineer, you will focus on the operational excellence of our .NET/C# services running on Windows, helping to build robust infrastructure that directly impacts our customer experience.
Key responsibilities
- Participate in on-call rotations for production trading systems and lead incident response during service disruptions.
- Investigate production incidents, perform root cause analysis, and implement preventive measures to eliminate recurring issues.
- Build and maintain Grafana dashboards and Prometheus alerts to monitor application and infrastructure health.
- Instrument .NET services to improve telemetry, logging, and visibility into service performance.
- Define and monitor Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to maintain high availability.
Requirements
- Strong experience debugging and supporting .NET/C# applications within Windows Server environments.
- Proficiency in PowerShell scripting and experience with Python or Bash for automation.
- Solid understanding of observability tools such as Grafana, Prometheus, and Loki.
- Hands-on experience with AWS infrastructure and Infrastructure as Code tools like Terraform.
- Practical knowledge of incident response, root cause analysis, and alert design.
What we offer
- Opportunity to work on mission-critical trading infrastructure with real-time impact.
- A chance to solve complex reliability and scalability challenges in a modern engineering culture.
- Collaboration with experienced engineers to influence reliability strategy and best practices.
- The ability to build world-class observability, automation, and deployment workflows.