Senior Site Reliability Engineer - Cloud Platform
Job Description
[Up to c. $300k Comp Package | Hybrid Working - 4 Days In Office]
Role Overview
We’re representing a global commodities trading and investment organisation expanding the infrastructure capability behind its firmwide investment, trading and analytics technology platform. The environment combines cloud infrastructure, container platforms, internal applications, APIs and data-driven tooling used across business-critical workflows.
This hire will sit within a small Platform Infrastructure function and take substantial hands-on ownership across reliability engineering, AWS architecture, Kubernetes, infrastructure automation and production observability. A major part of the mandate is moving the platform towards more resilient, repeatable cloud-native operating patterns while improving how services are deployed, monitored and recovered. Alongside core SRE responsibilities, the role offers meaningful ownership of the organisation’s OpenTelemetry adoption, Kubernetes evolution and disaster recovery capability. It suits an experienced engineer who wants to remain deeply technical while influencing how critical infrastructure is designed and operated across the firm...
Role Snapshot
- Bring 6+ years of hands-on experience across Site Reliability Engineering, Platform Engineering, DevOps or production infrastructure, with evidence of owning business-critical systems rather than primarily supporting them
- Design and operate highly available AWS infrastructure, applying resilient patterns across compute, storage, networking, scaling, load balancing, backup and recovery
- Build and evolve Kubernetes-based platform infrastructure using strong practical experience with Kubernetes, Docker, Terraform and Git
- Own Infrastructure as Code and delivery automation, creating repeatable deployment patterns and improving CI/CD, testing, upgrades, patching and environment consistency
- Develop the observability platform across metrics, logs and distributed tracing, including continued adoption of OpenTelemetry, alongside tooling such as Datadog
- Establish measurable reliability practices covering SLOs, SLIs, error budgets, capacity, operational health and reduction of recurring production failure
- Strengthen incident management through effective alerting, troubleshooting, runbooks, root-cause analysis and engineering changes that prevent repeat incidents
- Design and validate high-availability and disaster recovery capabilities, including RTO/RPO targets, recovery procedures, failover testing and dependency planning
- Use Python, Go, Bash or comparable engineering automation to reduce manual infrastructure work, supported by strong Linux, networking and production troubleshooting fundamentals
- (Preferred) Experience with Argo CD, Helm, Karpenter, Crossplane, progressive delivery, serverless architectures, additional cloud platforms or security controls within regulated or financial-services environments
...
Apply for this role
All fields marked with * are required.