Senior DevOps Engineer - Kubernetes Platform
Job Description
[Up to c. $300k Comp Package (or equivalent) | Office-Led Working]
Role Overview
We’re representing a highly successful proprietary trading firm that is expanding a small, high-impact platform engineering team responsible for the core infrastructure supporting trading and software engineering across the business. The team has recently designed and built a modern on-premises platform from the ground up, centred around Kubernetes, GitOps, observability and developer self-service, and is now entering its next stage of scale and adoption. This is a senior individual contributor role with genuine end-to-end ownership across cluster lifecycle, platform reliability, automation and tooling, rather than a traditional ticket-driven DevOps position. You’ll work closely with trading desks and engineering teams, helping evolve the platform, solve complex infrastructure problems and improve how internal users deploy, observe and operate production workloads...
Key Responsibilities
- Own and evolve production Kubernetes environments across cluster provisioning, lifecycle management, upgrades, networking, storage, RBAC, autoscaling and disaster recovery
- Help develop a multi-cluster and bare-metal Kubernetes estate, improving provisioning workflows, resilience, capacity management and operational consistency across the platform
- Operate and scale the firm’s observability environment across metrics, logs and distributed tracing using technologies such as Prometheus, Grafana, OpenTelemetry, Mimir, Loki and Tempo
- Improve monitoring standards by building dashboards, refining alerting, increasing telemetry coverage and helping engineering teams adopt consistent observability practices
- Design and maintain reusable CI/CD and GitOps capabilities using platforms such as GitLab CI, ArgoCD and Artifactory, improving deployment consistency and engineering velocity
- Build infrastructure automation and internal tooling across provisioning, configuration management, secrets integration and developer self-service using Terraform, Ansible, Python or Go
- Partner directly with trading desks and software engineering teams to onboard services, troubleshoot complex production issues and improve the reliability of P&L-sensitive workloads
- Contribute to the long-term operating model through documentation, runbooks, platform standards and automation that enables the team to scale without increasing manual operational dependency
What You’ll Bring...
- Around 6-9 years of experience across DevOps, Site Reliability Engineering, platform engineering or production infrastructure within technically demanding environments
- Deep hands-on Kubernetes experience, including cluster lifecycle, upgrades, production troubleshooting, networking, storage, RBAC, Helm and operational reliability
- Strong Linux engineering fundamentals across system services, networking, storage, resource utilisation, performance analysis and low-level production troubleshooting
- Production experience with observability platforms such as Prometheus, Grafana and Alertmanager, alongside a strong understanding of metrics, logging, tracing and OpenTelemetry
- Proven experience building or maintaining GitOps and CI/CD workflows using technologies such as ArgoCD, Flux, GitLab CI, GitHub Actions or comparable tooling
- Strong infrastructure automation experience using Terraform and/or Ansible, combined with practical Python or Go development for tooling, scripting and operational automation
- The communication skills to work directly with developers, infrastructure engineers and trading users, translating complex platform concepts into pragmatic engineering solutions
- (Preferred) Experience within proprietary trading, hedge funds, market making or another latency-sensitive environment, alongside exposure to bare-metal Kubernetes, advanced CNI networking, Kafka, Airflow, Kubeflow or similar orchestration technologies
...
Apply for this role
All fields marked with * are required.