Lead Network Reliability Engineer
Job Description
[Up to c. £225k Comp Package | Office-Led Working]
Role Overview
We’re representing a quantitative research and technology organisation operating a highly engineered infrastructure environment where network availability, performance and operational resilience have a direct impact on the wider technology platform.
This hire will provide hands-on technical leadership across network reliability engineering, with particular ownership of SRE practices, observability and automation. The mandate is to move the network estate further away from repetitive operational intervention and towards measurable, automated and engineering-led service management across both datacentre and office environments. You’ll remain close to the technology - improving telemetry, engineering operational tooling, shaping event-driven automation and leading the response to difficult production failures. It should suit a senior network engineer who combines strong routing, switching and network-security fundamentals with a genuine appetite for treating infrastructure operations as an engineering problem...
Role Snapshot
- Senior or lead-level network engineering experience within a large-scale, technically complex environment, with evidence of hands-on technical ownership rather than coordination alone
- Lead the reliability engineering approach for network and network-security services, improving availability, resilience, operational performance and service quality
- Apply SRE principles to networking through automation, stronger observability, repeatable engineering practices and systematic reduction of manual operational work
- Bring strong Cisco and Arista routing and switching expertise alongside practical knowledge of firewalls, IDS/IPS, segmentation and wider network-security controls
- Develop and extend network automation using Python and supporting engineering tooling such as Ansible and Jenkins
- Improve telemetry and service visibility using technologies including Prometheus, Grafana and OpenTelemetry, turning monitoring data into actionable operational insight
- Help evolve Infrastructure as Code and event-driven approaches for network changes, operational workflows and automated remediation
- Take technical ownership during major network or security incidents, driving investigation through root-cause analysis, corrective engineering and prevention of repeat failures
- Build stronger operational tooling, response procedures and engineering standards while influencing infrastructure, security and adjacent technology teams
- Bring experience with, or a strong technical interest in, AI-assisted and agentic approaches to network diagnostics, automation and operational remediation
...
Apply for this role
All fields marked with * are required.