Senior Engineering Manager - Cloud Platform & SRE
Job Description
[Up to c. $260k Base Salary + Significant Equity | On-Site Working]
Role Overview
We’re representing a venture-backed technology company delivering software used within mission-critical US civil aviation infrastructure. As a major new platform moves towards production, the organisation is strengthening the engineering leadership responsible for the cloud foundation, reliability model and operational capability supporting it.
This person will take ownership of the cloud platform and SRE organisation supporting that programme, covering both the evolution of an existing production platform and the build-out of new infrastructure for a highly regulated environment. The immediate priority is delivery - establishing the architecture, engineering systems and operating model needed to move quickly without compromising resilience or availability. It is a senior player-coach position with substantial organisational scope. You will inherit teams in the US and Europe, grow the US organisation significantly, develop the 24/7 reliability function and remain technically credible enough to make architectural decisions and work directly with engineers while the team scales.
*The position is based in Boston five days per week, with funded relocation available.
Role Snapshot
- Typically 9+ years across software engineering, platform engineering or SRE, including at least 3 years of meaningful engineering-management responsibility
- Own the architecture, delivery and production operation of the AWS platform supporting a large-scale civil aviation programme, including both existing services and new infrastructure
- Bring deep practical AWS and Kubernetes expertise, including multi-region design, infrastructure-as-code, container orchestration and high-availability production architecture
- Establish an engineering model capable of sustaining approximately four-nines availability, with strong observability, failure handling, incident response, postmortems and continuous reliability improvement
- Build and scale the US platform and SRE organisation from an initial small team, while leading through existing managers and establishing effective metrics, planning, feedback loops and operational processes
- Create the SRE and on-call model required for continuous 24/7 production support, balancing operational readiness with the pace demanded by a major delivery programme
- Strengthen CI/CD and progressive-delivery practices across the platform, including automated validation, controlled releases, rollback mechanisms and safe production change
- Lead the roadmap towards a FedRAMP environment and multi-region production deployment, partnering across engineering and leadership to translate regulatory requirements into workable infrastructure
- Develop shared engineering tooling, including practical use of modern AI-assisted development for code generation, review and testing, without losing engineering rigour or accountability
- (Preferred) A strong software-engineering or computer-science foundation before moving into SRE/platform engineering, plus experience in hyperscale, regulated, defence, critical-infrastructure or similarly demanding production environments
...
Apply for this role
All fields marked with * are required.