Senior AI (Infrastructure) Engineer

United States, New York, Illinois, Chicago
Permanent
Job ID: 2566

Job Description


[Up to c. $400k Comp Package | Office-Led Working]


Role Overview

We’re representing a global proprietary trading firm building a centralised AI capability designed to give the business greater control over its models, infrastructure and long-term AI strategy. The environment is highly technical, well-capitalised and intentionally greenfield, with the firm looking to reduce its dependence on external frontier-model providers over time. This Senior AI Infrastructure Engineer will take ownership of the model layer - building the foundations for open-weight model training, self-hosted inference, evaluation and intelligent model routing. You will join a small dedicated AI team with significant influence over the architecture, standards and technical direction of the platform...


Key Responsibilities

  • Build and operate a production model gateway routing inference across open-weight and commercial models, with fallback, rate controls and visibility across quality, latency and cost
  • Design end-to-end distillation pipelines that use frontier-model outputs to create high-quality training datasets for smaller, task-specific open models
  • Fine-tune and productionise open-weight models such as Llama, Qwen, Mistral or equivalent for internal workloads where greater control, performance or cost efficiency is required
  • Design, deploy and maintain self-hosted inference environments using vLLM, TGI, Triton or comparable technologies, with a particular focus on reliability, throughput and operational performance
  • Build and optimise GPU infrastructure across Kubernetes, including provisioning, scheduling, utilisation monitoring, workload management and troubleshooting of production inference workloads
  • Develop model evaluation and regression frameworks covering output quality, latency, throughput, cost and behavioural changes across model or infrastructure releases
  • Establish clear technical criteria for selecting between open and closed models, balancing performance, economics, latency, data sensitivity, operational complexity and long-term maintainability
  • Partner closely with AI, DevOps and infrastructure engineers to ensure the model platform can support future agentic workloads while helping define wider AI engineering standards, adoption patterns and governance


What You’ll Bring…

  • Around 6-10 years of software, machine learning or AI engineering experience, including strong production-level Python and evidence of owning technically demanding systems rather than primarily integrating third-party APIs
  • Meaningful production experience fine-tuning, post-training or distilling open-weight models, with the ability to explain training data, methodology, evaluation, failure modes and resulting production trade-offs in depth
  • Hands-on experience serving large language models within self-managed infrastructure using technologies such as vLLM, TGI, Triton or equivalent high-performance inference frameworks
  • Strong practical experience managing GPU-backed production environments, including provisioning, scheduling, capacity planning, utilisation monitoring, performance optimisation and operational troubleshooting
  • Deep familiarity with Kubernetes as an execution platform for GPU-intensive machine learning workloads, alongside an understanding of the infrastructure requirements behind reliable model serving
  • Proven experience building robust model evaluation, benchmarking and regression-testing approaches that can identify changes in quality, latency, cost and production behaviour before deployment
  • (Preferred) Experience with quantisation, PEFT, LoRA or other techniques for reducing training and inference cost while preserving sufficient model quality for production workloads
  • (Preferred) Exposure to model gateways, inference proxies, Hugging Face and the broader open-model ecosystem, or experience operating AI systems within financial services or another sensitive-data environment


...


Apply for this role

All fields marked with * are required.

I confirm I have a pre-existing right to work in the role’s location *
I require visa sponsorship now or will require it in the future

Back to Job Listings