Senior ML Infrastructure Engineer

United KingdomUnited Kingdom·OxfordFull-timesenior
OtherMl Infrastructure Engineer
0 views0 saves0 applied

Quick Summary

Key Responsibilities

Build, operate, and continuously optimise our high-performance GPU training and inference clusters, focusing on robust, high-availability scheduling, isolation, and automated lifecycle management.

Requirements Summary

Proven experience leading the design, build, and operation of high-performance ML compute clusters at scale A proactive,

Technical Tools
OtherMl Infrastructure Engineer

At the Ellison Institute of Technology (EIT), we’re on a mission to translate scientific discovery into real world impact. We bring together visionary scientists, technologists, policy makers, and entrepreneurs to tackle humanity’s greatest challenges in four transformative areas:

  • Health, Medical Science & Generative Biology
  • Food Security & Sustainable Agriculture
  • Climate Change & Managing CO₂
  • Artificial Intelligence & Robotics

This is ambitious work - work that demands curiosity, courage, and a relentless drive to make a difference. At EIT, you’ll join a community built on excellence, innovation, tenacity, trust, and collaboration, where bold ideas become real-world breakthroughs. Together, we push boundaries, embrace complexity, and create solutions to scale ideas for lab to society. Explore more at www.eit.org

Join our SciComp team to build the cloud and compute foundation that enables scientific breakthroughs. Deliver reliable, secure platforms and self-service guardrails that accelerate experimentation and turn ideas into results - faster, at scale, and with confidence. 

Responsibilities

~1 min read
  • →Build, operate, and continuously optimise our high-performance GPU training and inference clusters, focusing on robust, high-availability scheduling, isolation, and automated lifecycle management. 
  • →Drive systems design and implementation for high-throughput data paths, optimising I/O, caching, and data locality across compute and storage (including our current Lustre implementation). 
  • →Proactively benchmark, profile, and resolve performance bottlenecks across the compute, network, and orchestration layers to maximise efficiency for distributed training and inference. 
  • →Establish comprehensive observability, resilience, and automated security controls to ensure compliance and robust operation of sensitive research environments. 
  • →Partner with Research, Data, and Applied teams to forecast capacity and cost for GPU and storage needs, setting quotas and streamlining ML experimentation pipelines. 

Requirements

~1 min read
  • Proven experience leading the design, build, and operation of high-performance ML compute clusters at scale 
  • A proactive, autonomous approach to systems design and the proven ability and desire to ideate, co-create and implement optimal solutions 
  • Exposure to migrating or transforming ML infrastructure from traditional schedulers to modern, containerised systems 
  • Expertise with high-throughput storage systems for ML/HPC workloads 
  • Expert-level understanding of GPU architecture, high-speed networking for distributed training, and performance profiling to resolve bottlenecks 
  • A solid grasp of IaC and CI/CD practices (e.g., Terraform, Argo CD)

What We Offer

~1 min read
✓Competitive salary (dependent on experience) + travel allowance + bonus
✓Enhanced holiday. Our annual leave allowance is 25 days plus 8 bank holidays and an additional 3 days between Christmas and New Year. You will also have the opportunity to purchase an additional 5 days annual leave in January and July.
✓Pension - Employer contribution 7.5%, minimum employee contribution 5%
✓Life Assurance.
✓Income Protection
✓Private Medical Insurance as standard for you, your partner and any dependents. Including hospital Cash Plan
✓Employee discounts
✓Electric car scheme
✓Nursery Salary Sacrifice scheme
✓Cycle to Work Scheme
✓Family Planning
✓Neurodiversity support including advise and assessments
✓Coaching & Therapy services

You must have the right to work permanently in the UK with a willingness to travel as necessary. In certain cases, we can consider sponsorship, and this will be assessed on a case-by-case basis.

You will live in, or within easy commuting distance of, Oxford/London (or be willing to relocate)

The SciComp team work to a hybrid working pattern of 3 days in the office, 2x in our Oxford Office and 1x in our London office. You must be able to commit to this should you be successfully appointed.

Location & Eligibility

Where is the job
Oxford, United Kingdom
On-site at the office

Listing Details

Posted
September 8, 2026
First seen
September 29, 2026
Last seen
September 29, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
16%
Scored at
September 30, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Senior ML Infrastructure Engineer