About the role
Ellison Institute of Technology's SciComp team is hiring a Senior ML Infrastructure Engineer to build and operate the high-performance GPU training and inference clusters underpinning EIT's research. Responsibilities include optimizing high-throughput data paths and I/O/caching across compute and storage, benchmarking and resolving bottlenecks across compute/network/orchestration layers, and establishing observability, resilience and security controls for research environments, plus GPU/storage capacity planning with other teams.
Requires demonstrated experience operating ML compute clusters at scale, migrating ML infrastructure to containerized systems, and advanced knowledge of GPU architecture, high-speed networking for distributed training, and infrastructure-as-code/CI-CD tooling (e.g. Terraform, Argo CD). Hybrid role based between EIT's Oxford and London offices (3 days/week in office).
Text as published by the employer. Always confirm details on the employer's site.