Isomorphic Labs is building some of the largest foundation models in biotech to accelerate drug discovery, and is hiring a Software Engineer for its Compute Infrastructure team to drive fleet-wide compute efficiency across large-scale distributed GPU clusters. The role covers designing observability and telemetry systems to monitor hardware health and utilization, identifying and eliminating compute waste, improving accelerator utilization and ML run reliability, and hardening both research and production cloud environments (primarily GCP) run on Kubernetes.
Candidates should have hands-on experience operating large-scale AI/ML infrastructure, strong cloud compute design skills, Kubernetes deployment at scale, familiarity with Nvidia GPUs, and a track record building production observability and telemetry systems. The role is based in London with an expectation of three days per week in the office.
Text as published by the employer. Always confirm details on the employer's site.