Cube
โ† All jobs
AN
Andromeda Bosbouw V.O.F.
San Francisco ยท United States

Staff SRE, AI Infrastructure

EngineeringFull TimeRemote-friendly
Found on jobs.ashbyhq.com ยท last checked yesterday

About this role

Description from Andromeda Bosbouw V.O.F.'s career page

Staff SRE, AI Infrastructure Location: North America Remote / San Francisco ยท Full-Time About Andromeda Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers. Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it's needed most. Our aim is to become a liquidity layer for global AI compute โ€” routing workloads across providers, GPU generations, and geographies the way financial markets route capital. We're a small, senior team where one engineer's judgment shapes every customer's experience. You'll join early enough to define how we run infrastructure at scale, work directly with the world's most demanding AI customers, and build a career operating at the frontier of what compute can do. The Role We're hiring a Staff SRE to own the reliability of Andromeda's infrastructure end to end โ€” from a node being racked and joined to a cluster, through the schedulers and control planes that place jobs on it, up to the customer-facing surface where a training run either succeeds or doesn't. We're looking for someone with multiple years of hands-on experience operating GPU infrastructure at scale. You read NVIDIA release notes the day they drop. You have war stories about NCCL, fabric topology choices, and what it takes to keep a multi-thousand-GPU run healthy. You move comfortably from a kernel-level perf trac
AI

More at Andromeda Bosbouw V.O.F.

Apply on company site โ†’