GPU Infrastructure SRE
Hanabi Ai Inc.
Other Engineering
Posted on Sep 15, 2026
About the Role
- Build and operate the infrastructure that keeps large-model training and online inference running reliably.
- You will support GPU clusters at 100+ card scale, respond to production incidents, and work closely with training and inference teams on stability, performance, and resource efficiency.
- This is a hands-on infrastructure role across Linux, Slurm, Kubernetes, Ceph, RDMA networking, GPU drivers, observability, and automation.
Responsibilities
- Provide infrastructure support for large-model training and online inference workloads, responding quickly to production incidents and operational issues.
- Deploy and operate 100+ GPU clusters, ensuring training and inference jobs run stably at scale.
- Maintain Slurm scheduling and Kubernetes platforms, optimizing resource allocation and multi-tenant isolation.
- Deploy, expand, and tune Ceph distributed storage for large-scale AI workloads.
- Operate RDMA networks such as InfiniBand and RoCE, plus GPU fleet components including DCGM, drivers, CUDA, and NCCL version management.
- Build automation tools, monitoring, and alerting systems that improve cluster stability and reduce debugging time.
Requirements
- Senior Linux operations background, with hands-on experience in GPU clusters or HPC environments.
- Experience operating 100+ GPU clusters that support large-model training or online inference workloads.
- Deep understanding of the Linux kernel, networking stack, and storage stack, with the ability to diagnose low-level performance bottlenecks.
- Expertise with Slurm for training workloads and production Kubernetes clusters for inference, including architecture design and performance tuning.
- Strong experience with Ceph distributed storage, including large-scale deployment, expansion, and performance tuning.
- Familiarity with RDMA networking such as InfiniBand or RoCE, and the ability to help debug NCCL collective communication issues.
- Strong Python and Shell scripting skills, with experience building automation platforms or operations tooling.
- Experience with Prometheus and Grafana monitoring systems, including large-scale cluster monitoring and alerting.
- Strong ownership for production reliability, including leading incident response, postmortems, and on-call responsibilities.
Nice to Have
- Ability to partner with training teams to diagnose performance bottlenecks across NCCL communication and storage I/O.
- Experience building GPU clusters from zero to one at 100+ to 1,000+ GPU scale.
- Familiarity with additional distributed storage systems such as MinIO, Weka, or Lustre.
- HPC operations experience and familiarity with parallel file systems.
- Open-source contributions or technical writing related to infrastructure, HPC, GPU operations, or reliability.
How to Apply
- Send your resume along with GitHub, personal project, or technical writing links to the contact below.
- For open-source contributions or past projects, direct links or short write-ups are welcome.
- Take-home tasks, if any, will be paid at a reasonable market rate.
- No requirements around years of experience or degree - we evaluate on technical depth and past work.