Member of Technical Staff Compute Platform
NY, London, CA
On-site
Permanent
13 Applications
Job description
You will build and maintain tooling for automated remediation, topology-aware scheduling, capacity planning, and hardware debugging. You will improve cluster management for large GPU fleets, implement cluster-wide monitoring and benchmarking, and prepare infrastructure for larger GPU deployments, storage replication, and network performance.
Responsibilities
- Build and maintain tools for automatic remediation, topology-aware scheduling, capacity planning, and hardware debugging
- Design and improve the cluster management stack for large multi-GPU fleets
- Implement cluster-wide monitoring and active performance benchmarking
- Prepare infrastructure for next-generation GPU deployments and larger clusters
- Develop multi-cloud storage, data replication, and GPU network capabilities
- Co-design fault tolerance, node health checks, and remediation strategies
Requirements
- Systems engineering experience focused on cluster-wide behavior and maintenance
- Strong coding ability in systems or GPU infrastructure
- Deep GPU hardware knowledge
- NCCL knowledge
- Kubernetes architecture experience
- Cloud storage expertise across data centers
- Experience handling datasets and checkpointing at scale
Benefits
- Stock options
- Medical, dental, vision, and life insurance
- Annual wellness allowance
- Daily in-office lunch and dinner
- 22 weeks of paid parental leave
- Unlimited paid time off in the U.S.
- 30 days of vacation in the U.K.
- Visa sponsorship support
- Regular off-sites, happy hours, and team celebrations
Personalized Google Job Alerts
Never miss high-paying UK jobs — get instant alerts on Google
See newly verified UK job openings and transparent salary benchmarks on Google before other candidates apply.
Get Instant Job Alerts on Google
1-Click on Google
•
No Sign-up
•
100% Free
Is there something wrong with this job listing? Let us know.