Member of Technical Staff - Pre-Training Infra
CA, London, NY
On-site
Permanent
19 Applications
Job description
You will build distributed systems for frontier-model pre-training and operate large-scale training runs. You will optimize throughput, stability, memory, communication, and GPU utilization; maintain pipelines for datasets and checkpointing; work with researchers to productionize experiments; and debug bottlenecks across training stacks and runtimes.
Responsibilities
- Build and scale distributed training systems
- Design and operate large-scale foundation model training runs
- Develop infrastructure for training across thousands of GPUs
- Optimize training throughput stability and efficiency
- Productionize experimental training workflows with researchers
- Improve communication memory usage and GPU utilization
- Build and maintain training pipelines for datasets checkpointing and experiments
- Debug distributed training and GPU performance bottlenecks
Requirements
- Distributed training
- Machine learning
- Megatron
- DeepSpeed
- Model parallelism
- GPU optimization
- NCCL
- GPU communication
- Debugging
- Data pipeline
- Foundation model
Benefits
- Stock options
- Medical insurance
- Dental insurance
- Vision insurance
- Life insurance
- Annual wellness allowance
- Daily office lunch and dinner
- 22 weeks of paid parental leave
- Unlimited paid time off in the U.S.
- 30 days of vacation in the U.K.
- Visa sponsorship
- Regular off-sites
- Happy hours
- Team celebrations
Personalized Google Job Alerts
Never miss high-paying UK jobs — get instant alerts on Google
See newly verified UK job openings and transparent salary benchmarks on Google before other candidates apply.
Get Instant Job Alerts on Google
1-Click on Google
•
No Sign-up
•
100% Free
Is there something wrong with this job listing? Let us know.