Member of Technical Staff - Mid-Training Infra
CA, London, NY
On-site
Permanent
21 Applications
Job description
You will design, build, and operate GPU infrastructure for high-throughput inference and mid-training workloads. You will develop distributed systems for synthetic data generation and reinforcement learning, optimize model execution and GPU utilization, support large-scale evaluations, and resolve performance bottlenecks across runtimes, kernels, networking, and distributed compute.
Responsibilities
- Design build and operate large-scale GPU infrastructure
- Develop systems for synthetic data generation and reinforcement learning pipelines
- Build high-performance inference platforms across thousands of GPUs
- Optimize inference throughput latency and GPU utilization
- Support distributed reinforcement learning and model evaluation workloads
- Improve model execution through kernel optimization model parallelism and GPU runtime improvements
- Diagnose and resolve performance bottlenecks across distributed systems
Requirements
- GPU infrastructure
- Model serving
- Inference
- GPU optimization
- SGLang
- Megatron
- Reinforcement learning
- Distributed system
- Synthetic data
- GPU kernel
- Networking
Benefits
- Stock options
- Medical insurance
- Dental insurance
- Vision insurance
- Life insurance
- Annual wellness allowance
- Daily office lunch and dinner
- 22 weeks of paid parental leave
- Unlimited paid time off in the U.S.
- 30 days of vacation in the U.K.
- Visa sponsorship
- Regular off-sites
- Happy hours
- Team celebrations
Personalized Google Job Alerts
Never miss high-paying UK jobs — get instant alerts on Google
See newly verified UK job openings and transparent salary benchmarks on Google before other candidates apply.
Get Instant Job Alerts on Google
1-Click on Google
•
No Sign-up
•
100% Free
Is there something wrong with this job listing? Let us know.