Member of Technical Staff Compute Platform
NY, London, CA • Permanent • Competitive
Back to Job Search

Member of Technical Staff Compute Platform

New Easy Apply
NY, London, CA On-site Permanent 13 Applications
Competitive
Full-time
Posted 23 Sep 2026
Expires 23 Oct 2026

Job description

You will build and maintain tooling for automated remediation, topology-aware scheduling, capacity planning, and hardware debugging. You will improve cluster management for large GPU fleets, implement cluster-wide monitoring and benchmarking, and prepare infrastructure for larger GPU deployments, storage replication, and network performance.

Responsibilities

  • Build and maintain tools for automatic remediation, topology-aware scheduling, capacity planning, and hardware debugging
  • Design and improve the cluster management stack for large multi-GPU fleets
  • Implement cluster-wide monitoring and active performance benchmarking
  • Prepare infrastructure for next-generation GPU deployments and larger clusters
  • Develop multi-cloud storage, data replication, and GPU network capabilities
  • Co-design fault tolerance, node health checks, and remediation strategies

Requirements

  • Systems engineering experience focused on cluster-wide behavior and maintenance
  • Strong coding ability in systems or GPU infrastructure
  • Deep GPU hardware knowledge
  • NCCL knowledge
  • Kubernetes architecture experience
  • Cloud storage expertise across data centers
  • Experience handling datasets and checkpointing at scale

Benefits

  • Stock options
  • Medical, dental, vision, and life insurance
  • Annual wellness allowance
  • Daily in-office lunch and dinner
  • 22 weeks of paid parental leave
  • Unlimited paid time off in the U.S.
  • 30 days of vacation in the U.K.
  • Visa sponsorship support
  • Regular off-sites, happy hours, and team celebrations
Is there something wrong with this job listing? Let us know.