Описание
We are looking for a **Senior Infrastructure Engineer (GPU Platform)**
Requirements:
• Production experience with multi-node GPU training infrastructure
• Strong Linux, containers, CUDA, and NVIDIA GPU stack knowledge
• Hands-on experience with NCCL and InfiniBand or RoCE/RDMA troubleshooting
• Deep experience with Kubernetes or Slurm
• Experience with infrastructure automation and observability
• Experience diagnosing issues across training workloads, networking, storage, and GPU hosts
• Strong incident leadership and provider-facing communication skills
• English – Upper-Intermediate or higher
Would be a plus:
• Experience in an AI lab, HPC environment, or specialist GPU cloud
• PyTorch, Megatron, DeepSpeed, or other distributed-training frameworks
• Experience with parallel storage and checkpoint optimization
• Experience working with multi-provider GPU platforms
📩 Send your CV:
Доступно в источнике