Hyphen Connect Limited

Hyphen Connect Limited

LLM Pre-training & Distributed Engineer (AI Infrastructure)

Boston, USA

Sponsorship not specifiedDetected 89 days ago
PythonC++Distributed SystemsKubernetesMachine LearningPyTorchLLMsSystems Engineering

About the role

  • This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure.
  • The ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure efficient and reliable training processes.

Responsibilities

  • Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors.

Requirements

  • Required Skills:

Skills

  • Deep expertise in 3D parallelism (Data, Tensor, Pipeline).
  • Experience managing SLURM or Kubernetes-based GPU clusters.
  • Strong systems engineering background (C++, CUDA, Python).

Company info

  • We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer.

This listing is sourced directly from Hyphen Connect Limited's careers page and normalized into a canonical job model.