Basis Research

Basis Research

ML Systems Engineer, Infrastructure & Cloud

New York Office · Full-time

Sponsorship not specifiedDetected 241 days ago
Cloud PlatformsMachine LearningPyTorchFirewallResearchCommunication

About the role

  • ML Systems Engineers at Basis ensure training and evaluation infrastructure is fast, reliable, and scalable.
  • We are looking for engineers who combine deep understanding of ML systems with operational excellence.
  • You will be the guardian of training stability, the optimizer of compute costs, and the enabler of reproducible research.

Responsibilities

  • Own distributed training infrastructure including job launchers, checkpointing systems, recovery mechanisms, and monitoring that ensures experiments run reliably at scale.
  • Profile and optimize training performance by identifying bottlenecks in data loading, gradient computation, communication overhead, and implementing solutions that improve step time.
  • Manage cloud infrastructure and costs including capacity planning, spot instance strategies, storage optimization, and building tools that give researchers visibility into resource usage.
  • Implement security and compliance measures including access controls, data encryption, audit logging, and ensuring infrastructure meets requirements for handling sensitive data.
  • Build evaluation and benchmarking infrastructure that enables consistent, reproducible measurement of model performance across different conditions and datasets.
  • Maintain development environments including containerization, dependency management, and tools that ensure researchers can reproduce results across different systems.
  • Document and share knowledge through runbooks, post-mortems, and training materials that help the team understand and operate ML infrastructure effectively.
  • Collaborate with researchers to understand requirements, suggest infrastructure solutions, and ensure systems support rather than constrain research goals.

Requirements

  • The second is to advance society's ability to solve intractable problems.
  • This means expanding the scale, complexity, and breadth of problems that we can solve today, and even more importantly, accelerating our ability to solve problems in the future.
  • Experience with on-premise GPU cluster management.
  • Knowledge of optimization theory and numerical methods.
  • Possess deep knowledge of distributed training frameworks including PyTorch/JAX distributed strategies (DDP, FSDP, ZeRO), gradient accumulation, mixed precision training, and checkpoint/recovery systems.

Skills

  • Contributions to ML frameworks or distributed training libraries.

Compensation

  • Non-Discrimination Notice
  • By submitting your application, you grant Basis permission to use your materials for both hiring evaluation and recruitment-related research and development purposes.
  • Your information may be processed in different countries, including the US.
  • You retain copyright while providing Basis a license to use these materials for the stated purposes.
  • Competitive salary.

Benefits

  • Develop monitoring and alerting systems that detect anomalies in training metrics, resource utilization, or system health, enabling rapid response to issues.

Company info

  • We are in the office four days a week.
  • Be prepared to attend multi-day Basis-wide in-person events.
  • In-person Policy: We are in the office four days a week. Be prepared to attend multi-day Basis-wide in-person events.

This listing is sourced directly from Basis Research's careers page and normalized into a canonical job model.