Basis Research
ML Systems Engineer, Infrastructure & Cloud
New York Office · Full-time
Sponsorship not specifiedDetected 241 days ago
Cloud PlatformsMachine LearningPyTorchFirewallResearchCommunication
About the role
- ML Systems Engineers at Basis ensure training and evaluation infrastructure is fast, reliable, and scalable.
- We are looking for engineers who combine deep understanding of ML systems with operational excellence.
- You will be the guardian of training stability, the optimizer of compute costs, and the enabler of reproducible research.
Responsibilities
- Own distributed training infrastructure including job launchers, checkpointing systems, recovery mechanisms, and monitoring that ensures experiments run reliably at scale.
- Profile and optimize training performance by identifying bottlenecks in data loading, gradient computation, communication overhead, and implementing solutions that improve step time.
- Manage cloud infrastructure and costs including capacity planning, spot instance strategies, storage optimization, and building tools that give researchers visibility into resource usage.
- Implement security and compliance measures including access controls, data encryption, audit logging, and ensuring infrastructure meets requirements for handling sensitive data.
- Build evaluation and benchmarking infrastructure that enables consistent, reproducible measurement of model performance across different conditions and datasets.
- Maintain development environments including containerization, dependency management, and tools that ensure researchers can reproduce results across different systems.
- Document and share knowledge through runbooks, post-mortems, and training materials that help the team understand and operate ML infrastructure effectively.
- Collaborate with researchers to understand requirements, suggest infrastructure solutions, and ensure systems support rather than constrain research goals.
Requirements
- The second is to advance society's ability to solve intractable problems.
- This means expanding the scale, complexity, and breadth of problems that we can solve today, and even more importantly, accelerating our ability to solve problems in the future.
- Experience with on-premise GPU cluster management.
- Knowledge of optimization theory and numerical methods.
- Possess deep knowledge of distributed training frameworks including PyTorch/JAX distributed strategies (DDP, FSDP, ZeRO), gradient accumulation, mixed precision training, and checkpoint/recovery systems.
Skills
- Contributions to ML frameworks or distributed training libraries.
Compensation
- Non-Discrimination Notice
- By submitting your application, you grant Basis permission to use your materials for both hiring evaluation and recruitment-related research and development purposes.
- Your information may be processed in different countries, including the US.
- You retain copyright while providing Basis a license to use these materials for the stated purposes.
- Competitive salary.
Benefits
- Develop monitoring and alerting systems that detect anomalies in training metrics, resource utilization, or system health, enabling rapid response to issues.
Company info
- We are in the office four days a week.
- Be prepared to attend multi-day Basis-wide in-person events.
- In-person Policy: We are in the office four days a week. Be prepared to attend multi-day Basis-wide in-person events.
Apply directly at Basis Research →Create a free account for alerts like thisView Basis Research immigration profile
This listing is sourced directly from Basis Research's careers page and normalized into a canonical job model.