Echo Labs
ML Infrastructure Engineer
San Francisco · Senior
Sponsorship not specifiedDetected 174 days ago
PythonJavaGoC++Distributed SystemsCloud PlatformsDockerKubernetesMachine LearningPyTorchData EngineeringNLPA/B TestingElectrical EngineeringHardware DesignResearchCommunicationCollaboration
About the role
- We are seeking a Senior Machine Learning Infrastructure Engineer to join our team.
- This person will have significant ownership over the ML R&D platform, working closely with domain experts to architect new cloud infrastructure, data pipelines, and modeling flows.
- The work will ultimately enable the development of cutting-edge models for neuroscientific discovery and neural decoding, empowering brain-computer interface technology to improve the lives of patients living with severe neurological conditions.
Responsibilities
- Design and build systems ML cloud infrastructure to enable massive-scale modeling and analytics
- Support diverse model exploration, hyperparameter optimization, pretraining, fine-tuning, and evaluation processes
- Design and optimize scalable distributed training pipelines, with support for features such model sharding, cross-GPU communication, and real-time training monitoring
- Create, operate, and maintain robust ML platforms and services across the model lifecycle
- Build diverse and scalable data platforms
- Create infrastructure and pipelines for ingesting internal and external datasets with varied shapes, formats, and associated metadata
- Design and assess custom data formats for efficient storage and slicing of high-dimensional time-series data
- Foster visibility and reproducibility within the company by maintaining extensive documentation of design decisions, evaluations of viable alternatives for selected solutions, pipeline assessments, etc.
- Support ML R&D operations while preparing for eventual incorporation into product pipelines
Requirements
- Bachelor's degree in Computer Science, Electrical Engineering, or a related technical discipline
- 5+ years of industry experience in software engineering, large-scale data infrastructure, or systems ML
- Extensive proficiency in Python
- Familiarity with PyTorch
- Experience working with distributed-training frameworks (e.g. FSDP, DeepSpeed, Megatron-LM, Ray, etc.)
- Experience having technical ownership over at least one successfully implemented collaborative project
Nice to have
- Advanced degree (MS or PhD) in Computer Science, Electrical Engineering, or a related technical discipline
- Proficiency in C++, Go, CUDA, Rust, and/or Java
- Experience in data engineering and systems ML for time-series data
- Deep understanding of the fundamentals of distributed systems, including scalability, fault tolerance, monitoring, observability, scheduling, performance tuning, and resource management
- Experience with cloud-native environments and orchestration (Kubernetes, Docker, etc.)
- Experience scaling foundation-model training infrastructure or multi-cluster computing environments
- An opportunity to work on exciting, cutting-edge projects to transform patients' lives in a highly collaborative work environment.
- 401(k) program with matching contributions.
Compensation
- Competitive compensation, including stock options.
Benefits
- Competitive compensation, including stock options.
- Comprehensive benefits package.
- Create flexible and performant ML infrastructure
- Design, build, and optimize massive-scale databases and data pipelines for scalable, flexible, and reliable data access
Equal opportunity
- Employer
- Echo Neurotechnologies is an Equal Opportunity Employer (EOE). We celebrate diversity and are committed to creating an inclusive environment for all employees.
- Confidentiality
- All applications will be treated confidentially. Applicants may be asked to sign an NDA after the initial stages of the interview process.
Apply directly at Echo Labs →Create a free account for alerts like thisView Echo Labs immigration profile
This listing is sourced directly from Echo Labs's careers page and normalized into a canonical job model.