Echo Labs

Echo Labs

ML Infrastructure Engineer

San Francisco · Senior

Sponsorship not specifiedDetected 174 days ago
PythonJavaGoC++Distributed SystemsCloud PlatformsDockerKubernetesMachine LearningPyTorchData EngineeringNLPA/B TestingElectrical EngineeringHardware DesignResearchCommunicationCollaboration

About the role

  • We are seeking a Senior Machine Learning Infrastructure Engineer to join our team.
  • This person will have significant ownership over the ML R&D platform, working closely with domain experts to architect new cloud infrastructure, data pipelines, and modeling flows.
  • The work will ultimately enable the development of cutting-edge models for neuroscientific discovery and neural decoding, empowering brain-computer interface technology to improve the lives of patients living with severe neurological conditions.

Responsibilities

  • Design and build systems ML cloud infrastructure to enable massive-scale modeling and analytics
  • Support diverse model exploration, hyperparameter optimization, pretraining, fine-tuning, and evaluation processes
  • Design and optimize scalable distributed training pipelines, with support for features such model sharding, cross-GPU communication, and real-time training monitoring
  • Create, operate, and maintain robust ML platforms and services across the model lifecycle
  • Build diverse and scalable data platforms
  • Create infrastructure and pipelines for ingesting internal and external datasets with varied shapes, formats, and associated metadata
  • Design and assess custom data formats for efficient storage and slicing of high-dimensional time-series data
  • Foster visibility and reproducibility within the company by maintaining extensive documentation of design decisions, evaluations of viable alternatives for selected solutions, pipeline assessments, etc.
  • Support ML R&D operations while preparing for eventual incorporation into product pipelines

Requirements

  • Bachelor's degree in Computer Science, Electrical Engineering, or a related technical discipline
  • 5+ years of industry experience in software engineering, large-scale data infrastructure, or systems ML
  • Extensive proficiency in Python
  • Familiarity with PyTorch
  • Experience working with distributed-training frameworks (e.g. FSDP, DeepSpeed, Megatron-LM, Ray, etc.)
  • Experience having technical ownership over at least one successfully implemented collaborative project

Nice to have

  • Advanced degree (MS or PhD) in Computer Science, Electrical Engineering, or a related technical discipline
  • Proficiency in C++, Go, CUDA, Rust, and/or Java
  • Experience in data engineering and systems ML for time-series data
  • Deep understanding of the fundamentals of distributed systems, including scalability, fault tolerance, monitoring, observability, scheduling, performance tuning, and resource management
  • Experience with cloud-native environments and orchestration (Kubernetes, Docker, etc.)
  • Experience scaling foundation-model training infrastructure or multi-cluster computing environments
  • An opportunity to work on exciting, cutting-edge projects to transform patients' lives in a highly collaborative work environment.
  • 401(k) program with matching contributions.

Compensation

  • Competitive compensation, including stock options.

Benefits

  • Competitive compensation, including stock options.
  • Comprehensive benefits package.
  • Create flexible and performant ML infrastructure
  • Design, build, and optimize massive-scale databases and data pipelines for scalable, flexible, and reliable data access

Equal opportunity

  • Employer
  • Echo Neurotechnologies is an Equal Opportunity Employer (EOE). We celebrate diversity and are committed to creating an inclusive environment for all employees.
  • Confidentiality
  • All applications will be treated confidentially. Applicants may be asked to sign an NDA after the initial stages of the interview process.

This listing is sourced directly from Echo Labs's careers page and normalized into a canonical job model.