Epsilon Health

Epsilon Health

Research Engineer - ML Infrastructure

San Francisco, CA

Sponsorship not specifiedDetected 264 days ago
PythonDistributed SystemsBigQuerySnowflakeDatabricksAWSGCPCloud PlatformsDockerKubernetesMachine LearningPyTorchSparkAirflowData EngineeringMLOpsA/B TestingComplianceHIPAARadiologyResearch

About the role

  • This role requires someone who can move fluidly between model training, data engineering, ML systems, and production deployment.

Responsibilities

  • Build centralized data storage solutions with standardized formats (e.g., protobufs) that enable efficient retrieval and training across the organization.
  • Create model inference pipelines and evaluation frameworks that work seamlessly across research experimentation and production deployment.
  • Collaborate with researchers to rapidly prototype new ideas and translate them into production-ready code.
  • Own end-to-end delivery of ML systems from experimentation through deployment and monitoring.

Requirements

  • Strong Python skills and expertise in PyTorch or JAX
  • Hands-on experience with data pipeline technologies (e.g., Spark, Airflow, BigQuery, Snowflake, Databricks, Chalk) and schema design
  • Experience with distributed systems, cloud infrastructure (AWS/GCP), and containerization (Docker/Kubernetes)
  • Ability to move quickly and handle competing priorities in a fast-paced environment
  • 5+ years building ML infrastructure, data pipelines, or ML systems in production
  • Track record of building scalable data systems and shipping production ML infrastructure
  • Experience with reinforcement learning training pipelines (e.g., RLHF, reward modeling, or online learning systems)
  • Support A/B testing and experimentation workflows for model rollouts, including monitoring statistical significance and managing canary deployments.
  • Familiarity with vision-language models (VLMs) or multimodal architectures
  • Experience with medical imaging formats (DICOM) and healthcare data standards

Nice to have

  • Background in distributed training frameworks (PyTorch Lightning, DeepSpeed, Accelerate)
  • Familiarity with MLOps practices and model deployment pipelines
  • Experience with privacy-preserving data systems and HIPAA compliance

Benefits

  • Build and optimize distributed ML infrastructure for training foundation models on large-scale medical imaging datasets.
  • Design and implement robust data pipelines to collect, process, and store large-scale multimodal medical imaging data from both production traffic and offline sources.

Company info

  • We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics.
  • Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes.
  • We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place.

This listing is sourced directly from Epsilon Health's careers page and normalized into a canonical job model.