ReflectionAI

ReflectionAI

Member of Technical Staff - Compute Platform

New York · Staff+

H1B sponsorship availableDetected 123 days ago
Node.jsKubernetesPlatform EngineeringResearch

About the role

  • Reflection's Compute Platform team specializes in keeping our compute layer healthy and highly available.
  • We run a K8s-based platform distributed across multiple neo-clouds.

Responsibilities

  • Cluster Management: Build and maintain tools for the automatic remediation, topology-aware scheduling, capacity planning and rapid hardware debugging.
  • Platform Engineering: Design and iterate on our cluster management stack for workloads across large, multi-GPU fleets
  • Monitoring & Observability: Implement comprehensive cluster-wide monitoring, focusing on durability and active performance benchmarking.
  • Roadmap Execution: Prepare the infrastructure for next-generation GPU deployments and increasingly larger cluster sizes. Long-term, you will help own multi-cloud storage, petabyte-scale data replication, and GPU-to-GPU network performance.
  • Reflection is a research lab making intelligence open and accessible for everyone to use, customize, and build on.
  • We build open models that let anyone control their intelligence and help shape the future of AI.
  • Design and iterate on our cluster management stack for workloads across large, multi-GPU fleets
  • Implement comprehensive cluster-wide monitoring, focusing on durability and active performance benchmarking.
  • Sponsorship support: We sponsor visas to help exceptional talent join our team and support long-term immigration pathways where applicable.
  • Team building: We have regular off-sites, happy hours, and team celebrations.

Requirements

  • Systems-level engineering experience with a focus on cluster-wide behavior and maintenance.
  • Cloud storage expertise, specifically managing high-performance data products (like VAST) across multiple data centers, connecting those storage environments together and handling datasets and checkpointing at scale.

Skills

  • Prepare the infrastructure for next-generation GPU deployments and increasingly larger cluster sizes.

Compensation

  • Salary and equity structured to recognize and retain our talent globally.
  • Health & wellness: Comprehensive medical, dental, vision, and life, with an annual wellness allowance.

Benefits

  • Top-tier compensation: Salary and equity structured to recognize and retain our talent globally.
  • Stock options: Everyone who joins and contributes to Reflection's success gets to share in the upside through stock options.
  • Health & wellness: Comprehensive medical, dental, vision, and life, with an annual wellness allowance.
  • Life & family: 22 weeks paid parental leave for all new birthing and non-birthing parents, including adoptive and surrogate journeys.
  • Vacation days: Unlimited paid time off in the U.S. and 30 days in the U.K.

Company info

  • make intelligence open and accessible to all.

Visa & Work Authorization

  • We sponsor visas to help exceptional talent join our team and support long-term immigration pathways where applicable.

This listing is sourced directly from ReflectionAI's careers page and normalized into a canonical job model.