Lambda

Lambda

Staff Product Manager - Observability

Bellevue Office · Staff+

Sponsorship not specifiedDetected 6 days ago
Node.jsDistributed SystemsCloud PlatformsPrometheusGrafanaDatadogSite Reliability EngineeringMachine LearningPyTorchProduct ManagementProduct StrategyResearchCommunication

About the role

  • A customer runs distributed training across 64 to 1,024+ GPUs (Graphics Processing Units).
  • When a job slows down, customers need one question answered instantly: is it my code, NCCL (NVIDIA Collective Communications Library) configuration issue, or a bad InfiniBand link?

Responsibilities

  • Own the Observability Roadmap: Define and drive the product strategy for observability across Lambda's cloud, from single On-Demand GPU Instances to 1-Click Clusters at 64 to 1,024+ GPU scale.
  • Ship Diagnostics That Answer the First Question: Build experiences that let a customer quickly distinguish their own code or configuration issues from platform problems such as a failing node or a degraded InfiniBand link.
  • Turn data and customer signal into a clear decision about what to build next, and can show examples where your insight changed a roadmap
  • Can write crisply, so a one-page document from you is enough to align a room
  • Are energized by ambiguity, and can be the first product manager to own a domain and give it shape
  • Build experiences that let a customer quickly distinguish their own code or configuration issues from platform problems such as a failing node or a degraded InfiniBand link.
  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

Requirements

  • Win over engineers, designers, executives, and customers without formal authority, and have shipped products that required cross-team adoption

Nice to have

  • Have shipped monitoring or observability products, such as those in the Datadog, Grafana, or Prometheus class
  • Have worked with HPC (high performance computing) or distributed systems telemetry, including interconnect and fabric-level metrics
  • Have built developer tools or API-first products
  • Have hands-on exposure to distributed training stacks such as PyTorch with NCCL, or to GPU fleet tooling such as DCGM (Data Center GPU Manager)
  • Have worked in a usage-based cloud infrastructure business
  • insight, influence, and execution.
  • But, a great idea doesn't mean anything in a vacuum.
  • That is where influence comes in.

Skills

  • Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence.
  • One person, one GPU.
  • is it my code, NCCL (NVIDIA Collective Communications Library) configuration issue, or a bad InfiniBand link?
  • Observability is how Lambda answers that question.

Compensation

  • The annual salary range for this position has been set based on market data and other factors.
  • However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
  • We offer generous cash & equity compensation

Benefits

  • GPU and cluster health, utilization, job-level telemetry, and InfiniBand fabric metrics.
  • Are comfortable in deeply technical conversations about distributed systems, and can hold your own with engineers on topics like metrics pipelines, hardware health, and failure modes

Company info

  • We offer generous cash & equity compensation
  • Health, dental, and vision coverage for you and your dependents
  • Wellness and commuter stipends for select roles
  • 401k Plan with 2% company match (USA employees)
  • Flexible paid time off plan that we all actually use
  • One of the key things customers want to know is if their application is causing issues or if it is Lambda.
  • If it is Lambda, they want to know if it is within the expected tolerance.
  • You will work with each product to define and productize SLIs, SLOs, and SLAs.
  • It lets customers distinguish their code, their configuration, and our platform in minutes, and that transparency is a core reason teams trust Lambda with their largest training runs.
  • You will partner daily with SRE (Site Reliability Engineering), fleet engineering, and the console and API (Application Programming Interface) teams to bring one coherent observability experience to customers and operators alike.
  • Insight means you look at the data, determine what it means for customers and business, and then figure out what to do about it.
  • But a great idea that everyone is excited about doesn't matter unless it is delivered to customers.
  • If you love turning raw telemetry into products customers rely on, and you want your work to be the reason a research team trusts their 1,024-GPU training run, we'd love to hear from you.

Equal opportunity

  • Lambda is an Equal Opportunity employer.
  • Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Visa & Work Authorization

  • ational origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law

This listing is sourced directly from Lambda's careers page and normalized into a canonical job model.