CoreWeave

CoreWeave

Senior Software Engineer - AI Infrastructure Performance Insights & Observability

Sunnyvale, CA / Bellevue, WA · Senior · Full-time

Sponsorship not specified$182k-$242kDetected 1 day ago
PythonGoC++SpringDistributed SystemsKubernetesCI/CDPrometheusGrafanaMachine LearningSparkData EngineeringData VisualizationLLMsCollaborationMicrosoft Office

Stay score

odds of building a lasting career here

61Sponsors, lottery-bound
Cap-exempt (no lottery)0
Sponsors this role90
Entry-level history0
PERM / green-card track0
Lottery odds (Level IV)94
Fits your clock70

Sponsors, but it's cap-subject — you still face the weighted lottery (~61% per draw at Level IV). Good if you win; have a cap-exempt backup on your list.

Lottery odds assume a STEM candidate.

Personalize to your clock →

Employer immigration record

from this employer's Department of Labor filings

Files H-1B transfers

53 transfer filings in the last year, covering 53 workers. Median labor-condition decision: 7 days. An employer that already files transfers is one that can take over an existing H-1B.

Sourced from Department of Labor LCA, PERM and prevailing-wage disclosure data. Employer matching is by name, so figures may be split across an employer's legal entities. Absence of a filing means none appears in our copy of the data, not that none exists.

Community outcomes

No reports yet — be the first to help the next applicant.

About the role

  • This is not a straightforward data engineering or BI role.

Responsibilities

  • Time-Series & Metrics Infrastructure - Own and extend our time-series database (TSDB) layer as the backbone of real-time observability.
  • Write and optimize PromQL/MetricsQL queries that power alerting, anomaly detection, and trend analysis across thousands of GPUs and hundreds of benchmark runs.
  • Query Optimization & Performance - Profile and tune query engines against columnar and time-series stores so that the observability layer meets its own strict P99 latency and freshness SLAs. Benchmark the benchmarking infrastructure itself.
  • BI & Reporting (secondary) - Where needed, build self-service views (Grafana, Looker, or similar) for engineers, product managers, and executives, but as a downstream output of the insight and observability work above, not the primary deliverable.
  • Working knowledge of time-series databases and fluency in PromQL or MetricsQL for building real-time alerting and anomaly detection, not only historical dashboards.
  • Familiarity with data lake architectures and modern table formats (Iceberg, Parquet, Avro) sufficient to support an insights platform, though this is not the primary skill this role is hiring for.
  • Strong communicator comfortable collaborating with cross-functional teams and external partners.
  • Build the detection and diagnosis logic that surfaces anomalies and regressions before they become incidents, not just dashboards that report on them after the fact.
  • Design and build the systems that continuously assess
  • Own and extend our time-series database (TSDB) layer as the backbone of real-time observability.

Requirements

  • 5+ years of experience building distributed systems, observability platforms, or performance engineering tooling, ideally for infrastructure or ML systems rather than general-purpose BI.
  • Hands-on experience with Kubernetes at production scale, CI/CD, and observability stacks (Prometheus, Grafana, OpenTelemetry) used to monitor and diagnose infrastructure, not just report on it.

Nice to have

  • Experience with time-series databases, LSM-based storage engines, or custom telemetry pipelines.
  • Experience running MLPerf submissions or similar large-scale audited benchmarks.
  • Contributions to OSS projects such as Apache Iceberg, Apache Spark, Trino, llm-d, vLLM, or PyTorch.
  • Direct experience benchmarking or monitoring large GPU fleets or multi-region clusters.
  • Experience with CUDA kernels, NCCL/SHARP, RDMA/NUMA, or GPU interconnect topologies.
  • Familiarity with data cataloging, lineage tools, or data governance frameworks.
  • We're in an exciting stage of hyper-growth that you will not want to miss out on.
  • Be Curious at Your Core

Skills

  • CoreWeave is The Essential Cloud for AI™.
  • CRWV) in March 2025.
  • Learn more at www.coreweave.com.

Compensation

  • The base salary range for this role is $182,000 to $242,000.

Benefits

  • In addition to a competitive salary, we offer a variety of benefits to support your needs.
  • The benefits below reflect our US-based offerings for full-time employees
  • Medical, dental, and vision insurance - 100% paid for by CoreWeave
  • Company-paid Life Insurance
  • Voluntary supplemental life insurance
  • Short and long-term disability insurance
  • Flexible Spending Account
  • Health Savings Account
  • Tuition Reimbursement
  • Ability to Participate in Employee Stock Purchase Program (ESPP)
  • Mental Wellness Benefits through Spring Health
  • Paid Parental Leave
  • for roles in other locations, benefits vary and are shared during the hiring process.
  • The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process.

Company info

  • Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025.
  • At CoreWeave, we work hard, have fun, and move fast!

Visa & Work Authorization

  • Export Control Compliance

This listing is sourced directly from CoreWeave's careers page and normalized into a canonical job model.