Lambda

Lambda

Data Center Facility Telemetry & Controls Engineer

Remote, USA

Sponsorship not specified$136k-$380kDetected 16 days ago
PythonCloud PlatformsTerraformAnsiblePrometheusGrafanaRESTMachine LearningIncident ResponseStakeholder ManagementProcurementThermal AnalysisNetwork MonitoringResearch

About the role

  • The Data Center Facility Telemetry & Controls Management Engineer is a critical technical role responsible for the design, deployment, integration, and ongoing operation of Building Management Systems (BMS), Data Center Infrastructure Management (DCIM) platforms, and facility telemetry pipelines across Lambda's growing data center portfolio.
  • This engineer ensures that all facility systems - power, cooling, thermal, and environmental - are continuously monitored, alarmed, and controllable in real time, supporting the safe and efficient operation of high-density GPU deployments at rack densities of 136-380 kW per rack.

Responsibilities

  • Architect and manage BMS integration across colocation and Lambda-owned facilities, covering chillers, CRAHs, CDUs (Coolant Distribution Units), cooling towers, UPS systems, PDUs, and automatic transfer switches.
  • Collaborate with colocation partners (Equinix, Digital Realty, and others) to ensure telemetry data flows from provider BMS/EPMS into Lambda's monitoring stack.
  • Own the DCIM platform strategy and roadmap - evaluating, selecting, and implementing tooling for asset management, capacity planning, environmental monitoring, and power chain visibility.
  • Develop and maintain real-time dashboards for PUE, thermal performance, stranded capacity, and cooling system efficiency across all Lambda sites.
  • Build and maintain telemetry pipelines ingesting data from BMS, PDUs, in-rack sensors, CDUs, and network devices into centralized monitoring and alerting platforms (e.g., Prometheus, Grafana, InfluxDB, or equivalent).
  • Develop control strategies and setpoint frameworks for TCS (Thermal Control System) loops supporting direct liquid cooling at densities of 220-380 kW per rack.
  • Support design and construction coordination for liquid cooling infrastructure in new data center buildouts, ensuring BMS and controls readiness at Day 1.
  • Establish and maintain facility event management processes, including on-call response protocols for facility telemetry anomalies.
  • Lead root cause analysis for facility system failures and implement corrective actions to prevent recurrence.
  • Partner with the data center operations team to maintain and refine emergency response runbooks tied to BMS alerts and automated controls.

Requirements

  • 7+ years of experience in data center infrastructure engineering, with at least 4 years focused on BMS, DCIM, or controls systems in a hyperscale, colocation, or AI/HPC environment.
  • Hands-on experience designing and integrating BMS for mission-critical facilities including UPS, PDU, CRAH/CRAC, chiller plant, cooling tower, and liquid cooling (CDU/in-row) systems.
  • Strong working knowledge of industrial control protocols: BACnet IP/MS-TP, Modbus TCP/RTU, SNMP, DNP3, and modern API-based integrations.
  • Demonstrated experience with DCIM platforms (Nlyte, Sunbird, Vertiv TRELLIS, or equivalent) including deployment, configuration, and ongoing administration.
  • Experience with real-time telemetry stacks (Prometheus, InfluxDB, Grafana, or similar) applied to infrastructure monitoring use cases.
  • Strong understanding of data center power and cooling systems, including PUE optimization, thermal management, and redundancy architectures (2N, N+1).

Nice to have

  • Direct experience with direct liquid cooling (DLC) systems, CDU controls integration, and TCS loop management for high-density AI GPU deployments (100+ kW per rack).
  • Familiarity with OCP (Open Compute Project) hardware and telemetry standards.
  • Experience working with major colocation providers (Equinix, Digital Realty, CyrusOne, etc.) on BMS/EPMS integration and data sharing agreements.
  • Exposure to modular or edge data center deployments and associated controls considerations.
  • Background in scripting and automation (Python, Ansible, Terraform) applied to infrastructure management workflows.
  • Experience operating data centers at international scale, including Asia-Pacific or Southeast Asian markets.
  • Relevant certifications: CDCP, CDCE, ETA Data Center Specialist, or vendor-specific BMS/controls certifications.

Skills

  • Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence.
  • One person, one GPU.

Compensation

  • The annual salary range for this position has been set based on market data and other factors.
  • However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
  • We offer generous cash & equity compensation

Benefits

  • Competitive compensation including salary, equity, and comprehensive benefits.
  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
  • Drive continuous improvement in MTTR for facility-related events through better telemetry coverage and automated remediation.
  • Health, dental, and vision coverage for you and your dependents
  • Wellness and commuter stipends for select roles
  • Flexible paid time off plan that we all actually use

Company info

  • We offer generous cash & equity compensation
  • 401k Plan with 2% company match (USA employees)
  • Founded in 2012, with 500+ employees, and growing fast

Equal opportunity

  • Lambda is an Equal Opportunity employer.
  • Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
  • Equal Opportunity Employer

Visa & Work Authorization

  • ational origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law

This listing is sourced directly from Lambda's careers page and normalized into a canonical job model.