LVT

LVT

Staff Data Engineer

Seattle, Washington, United States · Staff+ · Full-time

No sponsorship$172k-$221kDetected 30 days ago
PythonKafkaMachine LearningPyTorchSparkData EngineeringLLMsAgentic AIMLOpsEHR/EMRResearchLeadershipMentoringPlain Language

About the role

  • LVT's AI systems are only as good as the data behind them.
  • As we move toward Physical AI, the binding constraint shifts from model architecture to the data flywheel.
  • Every AI team trains and evaluates from a single stack that transforms data from the raw source through standardized, versioned, governed datasets.

Responsibilities

  • LVT is redefining how businesses operate in the physical world, moving beyond traditional security solutions to deliver AI-driven, actionable intelligence that makes sites smarter, safer, and more secure.
  • This is your chance to join a cutting-edge team that isn't just watching the world change, but actively building the technology that is changing it.
  • Own the end-to-end loop that converts raw edge telemetry and video into labeled training data, frozen evaluation sets and feeds model outputs back into the next round.
  • Build and own the pipelines that register raw source data, standardize it into a single well-defined schema, and join and aggregate it into curated datasets so every team trains, validates, and benchmarks from one consistent store through one reader, rather than copying and reformatting data per use case.
  • Own how labels and semantic annotations are appended to datasets without rewriting source data, then versioned, quality-checked, and served, partnering with annotation and data-operations teams on label production and verification while you own the dataset, storage, and serving side.
  • Own schema and content versioning so producers can evolve datasets without breaking consumers opt-in versions, append-without-rewrite for new fields, and the reader/writer indirection that lets data migrate underneath clients on a controlled rollout instead of forced lockstep migrations.
  • You own the data side of the contract that defines what a model consumes and emits and annotation, edge, and infrastructure teams.
  • You should be equally comfortable discussing dataset schema design, storage and partitioning trade-offs for multimodal data, versioning and migration strategy, and the governance controls that keep sensitive video and sensor data safe.
  • Data Flywheel Ownership: Own the end-to-end loop that converts raw edge telemetry and video into labeled training data, frozen evaluation sets and feeds model outputs back into the next round.

Requirements

  • Strong experience with medallion-style layered data architectures and modern table/lake formats (e.g. Iceberg, Delta, Parquet, or comparable), including schema evolution and dataset versioning.
  • Experience with large multimodal data video, image, sensor/telemetry and the storage and access patterns that make it queryable at scale (denesting, repartitioning, binary-inline vs. reference storage).
  • A track record of setting data-engineering direction and leveling up engineers (technical leadership
  • formal management not required).
  • Bachelor's or Master's in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • formal people management is not required.
  • Lakehouse Architecture: Strong experience with medallion-style layered data architectures and modern table/lake formats (e.g. Iceberg, Delta, Parquet, or comparable), including schema evolution and dataset versioning.
  • Technical Leadership: A track record of setting data-engineering direction and leveling up engineers (technical leadership

Nice to have

  • Streaming or near-real-time ingestion from edge/IoT sources into a data lake (e.g. Kafka, Lambda, EMR, or similar).
  • Append-without-rewrite and hash-indexed dataset techniques on open table formats, and dataset/feature-versioning systems.
  • Generative-AI data work: fine-tuning and evaluation dataset curation for LLMs/VLMs.
  • Exposing datasets to AI agents through MCP-style query interfaces, with semantic schema and plain-language documentation for retrieval.

Skills

  • Hands-on with the data side of ML frameworks PyTorch/Lightning dataloaders and Spark and strong Python knowledge.
  • Framework Integration: Hands-on with the data side of ML frameworks PyTorch/Lightning dataloaders and Spark and strong Python knowledge.

Compensation

  • The beginning annual salary range for this role is $171,900 - $221,000 USD and is determined by location, job-related experience, and education/training.

Benefits

  • We invest in our crew's health, families, and financial futures with a benefits package designed to support you inside and outside the office.
  • Full-time benefits include, but not limited to: Comprehensive health, dental and vision coverage, retirement benefits (401k match up to 4%), and flexible PTO.
  • All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.
  • Computer-vision / video annotation tooling and workflows (e.g. Encord, Labelbox, or similar).

Company info

  • Named one of the Financial Times' Fastest Growing Companies 2025 and #10 on the Inc.
  • 5000 Rocky Mountain Regional list for 2025.
  • Innovative Leadership: Our CEO, Ryan Porter, was named an EY Entrepreneur of the Year 2025, and our CTO, Steve Lindsey, was inducted into the Silicon Slopes CTO Hall of Fame in 2024.
  • Product & Software Excellence: We were named one of The Software Report's Top 100 Software Companies of 2023 and are a winner of the Security Today Govies Award for 2025.
  • We are seeking a Staff Data Engineer to own that flywheel end to end including logs, sensor telemetry, labels and annotations, evaluation and benchmark sets.
  • This is a senior individual-contributor and technical-leadership role; formal people management is not required.
  • You will partner closely with AI/ML research, the ML platform / MLOps function.
  • Layered Dataset Pipelines: Build and own the pipelines that register raw source data, standardize it into a single well-defined schema, and join and aggregate it into curated datasets so every team trains, validates, and benchmarks from one consistent store through one reader, rather than copying and reformatting data per use case.

Equal opportunity

  • EQUAL OPPORTUNITY EMPLOYER.
  • Must be authorized to work in the U.S. If reasonable accommodation is needed to participate in the job application or interview process, and/or to perform essential job functions, please reach out to your recruiter.

Visa & Work Authorization

  • Dataset Versioning: Own schema and content versioning so producers can evolve datasets without breaking consumers opt-in versions, append-without-rewrite for new fields, and the reader/writer indirection that lets data migrate underneath

This listing is sourced directly from LVT's careers page and normalized into a canonical job model.