Torc Robotics

Torc Robotics

Senior, ML Engineer - VLM

Ann Arbor, MI, Remote - US · Senior

Sponsorship not specifiedDetected 14 days ago
DatabricksCloud PlatformsRESTMachine LearningDeep LearningPandasComputer VisionCadenceRoboticsElectrical EngineeringSensorsResearch

About the role

  • Meet The Team Torc is marching toward its AV 3.0 strategy, where end-to-end Vision-Language-Action (VLA) models perceive, reason, and act directly from sensor data.

Responsibilities

  • Own the offline dataset pipeline - design, implement, test, and deploy Cloud-based pipelines that convert logged multi-sensor data into VLM/VLA training datasets, spanning geometric labels (3D/2D detection, tracking, segmentation, depth) through semantic, scenario-level, and action/trajectory-grounded annotations.
  • Build VLM-assisted auto-labeling - develop open-vocabulary detection, dense captioning, semantic enrichment, and scene/scenario description generation that move beyond closed-set bounding boxes, using foundation models to scale annotation and cut manual labeling cost.
  • Generate reasoning-grounded labels - produce language-grounded reasoning and chain-of-causation style annotations, temporally aligned to ego-motion and trajectories, to support VLA training and explainable driving behavior.
  • Mine and curate the long tail - surface rare, difficult, and high-uncertainty scenarios, and build curated datasets that measurably improve downstream VLM/VLA model metrics rather than simply adding volume.
  • Partner with the end-to-end model team - co-define dataset specifications with VLM/VLA model developers, own the quality bar and delivery cadence, and operationalize a continuous dataset delivery loop into their training pipelines.
  • Scale on cloud infrastructure - build distributed, reproducible pipelines using columnar data formats and distributed compute, with disciplined software practices, version control, and documentation.
  • Lead and mentor - serve as project lead, guide less-experienced engineers, run design reviews, set coding and annotation standards, and drive alignment across team interfaces to the rest of the organization.
  • Scope of Influence: Expected to drive alignment across team interfaces to the rest of the organization. Designs, maintains, and owns team technical solutions and drives consensus. Mentors and guides engineers within the group.
  • Model Data Curation - building targeted datasets that measurably improve downstream model performance; large-scale Parquet data processing (Databricks, Daft, Pandas, etc.).
  • We run a continuous data flywheel - mine long-tail and failure cases, auto-label at scale, validate quality, and feed curated datasets directly into Torc's end-to-end VLM/VLA model development.

Nice to have

  • Bachelor's Degree in Computer Science, Robotics, Electrical Engineering, or related technical field plus competences typically acquired through 6+ years of experience
  • OR Master's Degree in a related technical field plus competences typically acquired through 3+ years of experience.

Benefits

  • Join us and catapult your career with the company that helped pioneer autonomous technology, and the first AV software company with the vision to partner directly with a truck manufacturer.
  • Torc is marching toward its AV 3.0 strategy, where end-to-end Vision-Language-Action (VLA) models perceive, reason, and act directly from sensor data.
  • Computer Vision & Deep Learning - model training and at least two of: 2D/3D Object Detection, Tracking, Sensor Fusion, Semantic Segmentation, BEV, Depth Estimation.
  • Multimodal / VLM experience - hands-on work with vision-language models, open-vocabulary or zero-shot recognition, dense captioning, or semantic embeddings / search applied to perception data.

Company info

  • At Torc, we have always believed that autonomous vehicle technology will transform how we travel, move freight, and do business.
  • Meet The Team
  • Sitting within Offline Perception, this team turns petabytes of logged multi-modal fleet data (images, kinematics) into VLM/VLA-ready datasets: geometric annotations, scenario-level semantic descriptions, action- and trajectory-grounded labels, and reasoning traces that explain why a maneuver was taken.

This listing is sourced directly from Torc Robotics's careers page and normalized into a canonical job model.