Stack AV

Stack AV

Staff Software Engineer, ML Platform

Pittsburgh, PA or Remote · Staff+ · Contract

Sponsorship not specifiedDetected 85 days ago
C++Machine LearningSparkAirflowData Engineering

About the role

  • In the ML Data Understanding team, our mission is to provide trusted and useful data to efficiently power all of Stack's ML applications end-to-end from mining to training to safety evaluation.
  • We work hand in hand with AV autonomy teams to provide cutting edge solutions to all their data needs, working across data engineering, mining, modeling and infrastructure.
  • In particular, we provide services to find (data mining), curate (datasets), annotate (data labeling), search and serve (high throughput data access) data for all ML needs.

Responsibilities

  • Build state-of-art multimodal data mining and semantic search solutions to power AV product development.
  • Develop data understanding platform infrastructure for real-time querying/vector databases and batch/stream processing using technologies like Ray, Spark, Lance, or similar.
  • Deliver end-to-end data mining solutions that span onboard (C++) and offboard (ML & Data Infra) infrastructure to accelerate AV product development.
  • Develop e2e solution for real-time semantic search services (text/images/videos) and vector DBs.
  • Build low latency/high throughput batch or stream processing pipelines.
  • Drive technical discussions across multiple orgs and deliver solutions on a timely basis.

Requirements

  • Experience in building ML models or infrastructure in domains such as autonomous vehicles, perception, and decision-making (desirable but not required).
  • Experience with model training, model optimization, or large data processing pipelines.
  • 6+ years of experience with:
  • Experience with both ML platforms and building ML-based applications (modeling experience is a bonus).
  • Proven track record of building scalable, reliable infrastructure in a fast-paced environment.
  • Ability to collaborate effectively across teams.
  • Experience building or using ML infrastructure for a large number of customer teams.
  • Deep understanding of design trade-offs with the ability to articulate those trade-offs and achieve alignment with others.
  • Multimodal data indexing and inference pipelines.
  • Building semantic search service, embedding generation for video/images and vector DB.
  • Large scale ML pipelines (Airflow/Flyte) and model optimization.
  • We are proud to be an equal opportunity workplace. We believe that diverse teams produce the best ideas and outcomes. We are committed to building a culture of inclusion, entrepreneurship, and innovation across gender, race, age, sexual orientation, religion, disability, and identity.

Nice to have

  • Prior experience in autonomous vehicles (AV) is a plus.

Benefits

  • We are building state of the art infrastructure to support machine learning training and inference workloads using OSS components such as Ray, Spark, Lance and Iceberg.

Company info

  • Semantic Search for Data Mining: We are building the infrastructure of a highly scalable semantic search service for multimodal data to find interesting events quickly and flexibly.
  • We are building the infrastructure of a highly scalable semantic search service for multimodal data to find interesting events quickly and flexibly.
  • As part of this mission, you would be setting the direction for and helping us build an inference service using the latest AI models & approaches.

Visa & Work Authorization

  • As such, this position may be contingent upon Stack AV verifying a candidate's residence, U.S. person status, and/or citizenship status

This listing is sourced directly from Stack AV's careers page and normalized into a canonical job model.