Stack AV
Staff Software Engineer, ML Platform
Pittsburgh, PA or Remote · Staff+ · Contract
Sponsorship not specifiedDetected 85 days ago
C++Machine LearningSparkAirflowData Engineering
About the role
- In the ML Data Understanding team, our mission is to provide trusted and useful data to efficiently power all of Stack's ML applications end-to-end from mining to training to safety evaluation.
- We work hand in hand with AV autonomy teams to provide cutting edge solutions to all their data needs, working across data engineering, mining, modeling and infrastructure.
- In particular, we provide services to find (data mining), curate (datasets), annotate (data labeling), search and serve (high throughput data access) data for all ML needs.
Responsibilities
- Build state-of-art multimodal data mining and semantic search solutions to power AV product development.
- Develop data understanding platform infrastructure for real-time querying/vector databases and batch/stream processing using technologies like Ray, Spark, Lance, or similar.
- Deliver end-to-end data mining solutions that span onboard (C++) and offboard (ML & Data Infra) infrastructure to accelerate AV product development.
- Develop e2e solution for real-time semantic search services (text/images/videos) and vector DBs.
- Build low latency/high throughput batch or stream processing pipelines.
- Drive technical discussions across multiple orgs and deliver solutions on a timely basis.
Requirements
- Experience in building ML models or infrastructure in domains such as autonomous vehicles, perception, and decision-making (desirable but not required).
- Experience with model training, model optimization, or large data processing pipelines.
- 6+ years of experience with:
- Experience with both ML platforms and building ML-based applications (modeling experience is a bonus).
- Proven track record of building scalable, reliable infrastructure in a fast-paced environment.
- Ability to collaborate effectively across teams.
- Experience building or using ML infrastructure for a large number of customer teams.
- Deep understanding of design trade-offs with the ability to articulate those trade-offs and achieve alignment with others.
- Multimodal data indexing and inference pipelines.
- Building semantic search service, embedding generation for video/images and vector DB.
- Large scale ML pipelines (Airflow/Flyte) and model optimization.
- We are proud to be an equal opportunity workplace. We believe that diverse teams produce the best ideas and outcomes. We are committed to building a culture of inclusion, entrepreneurship, and innovation across gender, race, age, sexual orientation, religion, disability, and identity.
Nice to have
- Prior experience in autonomous vehicles (AV) is a plus.
Benefits
- We are building state of the art infrastructure to support machine learning training and inference workloads using OSS components such as Ray, Spark, Lance and Iceberg.
Company info
- Semantic Search for Data Mining: We are building the infrastructure of a highly scalable semantic search service for multimodal data to find interesting events quickly and flexibly.
- We are building the infrastructure of a highly scalable semantic search service for multimodal data to find interesting events quickly and flexibly.
- As part of this mission, you would be setting the direction for and helping us build an inference service using the latest AI models & approaches.
Visa & Work Authorization
- As such, this position may be contingent upon Stack AV verifying a candidate's residence, U.S. person status, and/or citizenship status
Apply directly at Stack AV →Create a free account for alerts like thisView Stack AV immigration profile
This listing is sourced directly from Stack AV's careers page and normalized into a canonical job model.