Procore Technologies
Staff ML Data Engineer (Datagrid)
US - California - San Francisco · Staff+
Sponsorship not specified$227k-$313kDetected 7 days ago
PythonCode ReviewSQLDatabricksAWSGCPCI/CDKafkaMachine LearningSparkAirflowData EngineeringA/B TestingResearchLeadershipCommunicationMentoring
About the role
- The primary goal of this role is to ensure that researchers and engineers can reliably discover, curate, transform, and operate on large‑scale datasets that move from experimentation to production.
- As a Staff ML Data Engineer, you'll work closely with ML researchers, applied ML engineers, and system architects to turn ambiguous research needs into scalable, production‑ready data pipelines.
- You'll remain deeply hands‑on while providing technical leadership in data architecture, quality, and operational excellence.
Responsibilities
- Act as the technical lead for data engineering efforts supporting frontier model research and applied ML systems.
- Design, build, and maintain scalable batch and streaming pipelines for multimodal data (e.g., documents, images, spatial metadata).
- Partner closely with researchers and architects to translate experimental workflows into reliable, repeatable data systems.
- Lead the development of dataset curation, versioning, and lineage workflows that support rapid experimentation and reproducibility.
- Mentor other engineers through code reviews, design discussions, and hands‑on collaboration.
Requirements
- Bachelor's or Master's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 8+ years of experience designing and operating complex data systems in production or research‑adjacent environments.
- Strong proficiency in SQL and Python
- experience with data‑intensive or distributed systems.
- Comfort operating in highly ambiguous problem spaces and collaborating closely with researchers and architects.
- Strong communication skills, with the ability to explain technical tradeoffs to both research and engineering audiences.
Nice to have
- Nice to have experience with technologies such as:
Skills
- Databricks, Spark, lakehouse architectures, cloud data warehouses
- Airflow, Dagster, data quality and lineage tools
- AWS or GCP, containerized data workloads, CI/CD, infrastructure‑as‑code
Compensation
- 227,332.00 - 312,581.50 USD Annual
- This role may also be eligible for Equity Compensation and/or Bonus Incentive Compensation.
- Procore is committed to offering competitive, fair, and commensurate compensation.
- Actual compensation will be based on a candidate's job-related skills, experience, education or training, and location.
Benefits
- Proven experience building scalable data pipelines that support machine learning training, evaluation, or inference workflows.
Company info
- Procore will consider for employment all qualified applicants, including those with arrest or conviction records, in accordance with the requirements of applicable federal, state, and local laws, including the City of Los Angeles' Fair Chance Initiative for Hiring Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act.
- A criminal history may have a direct, adverse, and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment: 1. appropriately managing, accessing, and handling confidential information including proprietary and trade secret information, as well as accessing Procore's information technology systems and platforms; 2. interacting with and occasionally having unsupervised contact with internal/external customers, stakeholders, and/or colleagues; and 3. exercising sound judgment.
- What we're looking for
Apply directly at Procore Technologies →Create a free account for alerts like thisView Procore Technologies immigration profile
This listing is sourced directly from Procore Technologies's careers page and normalized into a canonical job model.