Matter Intelligence
Data/ML Infrastructure Engineer
San Francisco · Full-time
Sponsorship not specifiedDetected 117 days ago
PythonDistributed SystemsSQLPostgreSQLRedisAWSDockerKubernetesTerraformMachine LearningA/B TestingDesign SystemsResearch
About the role
- This role spans ingestion, processing, storage, compute, and serving, with a strong emphasis on reliability, observability, performance, and cost.
- You will work closely with research and product engineering to shorten iteration cycles, improve reproducibility, and raise the quality bar for production systems.
- You will define clear interfaces and operational standards that keep the platform trustworthy as data volume, model complexity, and product usage scale.
Responsibilities
- Design, build, and operate scalable data and ML infrastructure on AWS, including workloads running on Kubernetes
- Build and maintain systems for ingestion, processing, storage, and serving, with strong guarantees around data quality, correctness, and operational safety
- Partner closely with research to support perception model training and evaluation workflows, enabling faster experimentation and more reproducible iteration
- Build platform primitives for observability, data versioning, lineage, evaluation, reproducibility, and operational excellence
- Partner with product engineering to ensure data- and model-derived insights are accessible through reliable, low-latency serving and retrieval interfaces
- Design systems that enable efficient access patterns for customer-facing products, including search, indexing, and large-scale querying
- Meaningful experience building production data infrastructure, ML infrastructure, or distributed systems
- Experience building and operating systems on AWS
Requirements
- You have strong software engineering fundamentals and have built production systems where reliability, cost, and performance matter.
- You can reason clearly about distributed systems tradeoffs, and you have experience designing data-intensive infrastructure that other engineers depend on.
- You are comfortable working across data platform and ML platform concerns, and you understand how tightly coupled they become in production.
- You care about reproducibility, debuggability, and developer experience because you have seen how quickly they become bottlenecks.
- You work effectively across research and product teams.
- Familiarity with modern infrastructure and platform tooling, including Kubernetes, Docker, and Terraform
- Experience working with production storage and serving systems such as Postgres and Redis
- Familiarity with data and ML workflow tooling such as Metaflow
- You are energized by working close to the data, close to the models, and close to the product.
Nice to have
- Experience supporting ML training, evaluation, batch inference, or model deployment in production
- Familiarity with modern large-scale data patterns and tooling, including streaming, backfills, partitioning strategy, and schema evolution
- Exposure to perception, multimodal, or geospatial systems, especially where data originates from real sensors and is used in real products
- This is a full-time role based in San Francisco, CA.
- To comply with U.S. export regulations, applicants must be one of the following:
- A U.S. citizen or national
- A lawful permanent resident (green card holder)
- Eligible to obtain required authorizations from the U.S. Department of State
Compensation
- Our compensation and benefits package includes:
Benefits
- Early-stage equity package
- 100% employer-paid health, dental, and vision coverage
Visa & Work Authorization
- citizen or national
Apply directly at Matter Intelligence →Create a free account for alerts like thisView Matter Intelligence immigration profile
This listing is sourced directly from Matter Intelligence's careers page and normalized into a canonical job model.