Saviynt
Principal Software Engineer, AI Platform Engineering
El Segundo, CA · Principal
Sponsorship not specified$240k-$260kDetected 84 days ago
ScalaSQLRedisVector DatabasesKubernetesPlatform EngineeringgRPCMachine LearningSparkAirflowdbtData EngineeringRAGCadenceCommunication
About the role
- You set the architectural direction for how training data flows, evolves, and is governed across the AI Platform.
- AI Data Lake on GCS: bucket layout, raw → silver → gold tier separation, CMEK encryption, lifecycle rules
- Batch pipelines: Spark on Dataproc for TB-scale feature backfills, Iceberg compaction, and daily S3→GCS incremental sync
Responsibilities
- build embedding generation pipelines that chunk, encode, and upsert document embeddings into the vector store; own the data refresh cadence and staleness SLAs for retrieval context
- You define the standards ML engineers and scientists build on, and ensure every training signal is tenant-isolated, PII-free, and traceable from source to model.
- RAG data pipeline: build embedding generation pipelines that chunk, encode, and upsert document embeddings into the vector store; own the data refresh cadence and staleness SLAs for retrieval context
Requirements
- expose data platform services (feature serving, embedding upsert, schema validation) over HTTPS with mTLS and gRPC where low-latency streaming is required
- 8+ years of data engineering at production scale across multiple companies
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience or equivalent military experience
Nice to have
- Differential privacy or k-anonymity for ML training datasets
- Open source contributions: Feast, Great Expectations, Apache Beam, or dbt
- Familiarity with IAM / access governance data: entitlements, provisioning events, access graphs
- Iceberg or Delta Lake at petabyte scale
- Work on a large-scale, Kubernetes-based SaaS platform
- Solve challenging cloud and reliability problems at scale
- Flyte, Kubeflow Pipelines, Airflow, or Prefect - operated in production, ideally benchmarked two
- Orchestration at scale: Flyte, Kubeflow Pipelines, Airflow, or Prefect - operated in production, ideally benchmarked two
Skills
- Pgvector or Qdrant in production - index tuning, ANN search, embedding upsert pipelines
- Vector databases: Pgvector or Qdrant in production - index tuning, ANN search, embedding upsert pipelines
Compensation
- Competitive compensation, benefits, and growth opportunities
Apply directly at Saviynt →Create a free account for alerts like thisView Saviynt immigration profile
This listing is sourced directly from Saviynt's careers page and normalized into a canonical job model.