Saviynt

Saviynt

Principal Software Engineer, AI Platform Engineering

El Segundo, CA · Principal

Sponsorship not specified$240k-$260kDetected 84 days ago
ScalaSQLRedisVector DatabasesKubernetesPlatform EngineeringgRPCMachine LearningSparkAirflowdbtData EngineeringRAGCadenceCommunication

About the role

  • You set the architectural direction for how training data flows, evolves, and is governed across the AI Platform.
  • AI Data Lake on GCS: bucket layout, raw → silver → gold tier separation, CMEK encryption, lifecycle rules
  • Batch pipelines: Spark on Dataproc for TB-scale feature backfills, Iceberg compaction, and daily S3→GCS incremental sync

Responsibilities

  • build embedding generation pipelines that chunk, encode, and upsert document embeddings into the vector store; own the data refresh cadence and staleness SLAs for retrieval context
  • You define the standards ML engineers and scientists build on, and ensure every training signal is tenant-isolated, PII-free, and traceable from source to model.
  • RAG data pipeline: build embedding generation pipelines that chunk, encode, and upsert document embeddings into the vector store; own the data refresh cadence and staleness SLAs for retrieval context

Requirements

  • expose data platform services (feature serving, embedding upsert, schema validation) over HTTPS with mTLS and gRPC where low-latency streaming is required
  • 8+ years of data engineering at production scale across multiple companies
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience or equivalent military experience

Nice to have

  • Differential privacy or k-anonymity for ML training datasets
  • Open source contributions: Feast, Great Expectations, Apache Beam, or dbt
  • Familiarity with IAM / access governance data: entitlements, provisioning events, access graphs
  • Iceberg or Delta Lake at petabyte scale
  • Work on a large-scale, Kubernetes-based SaaS platform
  • Solve challenging cloud and reliability problems at scale
  • Flyte, Kubeflow Pipelines, Airflow, or Prefect - operated in production, ideally benchmarked two
  • Orchestration at scale: Flyte, Kubeflow Pipelines, Airflow, or Prefect - operated in production, ideally benchmarked two

Skills

  • Pgvector or Qdrant in production - index tuning, ANN search, embedding upsert pipelines
  • Vector databases: Pgvector or Qdrant in production - index tuning, ANN search, embedding upsert pipelines

Compensation

  • Competitive compensation, benefits, and growth opportunities

This listing is sourced directly from Saviynt's careers page and normalized into a canonical job model.