C THE Signs

C THE Signs

Senior MLOps Engineer

United States · Senior

Sponsorship not specifiedDetected 141 days ago
PythonGCPDockerKubernetesCI/CDPlatform EngineeringMachine LearningLLMsRAGMLOpsAI OrchestrationDesign SystemsBudgetingManual TestingHIPAA

About the role

  • This role is ideal for someone who has shipped ML systems in production and is excited about LLM orchestration, RAG, evaluations, guardrails, and observability in a regulated environment.

Responsibilities

  • MLOps & ML Platform Design and operate ML platforms that support end-to-end workflows: data ingestion, feature engineering, training, evaluation, deployment, and monitoring.
  • Build and maintain CI/CD for ML (testing, packaging, versioning, reproducibility, automated rollbacks, approvals).
  • Implement MLOps best practices: model registry, experiment tracking, lineage, governance, and reproducible training environments.
  • Develop scalable training infrastructure (distributed training, GPU scheduling, cost controls, auto-scaling).
  • Create and maintain feature pipelines / feature stores, ensuring consistency between training and inference (training-serving skew prevention).
  • Build and own end-to-end LLM delivery pipelines: prompt/versioning, retrieval, orchestration, evaluation, deployment, monitoring, and iterative improvement.
  • Create robust LLM evaluation harnesses (offline + online): golden datasets, automated regression testing, human-in-the-loop review workflows, and risk scoring.
  • Build cost controls: token/cost budgeting, caching strategies, autoscaling, and performance tuning.
  • Deployment, reliability, and operations Productionize ML Models on GCP using containers and orchestration (e.g., GKE, Cloud Run), and build CI/CD for ML/LLM systems with automated tests and safe rollouts.

Requirements

  • Strong experience with GCP services and cloud-native patterns.
  • Experience with Vertex AI (pipelines, endpoints, feature store, model registry, evaluation) and/or managed vector search on GCP.
  • Experience with containerization and orchestration (Docker, Kubernetes/GKE and/or Cloud Run).

Nice to have

  • training pipelines, evaluation, deployment patterns, monitoring, and iteration loops.
  • Demonstrated hands-on experience with

Skills

  • token/cost budgeting, caching strategies, autoscaling, and performance tuning.

Benefits

  • Competitive salary and benefits package.
  • The opportunity to work on life-changing AI technology that directly impacts patient outcomes.
  • Join a team that combines cutting-edge innovation with a mission to save lives and improve health equity.
  • Continuous learning opportunities with access to the latest tools and advancements in AI and healthcare.
  • tracing, metrics, logs, dashboards, alerting for model/system health (latency, token usage, error rates, retrieval quality, hallucination indicators, drift where relevant).
  • Data, governance, and compliance (Healthcare) Design systems with security and privacy by default: IAM, least privilege, secrets management, audit logs, encryption, data retention, and PHI/PII handling.
  • Implement governance: model/prompt lineage, dataset provenance, evaluation traceability, and approval workflows aligned with healthcare compliance expectations.
  • Flexible working arrangements (remote or hybrid options available).

This listing is sourced directly from C THE Signs's careers page and normalized into a canonical job model.