BrightAI

BrightAI

Staff MLOps Engineer – ML Platform

Palo Alto, CA · Staff+

Sponsorship not specifiedDetected 64 days ago
PythonFastAPIGitElasticsearchRedshiftVector DatabasesAWSCloud PlatformsDockerTerraformCI/CDGitHub ActionsPrometheusGrafanaMachine LearningPyTorchSparkAirflowData AnalysisData EngineeringRAGLLMOpsMLOpsAI Orchestration

About the role

  • Bright.AI is a high-growth Physical AI company transforming how infrastructure businesses interact with the physical world through intelligent automation.
  • You'll work at the intersection of ML engineering, cloud infrastructure, and developer experience, designing scalable data/model workflows, CI/CD for ML, observability, and governance that turn ideas into durable, monitored ML services.

Responsibilities

  • Design, build, and operate our ML/AI development platform on AWS -including Amazon SageMaker AI (Studio/Notebooks, Training/Processing/Batch Transform, Real‑Time & Async Inference, Pipelines, Feature Store) and supporting services.
  • Implement Infrastructure‑as‑Code (e.g., Terraform) and workflow orchestration ( Step Functions, Airflow)
  • optionally support EKS for training/inference.
  • Build automated data pipelines with S3, Glue, EMR/Spark (PySpark), Athena/Redshift
  • Implement CI/CD for ML (CodeBuild/CodePipeline or GitHub Actions): unit/integration tests, data contracts, model tests, canary/shadow deployments, and safe rollback.
  • set SLOs and autoscaling, and optimize for cost/performance.
  • Build monitoring & observability for production models and services (drift, performance, bias with SageMaker Model Monitor
  • Implement Infrastructure‑as‑Code (e.g., Terraform) and workflow orchestration ( Step Functions, Airflow); optionally support EKS for training/inference.
  • Build automated data pipelines with S3, Glue, EMR/Spark (PySpark), Athena/Redshift; add data quality (Great Expectations/Deequ) and lineage.
  • Ship real‑time endpoints (SageMaker endpoints/FastAPI on Lambda/ECS/EKS) and batch jobs; set SLOs and autoscaling, and optimize for cost/performance.

Requirements

  • Experience with experiment tracking & model registry (e.g., SageMaker Experiments/Model Registry or MLflow) and data versioning.

Nice to have

  • Distributed training at scale (SageMaker Training, PyTorch DDP, Hugging Face on SageMaker).
  • Data engineering at scale (e.g., Spark/EMR, Glue, Redshift).
  • Observability stacks (e.g., Grafana), performance tuning, and capacity planning for ML services.
  • LLMOps/RAG (Bedrock, vector databases, evals) as optional capabilities.
  • least‑privilege IAM, VPC isolation/PrivateLink, encryption, secret management.
  • Help integrate GenAI/Bedrock services where appropriate
  • B.S. or M.S. in Computer Science, Electrical/Computer Engineering, or related field
  • advanced degree a plus.

Benefits

  • Bonus Qualifications
  • Educational Background
  • Strong foundation in machine learning systems, distributed computing, and data engineering; applied experience building production grade ML platforms.

This listing is sourced directly from BrightAI's careers page and normalized into a canonical job model.