BrightAI
Staff MLOps Engineer – ML Platform
Palo Alto, CA · Staff+
Sponsorship not specifiedDetected 64 days ago
PythonFastAPIGitElasticsearchRedshiftVector DatabasesAWSCloud PlatformsDockerTerraformCI/CDGitHub ActionsPrometheusGrafanaMachine LearningPyTorchSparkAirflowData AnalysisData EngineeringRAGLLMOpsMLOpsAI Orchestration
About the role
- Bright.AI is a high-growth Physical AI company transforming how infrastructure businesses interact with the physical world through intelligent automation.
- You'll work at the intersection of ML engineering, cloud infrastructure, and developer experience, designing scalable data/model workflows, CI/CD for ML, observability, and governance that turn ideas into durable, monitored ML services.
Responsibilities
- Design, build, and operate our ML/AI development platform on AWS -including Amazon SageMaker AI (Studio/Notebooks, Training/Processing/Batch Transform, Real‑Time & Async Inference, Pipelines, Feature Store) and supporting services.
- Implement Infrastructure‑as‑Code (e.g., Terraform) and workflow orchestration ( Step Functions, Airflow)
- optionally support EKS for training/inference.
- Build automated data pipelines with S3, Glue, EMR/Spark (PySpark), Athena/Redshift
- Implement CI/CD for ML (CodeBuild/CodePipeline or GitHub Actions): unit/integration tests, data contracts, model tests, canary/shadow deployments, and safe rollback.
- set SLOs and autoscaling, and optimize for cost/performance.
- Build monitoring & observability for production models and services (drift, performance, bias with SageMaker Model Monitor
- Implement Infrastructure‑as‑Code (e.g., Terraform) and workflow orchestration ( Step Functions, Airflow); optionally support EKS for training/inference.
- Build automated data pipelines with S3, Glue, EMR/Spark (PySpark), Athena/Redshift; add data quality (Great Expectations/Deequ) and lineage.
- Ship real‑time endpoints (SageMaker endpoints/FastAPI on Lambda/ECS/EKS) and batch jobs; set SLOs and autoscaling, and optimize for cost/performance.
Requirements
- Experience with experiment tracking & model registry (e.g., SageMaker Experiments/Model Registry or MLflow) and data versioning.
Nice to have
- Distributed training at scale (SageMaker Training, PyTorch DDP, Hugging Face on SageMaker).
- Data engineering at scale (e.g., Spark/EMR, Glue, Redshift).
- Observability stacks (e.g., Grafana), performance tuning, and capacity planning for ML services.
- LLMOps/RAG (Bedrock, vector databases, evals) as optional capabilities.
- least‑privilege IAM, VPC isolation/PrivateLink, encryption, secret management.
- Help integrate GenAI/Bedrock services where appropriate
- B.S. or M.S. in Computer Science, Electrical/Computer Engineering, or related field
- advanced degree a plus.
Benefits
- Bonus Qualifications
- Educational Background
- Strong foundation in machine learning systems, distributed computing, and data engineering; applied experience building production grade ML platforms.
Apply directly at BrightAI →Create a free account for alerts like thisView BrightAI immigration profile
This listing is sourced directly from BrightAI's careers page and normalized into a canonical job model.