Voleon
Senior Cluster Site Reliability Engineer
Berkeley, CA · Senior
Sponsorship not specifiedDetected 96 days ago
PythonRubyAWSGCPCloud PlatformsDockerKubernetesTerraformAnsiblePrometheusGrafanaDevOpsSite Reliability EngineeringMachine LearningTensorFlowPyTorchSparkAirflowForecastingZero TrustResearchCollaboration
About the role
- We have become a multibillion-dollar asset manager, and we have ambitious goals for the future.
- Your colleagues will include internationally recognized experts in artificial intelligence and machine learning research as well as highly experienced finance and technology professionals.
Responsibilities
- Help software and research teams design policies around fair cluster usage, and help develop enforcement mechanisms for said policies
- Assist in forecasting cluster growth, and help select appropriate scale-up strategies. Help optimize operations across dimensions of cost and usability
- Ability to develop scripts and utilities of moderate complexity in a common scripting language (Python, Ruby, etc.)
- Build out custom observability mechanisms when off-the-shelf ones won't do
Requirements
- 5+ years of experience in SRE or DevOps roles, preferably working as a senior engineer or tech lead
- Familiarity with infrastructure-as-code and configuration management tools (Terraform, Ansible)
- Experience with cloud infrastructure (AWS or GCP)
- Familiarity designing and implementing modern observability stacks (Prometheus, Grafana, Loki, ELK, OpenTelemetry)
- Experience with distributed storage technologies (Lustre, Ceph, S3)
- Bachelor degree in computer science
- Ensure a high degree of cluster uptime (measured in multiple nines), and define + track SLAs to quantify reliability
Nice to have
- Hands-on experience with HPC frameworks (Slurm, Grid Engine) and Kubernetes-based job orchestrators (Airflow, Kueue, Kubeflow Pipelines), along with other distributed computing frameworks (Ray, Modin, Dask, Spark)
- Familiarity with ML frameworks (PyTorch/Tensorflow, JAX, Horovod, DeepSpeed)
- Experience with containerization (Docker, Podman, Singularity), particularly for HPC/batch compute environments
- Experience with HPC networking (InfiniBand, RDMA)
- Solid security/IAM foundations (Identity management systems, AWS/GCP IAM, Zero Trust)
- "FRIENDS OF VOLEON" CANDIDATE REFERRAL PROGRAM
Skills
- Voleon is a technology company that applies state-of-the-art AI and machine learning techniques to real-world problems in finance.
- Your work will provide a world-class HPC platform for researchers to focus on cutting-edge machine learning problems at scale.
Compensation
- In addition to our enriching and collegial working environment, we offer highly competitive compensation and benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.
Benefits
- Develop robust metrics and observability for cluster health and use those metrics to inform your work.
- Knowledge of HPC/batch compute frameworks (Slurm, Kueue, AWS/GCP Batch) and/or machine learning training systems (Kubeflow, MLflow, Horovod)
Equal opportunity
- The Voleon Group is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
- The Voleon Group is an Equal Opportunity employer.
- Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
Visa & Work Authorization
- ational origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law
This listing is sourced directly from Voleon's careers page and normalized into a canonical job model.