Mistral AI

Mistral AI

Site Reliability Engineer - NYC

New York

Sponsorship not specifiedDetected 92 days ago
PythonGoDistributed SystemsDockerKubernetesTerraformCI/CDPrometheusGrafanaDevOpsSite Reliability EngineeringMachine LearningIncident ResponseResearchCommunicationCollaborationProblem Solving

About the role

  • We democratize AI through high-performance, optimized, open-source and cutting-edge models, products and solutions.
  • Our offerings include le Chat, the AI assistant for life and work.
  • Our diverse workforce thrives in competitive environments and is committed to driving innovation.

Responsibilities

  • Design, build, and maintain scalable, highly available and fault-tolerant infrastructures to support our web services and ML workloads
  • Implement and improve monitoring, alerting, and incident response systems to ensure optimal system performance and minimize downtime
  • Implement and maintain workflows and tools (CI/CD, containerization, orchestration, monitoring, logging and alerting systems) for both our client-facing APIs and large training runs
  • Participate occasionally in on-call rotations to respond to incidents and perform root cause analysis to prevent future occurrences Development
  • Drive continuous improvement in infrastructure automation, deployment, and orchestration using tools like Kubernetes, Flux, Terraform
  • Collaborate with AI/ML researchers to develop and implement solutions that enable safe and reproducible model-training experiments
  • Build a cloud-agnostic platform offering an abstraction layer between science and infrastructure
  • Design and develop new workflows and tooling to improve to the reliability, availability and performance of our systems (automation scripts, refactoring, new API-based features, web apps, dashboards, etc.)
  • Collaborate with the security team to ensure infrastructure adheres to best security practices and compliance requirements
  • Document processes and procedures to ensure consistency and knowledge sharing across the team

Requirements

  • Master's degree in Computer Science, Engineering or a related field
  • 7+ years of experience in a DevOps/SRE role
  • Strong experience with cloud computing and highly available distributed systems
  • Experience working against reliability KPIs (observability, alerting, SLAs)
  • Hands-on experience with CI/CD, containerization and orchestration tools (Docker, Kubernetes...)
  • Knowledge of monitoring, logging, alerting and observability tools (Prometheus, Grafana, ELK Stack, Datadog...)
  • Familiarity with infrastructure-as-code tools like Terraform or CloudFormation
  • Proficiency in scripting languages (Python, Go, Bash...) and knowledge of software development best practices
  • Strong understanding of networking, security, and system administration concepts

Compensation

  • What we offer ๐Ÿ’ฐ Competitive salary and equity ๐Ÿš‘ Healthcare: Medical/Dental/Vision covered for you and your family ๐Ÿ‘ด๐Ÿป 401K: 6% matching ๐Ÿ๏ธ PTO: 18 days ๐Ÿš— Transportation: Reimburse office parking charges, or $120/month for public transp

Benefits

  • Medical/Dental/Vision covered for you and your family ๐Ÿ‘ด๐Ÿป 401K: 6% matching ๐Ÿ๏ธ

Visa & Work Authorization

  • What we offer ๐Ÿ’ฐ Competitive salary and equity ๐Ÿš‘ Healthcare: Medical/Dental/Vision covered for you and your family ๐Ÿ‘ด๐Ÿป 401K: 6% matching ๐Ÿ๏ธ PTO: 18 days ๐Ÿš— Transportation: Reimburse office parking charges, or $120/month for public transp

This listing is sourced directly from Mistral AI's careers page and normalized into a canonical job model.