Vumedi Inc.

Vumedi Inc.

DevOps Engineer

US OAK · Principal

Sponsorship not specifiedDetected 63 days ago
PythonPostgreSQLAWSCloud PlatformsDockerKubernetesTerraformCI/CDLinuxPrometheusDatadogDevOpsSite Reliability EngineeringMachine LearningAirflowData EngineeringLLMsCybersecurityNetwork SecurityIncident ResponseLeadershipMentoring

About the role

  • We are looking for a DevOps Engineer to join our engineering team and take ownership of our infrastructure, deployment processes, and overall platform reliability.
  • In this role, you will focus on improving our CI/CD pipelines, system reliability, and developer experience, while helping scale our cloud infrastructure in a secure and cost-efficient way.
  • You will work extensively with AWS services (compute, storage, networking, IAM, monitoring) and help ensure our systems are reliable, observable, and well-architected.

Responsibilities

  • Own and improve our infrastructure, CI/CD pipelines, and deployment processes across multiple environments
  • Collaborate closely with backend and data teams to support production systems, data pipelines, and overall platform reliability
  • Troubleshoot production issues, participate in incident response, and implement long-term fixes to improve system stability
  • Identify and drive improvements in performance, scalability, and cost efficiency across the platform
  • Support and scale AI/ML and LLM-based systems, ensuring reliable infrastructure for data processing and content classification workloads

Requirements

  • You have 5+ years of experience in DevOps, SRE, or infrastructure engineering, with a strong focus on cloud-native environments (preferably AWS)
  • You have managed cloud infrastructure (networking, IAM, compute, storage) with a strong understanding of security best practices and cost optimization
  • You have a strong understanding of monitoring, logging, and observability (e.g., Datadog, Prometheus, CloudWatch), and proactively identifying and resolve issues
  • You are comfortable debugging production issues across systems and collaborating with engineering teams to resolve them
  • You are proactive, take ownership, and enjoy working in environments with high autonomy and evolving processes
  • You are curious and motivated to learn, especially in areas like AI/ML infrastructure and large-scale systems
  • 5+ years of experience in DevOps, Site Reliability Engineering, or infrastructure-focused roles
  • Proven experience designing and operating scalable, reliable, and secure cloud infrastructure (preferably AWS) in production environments
  • Strong understanding of cloud security best practices (IAM, network security, secrets management), preferably within AWS
  • Proficiency in Python for automation, scripting, and tooling

Nice to have

  • Experience supporting or scaling AI/ML or LLM-based systems in production
  • You have worked with containerized applications (Docker) and are familiar with orchestration concepts (Kubernetes or ECS is a plus)
  • You are familiar with Infrastructure as Code principles (e.g., Terraform) and have experience implementing Infrastructure as Code from scratch in existing environments
  • You have experience working with or supporting backend systems and data platforms (e.g., Postgres, Airflow is a plus)
  • Background in backend engineering or software development
  • Experience working in a fast-paced startup or scale-up environment
  • Experience leading and mentoring engineers, while contributing to team-wide best practices

Skills

  • Build technology that matters in a fast-scaling

Company info

  • Your work directly impacts how doctors across the world learn and make decisions that save lives.
  • Grow as we grow: Be part of a company in an accelerated growth phase, where expanding teams, products, and markets create real opportunities for ownership, leadership, and career progression.
  • Build with AI: Work on applied LLM systems - from intelligent search to AI-driven content agents - and shape how AI transforms medical knowledge delivery.
  • Own your craft end-to-end: Take full responsibility for building systems that scale globally and power mission-critical workflows.
  • Collaborate globally: Join a world-class team of passionate engineers on modern tech stack which will further drive your career development.
  • Have real product impact: Influence the direction of product development by collaborating closely with product and leadership teams.
  • Be part of a company in an accelerated growth phase, where expanding teams, products, and markets create real opportunities for ownership, leadership, and career progression.

This listing is sourced directly from Vumedi Inc.'s careers page and normalized into a canonical job model.