Mimecast

Mimecast

Senior ML Ops Engineer

United States of America - Ohio - Columbus · Senior

Sponsorship not specifiedDetected 30 days ago
PythonJavaBashSpringCode ReviewGitAWSCloud PlatformsDockerKubernetesTerraformCI/CDGitHub ActionsJenkinsGrafanaSite Reliability EngineeringMachine LearningPyTorchLLMsAgentic AILLMOpsMLOpsCybersecurityTest Automation

About the role

  • The AI Enablement Platform serves billions of requests per month across multiple regions, powering AI-driven capabilities in email security, insider risk, data loss prevention, and collaboration security for Mimecast's Human Risk Management platform.
  • This role sits at the intersection of infrastructure engineering and machine learning.
  • This is a senior individual contributor role.

Responsibilities

  • Self-Service Deployment Tooling: Design and build config-driven, validated workflows that enable ML Engineers to deploy models to AIP infrastructure without requiring hands-on ML Ops involvement for each release.
  • Platform Resilience and Scaling: Own the reliability and scalability of ML inference infrastructure.
  • Design and tune autoscaling policies against real production traffic patterns, implement rate limiting and backpressure mechanisms (HTTP 429, retry-after) at the API layer, and build request prioritization frameworks (real-time vs. batch) so the platform protects itself under load without manual intervention or consumer-side changes.
  • CI/CD and Automation: Design, implement, and maintain robust CI/CD pipelines for ML model and infrastructure deployments.
  • Infrastructure as Code: Manage all AIP infrastructure through Terraform and configuration management tooling. Maintain multi-region deployment capabilities and ensure infrastructure changes are reviewable, repeatable, and auditable.
  • Cost Optimization: Implement and enforce cost tagging and allocation at deployment time.
  • Optimize ML inference endpoints for cost-effectiveness, including right-sizing instance types, managing reserved capacity, and providing opinionated endpoint configuration recommendations based on model characteristics.
  • Agent and LLM Operations: Support the deployment and operational management of AI agents and LLM-based capabilities within the AIP's templatized agent framework.
  • Support compliance with regulatory requirements relevant to AI systems in the cybersecurity domain.
  • Participate in architectural reviews, contribute to platform governance, and drive engineering standards through documentation, code reviews, and design discussions.

Requirements

  • You are expected to drive architectural decisions, mentor other engineers, define standards, and operate with a high degree of autonomy.
  • Strong experience with AWS, particularly SageMaker (endpoints), EC2, ECS/EKS, SQS, S3, CloudWatch, and IAM.
  • Proficiency in Python, Java (Spring Boot), and Bash scripting.

Nice to have

  • Exposure to LLM serving infrastructure (model hosting, prompt management, token-level observability) and agentic AI deployment patterns.
  • Experience with cost allocation, FinOps, or cloud cost optimization for ML workloads.
  • Background in cybersecurity or experience operating AI systems in regulated or security-sensitive environments.
  • Familiarity with the Model Context Protocol (MCP) or similar agent-tool integration standards.
  • Join our AI Enablement Platform (AIP) team to accelerate your career journey, working with cutting-edge technologies and contributing to projects that have real customer impact.
  • You will be immersed in a dynamic environment that recognizes and celebrates your achievements.
  • Employees are expected to come to the office at least two days per week, because working together in person:
  • Introduces employees to priorities outside of their immediate realm.

Skills

  • This includes infrastructure for agent hosting, tool access configuration, and observability for agentic workloads.
  • Mentor ML Ops and ML Engineers on operational best practices.
  • Represent ML Ops in cross-functional planning with SRE, Cloud Platform, and consuming product teams.

Compensation

  • The base salary range for this position is $148,000−$222,000 plus benefits.

Benefits

  • Continuously monitor model performance, data drift, latency, error rates, and system health.
  • Fosters a culture of collaboration, communication, performance and learning.

Company info

  • your race, age, religion, sexual orientation, gender identity, ability, marital status, nationality, or any other protected characteristic won't affect your application.
  • If you require any adjustments or accommodations due to a disability, or any other reason that may help you in your interview process, please let us know by emailing careers@mimecast.com.
  • Due to certain obligations to our customers, an offer of employment will be subject to your successful completion of applicable background checks, conducted in accordance with local law.
  • It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment.
  • Our Mihra AI agent delivers 7x faster threat response for customers, and we're recognized as "Agents of Change" in Human Risk Management.

This listing is sourced directly from Mimecast's careers page and normalized into a canonical job model.