Mimecast
Senior ML Ops Engineer
United States of America - Ohio - Columbus · Senior
Sponsorship not specifiedDetected 51 days ago
PythonJavaBashSpringCode ReviewGitAWSCloud PlatformsDockerKubernetesTerraformCI/CDGitHub ActionsJenkinsGrafanaSite Reliability EngineeringMachine LearningPyTorchLLMsAgentic AILLMOpsMLOpsCybersecurityTest Automation
About the role
- The AI Enablement Platform serves billions of requests per month across multiple regions, powering AI-driven capabilities in email security, insider risk, data loss prevention, and collaboration security for Mimecast's Human Risk Management platform.
- This role sits at the intersection of infrastructure engineering and machine learning.
- This is a senior individual contributor role.
Responsibilities
- Self-Service Deployment Tooling: Design and build config-driven, validated workflows that enable ML Engineers to deploy models to AIP infrastructure without requiring hands-on ML Ops involvement for each release.
- Platform Resilience and Scaling: Own the reliability and scalability of ML inference infrastructure.
- Design and tune autoscaling policies against real production traffic patterns, implement rate limiting and backpressure mechanisms (HTTP 429, retry-after) at the API layer, and build request prioritization frameworks (real-time vs. batch) so the platform protects itself under load without manual intervention or consumer-side changes.
- CI/CD and Automation: Design, implement, and maintain robust CI/CD pipelines for ML model and infrastructure deployments.
- Infrastructure as Code: Manage all AIP infrastructure through Terraform and configuration management tooling. Maintain multi-region deployment capabilities and ensure infrastructure changes are reviewable, repeatable, and auditable.
- Cost Optimization: Implement and enforce cost tagging and allocation at deployment time.
- Optimize ML inference endpoints for cost-effectiveness, including right-sizing instance types, managing reserved capacity, and providing opinionated endpoint configuration recommendations based on model characteristics.
- Agent and LLM Operations: Support the deployment and operational management of AI agents and LLM-based capabilities within the AIP's templatized agent framework.
- Support compliance with regulatory requirements relevant to AI systems in the cybersecurity domain.
- Participate in architectural reviews, contribute to platform governance, and drive engineering standards through documentation, code reviews, and design discussions.
Requirements
- You are expected to drive architectural decisions, mentor other engineers, define standards, and operate with a high degree of autonomy.
- Strong experience with AWS, particularly SageMaker (endpoints), EC2, ECS/EKS, SQS, S3, CloudWatch, and IAM.
- Proficiency in Python, Java (Spring Boot), and Bash scripting.
Nice to have
- Exposure to LLM serving infrastructure (model hosting, prompt management, token-level observability) and agentic AI deployment patterns.
- Experience with cost allocation, FinOps, or cloud cost optimization for ML workloads.
- Background in cybersecurity or experience operating AI systems in regulated or security-sensitive environments.
- Familiarity with the Model Context Protocol (MCP) or similar agent-tool integration standards.
- Join our AI Enablement Platform (AIP) team to accelerate your career journey, working with cutting-edge technologies and contributing to projects that have real customer impact.
- You will be immersed in a dynamic environment that recognizes and celebrates your achievements.
- Employees are expected to come to the office at least two days per week, because working together in person:
- Introduces employees to priorities outside of their immediate realm.
Skills
- This includes infrastructure for agent hosting, tool access configuration, and observability for agentic workloads.
- Mentor ML Ops and ML Engineers on operational best practices.
- Represent ML Ops in cross-functional planning with SRE, Cloud Platform, and consuming product teams.
Compensation
- Drives innovation and creativity within and between teams.
- Introduces employees to priorities outside of their immediate realm.
- Ensures important interpersonal relationships and connections with one another and our community!
Benefits
- Continuously monitor model performance, data drift, latency, error rates, and system health.
- Fosters a culture of collaboration, communication, performance and learning.
Company info
- your race, age, religion, sexual orientation, gender identity, ability, marital status, nationality, or any other protected characteristic won't affect your application.
- If you require any adjustments or accommodations due to a disability, or any other reason that may help you in your interview process, please let us know by emailing careers@mimecast.com.
- Due to certain obligations to our customers, an offer of employment will be subject to your successful completion of applicable background checks, conducted in accordance with local law.
- It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment.
- Our Mihra AI agent delivers 7x faster threat response for customers, and we're recognized as "Agents of Change" in Human Risk Management.
Apply directly at Mimecast →Create a free account for alerts like thisView Mimecast immigration profile
This listing is sourced directly from Mimecast's careers page and normalized into a canonical job model.