Eqbank
Senior AI Platform Operations Engineer
Toronto · Senior
Sponsorship not specifiedDetected 13 days ago
PythonPowerShellGitAzureCloud PlatformsTerraformCI/CDGitHub ActionsDevOpsSite Reliability EngineeringMachine LearningIncident ResponseForecastingCollaboration
About the role
- Purpose of Job The Senior AI Platform Operations Engineer is accountable for the reliability, operability, and controlled enablement of the organization's AI platform. This role ensures that AI Platform services and solutions are production-ready, secure, observable, and compliant by executing disciplined operational practices, across platform management,
- monitoring, incident coordination, and governance control enforcement. The incumbent plays a key role in enabling the safe and scalable adoption of AI by ensuring that AI solutions are deployed, monitored, supported, and continuously improved in line with enterprise standards for reliability, security, and compliance
Responsibilities
- Lead operational triage, escalation coordination, and post-incident reviews to strengthen services stability and resilience.
Requirements
- Implement and maintain observability capabilities, including telemetry, logging, metrics, and traces required for enterprise AI operations.
- Maintain documentation and evidence required for audit, governance reviews, production readiness checkpoints, and control validation.
- Maintain operational visibility of AI platform assets required for monitoring, support, and cost alignment.
- University degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.
- 5-7 years of experience in platform operations, site reliability engineering, DevOps, cloud operations, or enterprise IT operations.
- Strong experience supporting production platforms and services, including monitoring, incident response, problem management, service restoration, and operational reporting.
- Expertise with observability tools such as Azure Monitor, Application Insights, and Grafana.
- Experience with CI/CD and automation tools such as Azure DevOps, GitHub Actions, and Logic Apps.
- Knowledge of configuration management and infrastructure-as-code tools such as Bicep, Terraform, Azure Policy, Key Vault, and relevant open-source technologies.
- Knowledge of integration and event-driven technologies such as API Management, open-source API tools, Service Bus, Event Grid, and Apache Kafka.
Benefits
- Monitor platform health using dashboards, logs, metrics, and alerts, and coordinate incident and service restoration activities.
This listing is sourced directly from Eqbank's careers page and normalized into a canonical job model.