Rbc
Site Reliability Engineer (SRE), Cloud Operations
TORONTO, Ontario, Canada · Full-time
Sponsorship not specifiedDetected 4 hours ago
PythonExpressAWSGCPAzureCloud PlatformsKubernetesTerraformAnsibleCI/CDPrometheusGrafanaDevOpsSite Reliability EngineeringPlatform EngineeringKafkaMachine Learning
> stay_score
odds of building a lasting career here
16Unrated
Cap-exempt (no lottery)0
Sponsors this role0
Entry-level history0
PERM / green-card track0
Lottery odds40
Fits your clock70
No strong sponsorship signal in the public record yet. In the full product we resolve the exact legal entity and show its filing history with a confidence score — treat as unverified until then.
Lottery odds assume a STEM candidate.
Personalize to your clock →> community_outcomes
No reports yet — be the first to help the next applicant.
About the role
- Job Description What is the opportunity?
- This role offers the chance to shape how the bank operates, monitors, and self-heals its private and public cloud platforms - from OpenShift clusters and Kafka environments to self-healing automation systems.
- If you want to move beyond traditional ops into the future of intelligent, autonomous infrastructure operations - this is the role.
Responsibilities
- reducing toil for NOC, Data Center, and Branch teams, building automation that eliminates manual work, and establishing reliable operational practices.
- Support highly scalable, secure, and highly available architectures across private and public cloud platforms (Kubernetes/OpenShift, ECE, Confluent Kafka).
- Participate in and lead design reviews for new platform features, infrastructure changes, and operational integration points, ensuring alignment with security, reliability, and regulatory requirements.
- Drive automation, CI/CD, and Infrastructure as Code practices across the team, leveraging Ansible and Terraform for deployment validation and self-healing remediation workflows.
- Participate in on-call rotation for platform support, incident management, and troubleshooting, triaging incidents via Grafana, Prometheus, Dynatrace, and PagerDuty.
- We thrive on the challenge to be our best, progressive thinking to keep growing, and working together to deliver trusted advice to help our clients thrive and communities prosper.
Requirements
- Hands-on experience with Ansible and Terraform for Infrastructure as Code and automation.
- Proficiency in Python scripting (core to infrastructure automation and platform development).
- Hands-on experience with monitoring and observability stacks (Prometheus, Grafana, ELK, or equivalent).
- Experience with incident management processes, on-call rotations, and post-incident review practices.
- Experience with AI/ML concepts applied to operations (anomaly detection, intelligent alerting, predictive capacity planning).
- Nice-to-have Experience with AI/ML concepts applied to operations (anomaly detection, intelligent alerting, predictive capacity planning).
- Experience with GPU/compute infrastructure for ML inference workloads What's in it for you?
Compensation
- A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicable Leaders who support your development through coaching and managing opportunities Ability to mak
This listing is sourced directly from Rbc's careers page and normalized into a canonical job model.