Intactfc
SRE specialist
Montréal, Quebec, CAN
No sponsorship$110k-$134kDetected 30 days ago
TypeScriptPythonGoAWSGCPAzureKubernetesTerraformCI/CDSite Reliability EngineeringKafkaRabbitMQDNSBGP/OSPFLeadershipCommunicationMentoringPublic SpeakingArgo CD
About the role
- We are seeking a hands-on Site Reliability Engineer within the Intelligent Operations Department's SRE & Resiliency team.
- This role operates across Azure, AWS, GCP, and on‑prem environments, embedded in the broader enterprise resiliency and production reliability strategy.
- Core responsibilities include deep investigations, advanced observability (OpenTelemetry, Dynatrace, Elastic), auto-healing tooling, SLI/SLO stewardship, and business-aligned reliability reporting.
Responsibilities
- Lead high‑severity investigations and RCA with App/Infra/Incident teams.
- Proactively find systemic risks and resilience gaps; drive durable fixes.
- Implement end‑to‑end traces/metrics/logs with consistent semantics.
- Build policy‑driven remediation (circuit breakers, throttling, retries).
- Develop reliability copilots; monitor AI systems for reliability and cost.
- The SRE will function as part of a special investigations unit that empowers and enables Applicative Support, Infrastructure Support, and the Incident Management team-coaching, guiding, and leading investigations into active incidents and proactive reliability improvements.
Requirements
- 8+ years of experience in SRE/Platform/Infrastructure/Software Engineering operating large-scale production systems across multi-cloud and on‑prem.
- No Canadian work experience required however must be eligible to work in Canada
Skills
- Define user‑centric SLIs/SLOs; enforce error budget policies.
- Publish reliability reports and scorecards; drive continuous improvement.
- Upskill support/incident teams; standardize playbooks and training.
- Promote automation‑first, data‑driven, resilience culture.
- Operate across Azure/AWS/GCP/on‑prem; GLB, DNS, TLS, CDN, failover.
- Improve K8s/mesh (AKS/EKS/GKE, Istio/Linkerd) and data/streaming resilience.
- Use AI for causal detection/anomalies to cut MTTR.
- Cloud & Platform Reliability
Compensation
- Annual bonus target, based on the base salary, with a potential payout of up to double the target (subject to personal and company performance):
- As part of our commitment to Win As A Team, we share our success with employees through our annual bonus plan and Employee Share Purchase Plan (ESPP) - with Intact matching 50% of your net shares.
- Salary for the candidate will be determined taking into consideration a number of factors including: experience, skills, qualifications, anticipated contribution to role, internal equity, etc.
- The salary range presented above is based on a 35-hour workweek and would represent a majority of different candidate profiles.
- 109,900 - 134,300
- Possibility to purchase up to 5 extra days off per year
Benefits
- Build insights and anomaly detection; create topology‑aware health models.
- Our pension offerings provide flexibility and long-term security for our employees beyond their careers.
- We are one of the few companies offering the opportunity to receive guaranteed income for life via our defined benefit pension plan.
- Flexible work arrangements and a hybrid work model
Visa & Work Authorization
- Please note that Intact does not provide sponsorship or other support for immigration-related matters including but not limited to employer-specific closed work permits.
Apply directly at Intactfc →Create a free account for alerts like thisView Intactfc immigration profile
This listing is sourced directly from Intactfc's careers page and normalized into a canonical job model.