Intactfc

Intactfc

SRE specialist

Montréal, Quebec, CAN

No sponsorship$110k-$134kDetected 30 days ago
TypeScriptPythonGoAWSGCPAzureKubernetesTerraformCI/CDSite Reliability EngineeringKafkaRabbitMQDNSBGP/OSPFLeadershipCommunicationMentoringPublic SpeakingArgo CD

About the role

  • We are seeking a hands-on Site Reliability Engineer within the Intelligent Operations Department's SRE & Resiliency team.
  • This role operates across Azure, AWS, GCP, and on‑prem environments, embedded in the broader enterprise resiliency and production reliability strategy.
  • Core responsibilities include deep investigations, advanced observability (OpenTelemetry, Dynatrace, Elastic), auto-healing tooling, SLI/SLO stewardship, and business-aligned reliability reporting.

Responsibilities

  • Lead high‑severity investigations and RCA with App/Infra/Incident teams.
  • Proactively find systemic risks and resilience gaps; drive durable fixes.
  • Implement end‑to‑end traces/metrics/logs with consistent semantics.
  • Build policy‑driven remediation (circuit breakers, throttling, retries).
  • Develop reliability copilots; monitor AI systems for reliability and cost.
  • The SRE will function as part of a special investigations unit that empowers and enables Applicative Support, Infrastructure Support, and the Incident Management team-coaching, guiding, and leading investigations into active incidents and proactive reliability improvements.

Requirements

  • 8+ years of experience in SRE/Platform/Infrastructure/Software Engineering operating large-scale production systems across multi-cloud and on‑prem.
  • No Canadian work experience required however must be eligible to work in Canada

Skills

  • Define user‑centric SLIs/SLOs; enforce error budget policies.
  • Publish reliability reports and scorecards; drive continuous improvement.
  • Upskill support/incident teams; standardize playbooks and training.
  • Promote automation‑first, data‑driven, resilience culture.
  • Operate across Azure/AWS/GCP/on‑prem; GLB, DNS, TLS, CDN, failover.
  • Improve K8s/mesh (AKS/EKS/GKE, Istio/Linkerd) and data/streaming resilience.
  • Use AI for causal detection/anomalies to cut MTTR.
  • Cloud & Platform Reliability

Compensation

  • Annual bonus target, based on the base salary, with a potential payout of up to double the target (subject to personal and company performance):
  • As part of our commitment to Win As A Team, we share our success with employees through our annual bonus plan and Employee Share Purchase Plan (ESPP) - with Intact matching 50% of your net shares.
  • Salary for the candidate will be determined taking into consideration a number of factors including: experience, skills, qualifications, anticipated contribution to role, internal equity, etc.
  • The salary range presented above is based on a 35-hour workweek and would represent a majority of different candidate profiles.
  • 109,900 - 134,300
  • Possibility to purchase up to 5 extra days off per year

Benefits

  • Build insights and anomaly detection; create topology‑aware health models.
  • Our pension offerings provide flexibility and long-term security for our employees beyond their careers.
  • We are one of the few companies offering the opportunity to receive guaranteed income for life via our defined benefit pension plan.
  • Flexible work arrangements and a hybrid work model

Visa & Work Authorization

  • Please note that Intact does not provide sponsorship or other support for immigration-related matters including but not limited to employer-specific closed work permits.

This listing is sourced directly from Intactfc's careers page and normalized into a canonical job model.