Tobogganlabs

Tobogganlabs

Ingénieur·e en fiabilité des sites - Site Reliability Engineer

Montréal, Quebec, Canada

Sponsorship not specifiedDetected 41 days ago
PythonBashGitAWSAzureCloud PlatformsKubernetesTerraformCI/CDGitHub ActionsJenkinsPrometheusGrafanaDatadogDevOpsSite Reliability EngineeringNetwork SecurityIncident ResponseLeadershipCommunicationAdaptability

About the role

  • Note that while we specialize in healthcare and regulated industries, not all our projects are in these fields, so you may work across different domains from time to time.
  • Automation and platform operations - Automate deployment pipelines, infrastructure provisioning, and operational runbooks to reduce toil and improve system resilience.
  • Supporting the team - Share SRE expertise with colleagues, contribute to internal tooling and documentation, mentor team members, and participate in the broader Toboggan community.

Responsibilities

  • We're seeking a Site Reliability Engineer (SRE) to help our clients build reliable, observable, and secure production systems.
  • Observability and reliability - Design and implement monitoring, alerting, and logging systems; lead incident response and post-mortem processes; define and track SLOs and SLIs.
  • On some projects, technical leadership - Own the reliability and infrastructure workstream, guide client engineering teams on SRE practices, and contribute to architectural decisions.

Requirements

  • Have 5+ years of experience in infrastructure, DevOps, or site reliability engineering;
  • Have hands-on experience with AWS or Azure infrastructure and infrastructure-as-code tools (Terraform, CloudFormation, or equivalents);
  • Have strong experience with CI/CD pipelines (GitHub Actions, ArgoCD, Jenkins, or equivalents) and deployment automation;
  • Have experience with observability tools (Prometheus, Grafana, Datadog, CloudWatch, or equivalents) and incident management processes;

Nice to have

  • Budget pour le bureau à domicile et la technologie;
  • Budget annuel de développement professionnel;
  • REER avec contribution de l'employeur après 1 an;
  • Assurance santé et dentaire payée à 100 % par l'employeur, incluant un montant annuel pour les soins complémentaires (acupuncture, ostéopathie, massothérapie, naturopathie, psychologie, etc.);
  • Assurance vie et assurance invalidité de courte et de longue durée;
  • Have experience in client-facing roles such as consulting, implementation engineering, or advisory work.
  • Have experience with container orchestration (Kubernetes, ECS) and cloud-native tooling.

Skills

  • Most of our clients run AWS, Terraform, GitHub Actions or similar CI/CD tooling.
  • Have excellent communication skills and can explain infrastructure and reliability concepts to varied stakeholders;
  • Have built infrastructure automation using scripting (Python, Bash) or workflow tools.
  • Hold relevant certifications (AWS DevOps Professional, AWS Solutions Architect, CKA, or similar).

Company info

  • Toboggan Labs is a boutique consultancy building at the intersection of AI and healthcare.
  • We solve challenging human problems by applying cutting-edge technology and domain understanding.
  • We prefer to hire in Quebec, but we are open to candi
  • We are a remote-first company with office space in Montreal.

This listing is sourced directly from Tobogganlabs's careers page and normalized into a canonical job model.