Ultimo Software Solutions Inc.

Ultimo Software Solutions Inc.

Senior Lead Site Reliability Engineer

Vancouver, British Columbia, Canada · Senior · Full-time

Sponsorship not specifiedDetected 12 days ago
PythonBashPowerShellNode.jsDistributed SystemsGitSQLAzureCloud PlatformsKubernetesTerraformAnsibleHelmCI/CDGitHub ActionsJenkinsPrometheusGrafanaDevOpsSite Reliability EngineeringCybersecurityIncident ResponseComplianceCollaboration

About the role

  • You will work in close partnership with R&D to embed operational excellence into the software delivery lifecycle, and you take full ownership of every system within Cloud Operations' purview.

Responsibilities

  • Design, implement, and continuously improve Azure-based infrastructure for high-availability, mission-critical SaaS services - owning the full lifecycle from architecture through to production operation.
  • Own, operate, and continuously improve CI/CD pipelines across Jenkins, Azure DevOps, and GitHub Actions - including pipeline architecture, build performance, deployment reliability, secrets handling, and migration work as we evolve our toolchain.
  • This is active ownership, not support.
  • Configure and maintain Ansible playbooks for configuration management, provisioning automation, and drift remediation across the infrastructure estate.
  • Build and maintain Infrastructure as Code using Terraform and/or ARM/Bicep, covering the full provisioning lifecycle - from initial environment build through to day-two operations and ongoing change management.
  • Work directly and continuously with R&D engineering teams to embed reliability, operability, and deployment quality into the software development lifecycle - including pipeline design reviews, pre-production environment ownership, release readiness, and incident learnings fed back into build practices.
  • Take full technical ownership of all systems within Cloud Operations' scope - infrastructure, tooling, pipelines, observability, and security controls. If it lives in our environment, you own its reliability, its documentation, and its improvement roadmap.
  • Lead root cause analysis on production incidents; author post-mortems with actionable engineering remediation, not just process changes.
  • Define, instrument, and own SLOs, SLIs, and error budgets for Azure-hosted SaaS services; use data to drive reliability investment decisions.
  • Evaluate emerging Azure services and features against real production requirements; build proof-of-concepts, validate at scale, and drive adoption where the engineering case is clear.

Requirements

  • You bring deep Azure and DevOps expertise, thrive in complex distributed environments, and raise the technical bar through the quality of your engineering work.
  • Contribute directly to the technical controls, evidence collection, and continuous compliance posture required to maintain SOC 2 Type II, ISO 27001, and ISO 9001 certification across the Cloud Operations environment.
  • Working experience with Ansible for configuration management and infrastructure automation.
  • Hands-on experience with Prometheus and Grafana in a production context - metric instrumentation, alerting rule design, and dashboard development, not just consumption.
  • Experience with Pingdom for synthetic monitoring and PagerDuty for incident alerting and on-call management - including configuration of escalation policies, alert routing, and participation in a 24/7 on-call rotation.
  • Proven, production-grade experience with Infrastructure as Code using Terraform and/or ARM/Bicep.

Benefits

  • Own cluster health, upgrade lifecycle, and capacity planning end-to-end.
  • Instrument Kubernetes workloads with Prometheus exporters and build Grafana dashboards that give engineering teams genuine operational visibility into service health, latency, error rates, and resource consumption.

This listing is sourced directly from Ultimo Software Solutions Inc.'s careers page and normalized into a canonical job model.