BIG Viking Games 3

BIG Viking Games 3

DevOps Engineer, Cloud Infrastructure & Live Games

Toronto, Ontario, Canada

Sponsorship not specified$95k-$115kDetected 7 days ago
GitPostgreSQLMySQLRedisSnowflakeAWSCloud PlatformsKubernetesTerraformCI/CDGitHub ActionsJenkinsCircleCILinuxPrometheusGrafanaDatadogDevOpsSite Reliability EngineeringMachine LearningData EngineeringCybersecurityIncident ResponseAccessibility

About the role

  • This is a hands-on role for someone who understands cloud infrastructure, automation, CI/CD, containers, monitoring, uptime, and production reliability - and who is comfortable working with legacy production systems alongside modern infrastructure patterns.
  • Our games have been running for over a decade; the infrastructure reflects that history, and the right person sees that as an interesting challenge rather than a dealbreaker.
  • The right person is practical, security-aware, automation-minded, and able to balance speed with reliability.

Responsibilities

  • Monitor, maintain, and improve cloud infrastructure across AWS, Netlify, Vercel, and related platforms that support our live games, data systems, internal tools, and AI-powered operational workflows.
  • Drive infrastructure modernization while maintaining uptime for live games with active player communities - every improvement ships while the plane is flying.
  • Build, maintain, and improve automation for deployments, environment management, provisioning, secrets rotation, and operational workflows - reducing manual toil and human error.
  • Implement and maintain Infrastructure as Code using tools such as Terraform, CloudFormation, CDK, or similar technologies.
  • Own secrets and credential lifecycle management across platforms - including API key rotation, access controls, environment variable governance, and least-privilege practices.
  • Support and improve the infrastructure that powers AI and automation tooling, including API integrations, MCP servers, serverless functions, webhook reliability, and orchestration platforms.
  • Support incident response, root cause analysis, remediation planning, and post-incident improvements.
  • Help manage cloud spend, infrastructure usage, resource tagging, and environment efficiency.
  • Create clear documentation, runbooks, SOPs, and repeatable processes for infrastructure and DevOps workflows.
  • You will work closely with engineering, product, QA, data, and live operations teams to improve how we build, deploy, monitor, and operate our systems.

Requirements

  • 5+ years of experience in DevOps, infrastructure engineering, cloud engineering, site reliability engineering, or a similar role.
  • Strong hands-on experience with AWS or similar cloud platforms.
  • Experience designing, maintaining, and improving production infrastructure - including comfort with legacy systems that predate modern cloud-native patterns.
  • Proficiency with Infrastructure as Code tools such as Terraform, CloudFormation, CDK, Pulumi, or similar.
  • Experience with containerized applications, especially Docker.
  • Experience with CI/CD tools, version control, deployment automation, and modern release workflows.
  • Strong understanding of Linux systems, networking, cloud security, monitoring, logging, and operational troubleshooting.
  • Experience supporting production systems where uptime, reliability, and performance matter - especially systems that cannot tolerate extended downtime.
  • Experience with relational databases (MariaDB, MySQL, Postgres) and comfort working adjacent to data pipelines and ETL processes.
  • Security-aware mindset with practical experience in secrets management, credential rotation, access control, vulnerability reduction, and least-privilege practices.

Nice to have

  • Experience in gaming, live-service products, SaaS, digital products, or other high-availability consumer platforms.
  • Experience supporting live games, virtual worlds, multiplayer systems, or real-time online products.
  • Experience with GitHub Actions, GitLab CI, Jenkins, CircleCI, Buildkite, or similar CI/CD tools.
  • Experience with Datadog, Grafana, Prometheus, CloudWatch, ELK, OpenTelemetry, or similar observability tools.
  • Experience with Redis, Memcached, queues, workers, or event-driven systems.
  • Experience with Snowflake, data warehouse connectivity, ETL monitoring, or data pipeline reliability.
  • Experience with serverless platforms (Netlify Functions, Vercel, AWS Lambda) and multi-platform hosting environments.
  • Experience with disaster recovery, backup strategies, incident management, load testing, and performance tuning.

Compensation

  • $95,000 to $115,000 determined based on experience, technical depth, infrastructure ownership breadth, and overall fit.

Equal opportunity

  • Accessibility and Accommodation Big Viking Games is committed to creating an inclusive and accessible environment for all candidates.
  • If you require accommodation during the hiring process, please contact hr@bigvikinggames.com so we can work with you to support your needs.

This listing is sourced directly from BIG Viking Games 3's careers page and normalized into a canonical job model.