BIG Viking Games 3
DevOps Engineer, Cloud Infrastructure & Live Games
Toronto, Ontario, Canada
Sponsorship not specified$95k-$115kDetected 7 days ago
GitPostgreSQLMySQLRedisSnowflakeAWSCloud PlatformsKubernetesTerraformCI/CDGitHub ActionsJenkinsCircleCILinuxPrometheusGrafanaDatadogDevOpsSite Reliability EngineeringMachine LearningData EngineeringCybersecurityIncident ResponseAccessibility
About the role
- This is a hands-on role for someone who understands cloud infrastructure, automation, CI/CD, containers, monitoring, uptime, and production reliability - and who is comfortable working with legacy production systems alongside modern infrastructure patterns.
- Our games have been running for over a decade; the infrastructure reflects that history, and the right person sees that as an interesting challenge rather than a dealbreaker.
- The right person is practical, security-aware, automation-minded, and able to balance speed with reliability.
Responsibilities
- Monitor, maintain, and improve cloud infrastructure across AWS, Netlify, Vercel, and related platforms that support our live games, data systems, internal tools, and AI-powered operational workflows.
- Drive infrastructure modernization while maintaining uptime for live games with active player communities - every improvement ships while the plane is flying.
- Build, maintain, and improve automation for deployments, environment management, provisioning, secrets rotation, and operational workflows - reducing manual toil and human error.
- Implement and maintain Infrastructure as Code using tools such as Terraform, CloudFormation, CDK, or similar technologies.
- Own secrets and credential lifecycle management across platforms - including API key rotation, access controls, environment variable governance, and least-privilege practices.
- Support and improve the infrastructure that powers AI and automation tooling, including API integrations, MCP servers, serverless functions, webhook reliability, and orchestration platforms.
- Support incident response, root cause analysis, remediation planning, and post-incident improvements.
- Help manage cloud spend, infrastructure usage, resource tagging, and environment efficiency.
- Create clear documentation, runbooks, SOPs, and repeatable processes for infrastructure and DevOps workflows.
- You will work closely with engineering, product, QA, data, and live operations teams to improve how we build, deploy, monitor, and operate our systems.
Requirements
- 5+ years of experience in DevOps, infrastructure engineering, cloud engineering, site reliability engineering, or a similar role.
- Strong hands-on experience with AWS or similar cloud platforms.
- Experience designing, maintaining, and improving production infrastructure - including comfort with legacy systems that predate modern cloud-native patterns.
- Proficiency with Infrastructure as Code tools such as Terraform, CloudFormation, CDK, Pulumi, or similar.
- Experience with containerized applications, especially Docker.
- Experience with CI/CD tools, version control, deployment automation, and modern release workflows.
- Strong understanding of Linux systems, networking, cloud security, monitoring, logging, and operational troubleshooting.
- Experience supporting production systems where uptime, reliability, and performance matter - especially systems that cannot tolerate extended downtime.
- Experience with relational databases (MariaDB, MySQL, Postgres) and comfort working adjacent to data pipelines and ETL processes.
- Security-aware mindset with practical experience in secrets management, credential rotation, access control, vulnerability reduction, and least-privilege practices.
Nice to have
- Experience in gaming, live-service products, SaaS, digital products, or other high-availability consumer platforms.
- Experience supporting live games, virtual worlds, multiplayer systems, or real-time online products.
- Experience with GitHub Actions, GitLab CI, Jenkins, CircleCI, Buildkite, or similar CI/CD tools.
- Experience with Datadog, Grafana, Prometheus, CloudWatch, ELK, OpenTelemetry, or similar observability tools.
- Experience with Redis, Memcached, queues, workers, or event-driven systems.
- Experience with Snowflake, data warehouse connectivity, ETL monitoring, or data pipeline reliability.
- Experience with serverless platforms (Netlify Functions, Vercel, AWS Lambda) and multi-platform hosting environments.
- Experience with disaster recovery, backup strategies, incident management, load testing, and performance tuning.
Compensation
- $95,000 to $115,000 determined based on experience, technical depth, infrastructure ownership breadth, and overall fit.
Equal opportunity
- Accessibility and Accommodation Big Viking Games is committed to creating an inclusive and accessible environment for all candidates.
- If you require accommodation during the hiring process, please contact hr@bigvikinggames.com so we can work with you to support your needs.
Apply directly at BIG Viking Games 3 →Create a free account for alerts like thisView BIG Viking Games 3 immigration profile
This listing is sourced directly from BIG Viking Games 3's careers page and normalized into a canonical job model.