bettermode
Senior Platform Systems Engineer
Toronto, Ontario · Senior · Full-time
Sponsorship not specified$160k-$180kDetected 20 days ago
TypeScriptGoRustDistributed SystemsPostgreSQLMongoDBAWSKubernetesTerraformHelmLinuxDevOpsPlatform EngineeringgRPCMachine LearningLLMsMLOpsCybersecurityIncident ResponseComplianceSystems EngineeringTCP/IPWebflowCollaboration
About the role
- Location type: Remote or Hybrid (3 days at the office in Downtown Toronto - Monday, Tuesday and Wednesday, for employees residing within 40km of the company headquarters)
- This is not a generic DevOps role, not a narrow tool-operator role, and not a vendor-certified specialist role.
- Our operating model includes a production on-call program: engineers participate in an every-other-week rotation for P0 incident response, post-incident learning, and production ownership.
Responsibilities
- Own Kubernetes platform patterns and Terraform/OpenTofu workflows that make environments reproducible, reviewable, secure, and recoverable, including promotion, drift control, and policy-aware infrastructure changes.
- Design AZ-aware and topology-aware improvements, starting with Aurora PostgreSQL routing/scalability and extending to other data-plane systems where traffic locality, availability, and cost matter.
- Build cost and workload observability that attributes AWS infrastructure spend, network transfer, CPU/memory usage, and cross-AZ patterns to services, workloads, and teams.
- Build production-grade platform components in Go, Rust, or TypeScript where appropriate, including Kubernetes controllers, Terraform plugins, telemetry collectors, bespoke proxies, and CLIs.
- Implement platform security and compliance controls aligned with SOC 2, OWASP, GDPR, IAM least privilege, secrets handling, encryption, network segmentation, auditability, and data protection.
- Support OLAP analytics infrastructure and the migration from Pinot to ClickHouse, with attention to ingestion topology, query performance, data correctness, cost, and operational safety.
- Design safe rollout, resilience, and DR patterns, including canaries, bypass modes, fast rollback, degraded-mode operation, backup/restore workflows, failover procedures, RTO/RPO trade-offs, and incident playbooks.
- Build an AZ-aware Aurora PostgreSQL routing/proxy component and topology-aware controls that reduce inter-AZ traffic, improve reader/writer behaviour, and provide predictable behaviour during scaling and failover.
- Build workload-level cost and resource intelligence capabilities across AWS and Kubernetes, including attribution for network traffic, cross-AZ patterns, CPU/GPU/memory utilization, and other signals needed to improve infrastructure efficiency.
Requirements
- Deep production experience with Kubernetes/EKS, Terraform/OpenTofu, AWS, and Cloudflare, including secure deployment patterns, environment promotion, drift control, and operational ownership.
- Have experience with service meshes, proxies, transport-aware systems, traffic steering, or network observability for cost-attribution.
- Have MLOps experience with KServe, Kubeflow, MLflow, model-serving infrastructure, GPU workloads, or other AI/ML platform systems.
Skills
- Eastern Standard Time
Compensation
- CA$160K-$180K/annually for Canada-based candidates
- AI Use: Large language models (LLM) might be used in the hiring process for this position to screen, assess or select job applicants.
- <!-- notionvc: 9a63c557-af0d-4431-b244-c0c5ce85d522 -->
Benefits
- 🩺 From your very first day, you and your family are covered by comprehensive Canadian health benefits-dental and vision included-so you can focus on what matters most.
- 😎 Enjoy unlimited paid vacation days, paid parental leave to support your family, and bereavement leave should you need it.
- 💡 We want you to thrive in your work: every team member receives a monthly Tech & Appreciation Stipend-perfect for testing new software or tools and improving your workflows as you see fit.
- The office features complimentary snacks, coffee, video games, and board games, as well as dedicated seating and a flexible environment that supports creativity, focus, and teamwork.
Company info
- At Bettermode, we are redefining how businesses streamline customer experiences and foster strong relationships. Our platform empowers businesses to seamlessly craft powerful web apps with engagement tools in its core tailored to their unique needs.
- Backed by Silicon Valley investors and trusted by brands like Monday.com, Webflow, and Nubank, we're proud to connect millions of end-users daily (check our Showcase page 😉).
- Join us as we continue building tools that redefine customer engagement!
- At Bettermode, we are redefining how businesses streamline customer experiences and foster strong relationships.
- Our platform empowers businesses to seamlessly craft powerful web apps with engagement tools in its core tailored to their unique needs.
- Remote or Hybrid (3 days at the office in Downtown Toronto - Monday, Tuesday and Wednesday, for employees residing within 40km of the company headquarters)
- At Bettermode, we're dedicated to empowering our team to thrive-both professionally and personally.
- Our culture is built on ownership and trust, giving you real influence over how we scale and succeed.
- We offer location-based, competitive compensation that reflects your expertise and impact, with annual reviews so you can grow with us.
Apply directly at bettermode →Create a free account for alerts like thisView bettermode immigration profile
This listing is sourced directly from bettermode's careers page and normalized into a canonical job model.