Element451

Element451

Senior Platform Engineer

Remote, Canada · Senior

Sponsorship not specifiedDetected 19 days ago
GitMongoDBSnowflakeAWSCloud PlatformsDockerTerraformCI/CDGitHub ActionsSite Reliability EngineeringPlatform EngineeringCybersecurityNetwork SecurityDetection EngineeringIncident ResponseComplianceCustomer SuccessFirewall

About the role

  • This is a hands-on senior IC role, and a deliberately broad one: reliability and operations are the core, with CI/CD and delivery, security, infrastructure, and data reliability all real, recurring parts of the work.
  • The work has range and a fair amount of unpredictability; the people who thrive in it tend to want exactly that.
  • Be the platform's hands-on security operator - IAM and least-privilege hygiene, secrets management, threat detection and response (WAF, GuardDuty), and vulnerability triage and remediation against SLA.

Responsibilities

  • This is a rare chance to own reliability, operations, and security at the level where it actually happens: keeping a real production system healthy, fast, and safe, and building the automation that keeps it that way.
  • You'll keep Element451's platform reliable, secure, and operable - and build the delivery systems that let it scale without scaling the firefighting.
  • You'll partner closely with our Director of Platform Engineering, who owns the platform strategy; you own a large share of the operational and delivery execution that brings it to life.
  • We hold a high bar - and we give you the ownership, context, and support to meet it.
  • Carry production operations day to day - participate in on-call, lead incident response, run blameless post-incident reviews, and drive issues to root cause rather than patching symptoms.
  • Treat operational toil as engineering work to eliminate - relentlessly automate remediation, sharpen alert quality, and drive down MTTD and MTTR rather than absorbing manual load.
  • Build the developer-facing automation and paved roads that cut friction for the product engineering team - treating them as the platform's customer and making the reliable, secure path the easy path.
  • This is the security function in practice today; you partner with the Director of Platform Engineering on strategy and standards, and you're trusted to set the operational bar where none exists yet.
  • Own the outcome. You're on the hook for the result, not the task - reliability, security, and operability end to end - and you're genuinely unsatisfied with "it mostly works."
  • Our process is rigorous and designed to be real signal - for you and for us.

Requirements

  • Operational experience with MongoDB Atlas or a comparable managed database platform - backup and recovery, performance tuning, and monitoring at scale.
  • Working knowledge of compliance operations - SOC 2 Type II and FERPA - and producing audit evidence (we use Vanta) as a byproduct of good engineering rather than a separate effort.
  • 7+ years across site reliability, operations, delivery, infrastructure, or platform engineering, with a track record of hands-on delivery - including real startup experience where you owned broad, shifting scope without a large team behind you.
  • A strong SRE foundation - SLI/SLO design, observability stack ownership, and incident response in a production environment with real customer impact.
  • Strong CI/CD and delivery-engineering chops - building and operating pipelines (GitHub Actions), Docker/ECR workflows, and ECS deployment automation, with progressive delivery and automated rollback - enough to co-own the delivery platform, not just consume it.
  • Security-operations depth you can own without a security team behind you - IAM governance and least-privilege, secrets management, network security and threat detection (WAF, GuardDuty), and vulnerability triage and remediation. You'll be the org's hands-on security practitioner, so the judgment to set an operational bar - not just follow one - matters here.
  • Deep, current AWS expertise - ECS/Fargate, Lambda, SQS/SNS, EventBridge, S3/CloudFront, VPC networking, IAM, and Secrets Manager - plus strong Terraform and infrastructure-as-code discipline across multi-environment systems.
  • Comfort operating as a high-output IC with broad domain ownership and a long execution horizon. Current familiarity with AI-assisted operations - intelligent alerting, anomaly detection, or AI-augmented incident response - is a plus.
  • Impactful not Immediate - We prioritize and invest in initiatives that will be most impactful.
  • Progress before Perfection - We are action-oriented people. We are empowered to make decisions and achieve our goals.
  • Learners before Masters - We are curious and humble people who strive to constantly improve.
  • Together not Alone - We rally behind each other and pitch in to support the greater whole.

Nice to have

  • Comfort operating as a high-output IC with broad domain ownership and a long execution horizon.
  • Current familiarity with AI-assisted operations - intelligent alerting, anomaly detection, or AI-augmented incident response - is a plus.

Compensation

  • Competitive pay and full benefits - a salary calibrated to the seniority of the role, plus comprehensive medical & dental coverage for you and your family.

Benefits

  • Competitive pay and full benefits - a salary calibrated to the seniority of the role, plus comprehensive medical & dental coverage for you and your family.
  • Time to recharge - flexible PTO, paid company holidays, tenure milestones that reward your commitment, and your birthday off.

Company info

  • Our values describe how we work, and we look for people who already operate this way:
  • Element451 is building the AI-powered platform reshaping how colleges and universities recruit, enroll, and support their students - and the reliability of that platform is what earns the trust of the institutions that run on it.

This listing is sourced directly from Element451's careers page and normalized into a canonical job model.