simplisafe

simplisafe

Staff Software Engineer, ML Infrastructure

Boston, MA · Staff+

Sponsorship not specified$147k-$215kDetected 6 days ago
PythonRustC++Distributed SystemsCode ReviewAWSKubernetesCI/CDKafkaMachine LearningComputer VisionLLMsMLOpsIncident ResponseLoad BalancingResearchLeadershipCommunicationCollaboration

About the role

  • This is a senior individual contributor role for a distributed systems expert who wants to apply that craft to one of the most demanding problem domains in the company.
  • The work spans two of our most demanding workloads: real-time computer vision inference that processes video from cameras and doorbells across our customer base, and LLM/GenAI infrastructure that will power our future generation of intelligent applications.
  • Both are, fundamentally, distributed systems problems - high-throughput, low-latency, multi-tenant, GPU-aware, and unforgiving of regressions.

Responsibilities

  • Drive architecture decisions for our Kubernetes-based ML platform - anchored on Ray for inference, alongside KServe, Triton, and vLLM - across real-time and batch workloads.
  • Lead deep technical reviews on system design, capacity planning, and reliability for the highest-stakes ML systems at SimpliSafe.
  • Build and operate real-time CV inference at scale
  • Own the design and evolution of cloud-side inference systems that process live video and events from SimpliSafe devices in real time.
  • Drive throughput, latency, and cost improvements (batching strategies, GPU utilization, autoscaling, multi-model serving) for production CV models.
  • Build the feedback loops between cloud inference, edge devices, and the data flywheel that improves model quality over time.
  • Partner with applied ML engineers to take new GenAI-powered product features from prototype to scaled deployment.
  • A mission- and values-driven culture and a safe, inclusive environment where you can build, grow and thrive
  • Employee Resource Groups (ERGs) that bring people together, give opportunities to network, mentor and develop, and advocate for change.

Requirements

  • Deep expertise in high-throughput, low-latency services - ad serving, recommendations, real-time APIs, online platforms, or similar - including the operational reality of running them at scale.
  • Strong production experience on Kubernetes and AWS (EKS, S3, IAM, networking) and with Kafka, containerized deployments, CI/CD, and infrastructure-as-code.
  • Demonstrated experience with the building blocks of high-scale systems: load balancing, autoscaling, batching, caching, multi-tenancy, queuing, and capacity planning.
  • Strong written and verbal communication - you can make complex technical tradeoffs legible to ML scientists, product, and other infra teams.
  • Hands-on experience with Ray, KServe, Triton, vLLM, or other ML serving stacks.
  • Hands-on experience with LLM serving in production (vLLM, TGI, TensorRT-LLM, SGLang) - KV cache management, continuous batching, speculative decoding, quantization for serving.
  • Experience operating GPU-based inference systems - GPU-aware scheduling, multi-model serving, accelerator utilization optimization.
  • Familiarity with ML fundamentals - how models are trained, evaluated, versioned, deployed, monitored, and rolled back in production.
  • Experience with model lifecycle tooling (MLflow, Weights & Biases, model registries, drift detection, shadow deployments).
  • Experience operating in environments with strong security and compliance requirements.
  • 8+ years of software engineering experience, with a clear track record of building and operating large-scale distributed systems in production.
  • Staff-level technical leadership: ability to drive ambiguous, cross-cutting initiatives, align senior stakeholders, and elevate the engineers around you without formal authority.
  • ML exposure is preferred - having deployed or operated production ML systems, worked closely with ML teams, or built ML-adjacent infrastructure. Exceptional distributed systems engineers without direct ML experience are encouraged to apply
  • Bonus Points
  • Experience building real-time video or streaming pipelines (Kafka, Kinesis, Flink, or similar) at scale.

Nice to have

  • Proficiency in Python is required
  • experience with a systems language (Go, C++, Rust) for performance-sensitive components is a plus.
  • Prior ML experience is a plus, not a prerequisite.
  • ML exposure is preferred - having deployed or operated production ML systems, worked closely with ML teams, or built ML-adjacent infrastructure.
  • we'll help you ramp.

Compensation

  • The target annual base pay range for this role is $146,600 to $215,100.
  • This target annual base pay range represents our good-faith estimate of what we expect to pay for this role.
  • We use a market-based compensation approach to set our target annual base pay ranges and make adjustments annually.

Benefits

  • A comprehensive total rewards package that supports your wellness and provides security for SimpliSafers and their families (For more information on our total rewards please click here )

Company info

  • About SimpliSafe
  • We're a high-tech home security company that's passionate about protecting the life you've built and our mission of keeping Every Home Secure.
  • And we've created a culture here that cares just as deeply about the career you're building.
  • Ours is a no ego culture of collaboration and innovation where those seeking their next challenge can find big opportunities and make a huge impact on the lives of all those who we protect.
  • We don't just want you to work here.
  • We want you to grow and thrive here.
  • We're embracing a hybrid work model that enables our teams to split their time between office and home.
  • Hybrid for us means we expect our teams to come together in our state-of-the-art office on two core days, typically Tuesday, Wednesday, or Thursday - working together in person and choosing where they work for the remainder of the week.
  • We all benefit from flexibility and get to use the best of both worlds to get our work done.
  • Why are we hiring?
  • Well, we're growing and thriving.
  • So, we need smart, talented, and humble people who share our values to join us as we disrupt the home security space and relentlessly pursue our mission of keeping Every Home Secure.
  • We regularly review our programs to ensure they remain competitive and aligned with our values.

This listing is sourced directly from simplisafe's careers page and normalized into a canonical job model.