Treeswift

Treeswift

Senior Site Reliability and Infrastructure Engineer

New York Office · Senior · Full-time

Sponsorship not specified$160k-$220kDetected 40 days ago
MongoDBAWSCloud PlatformsKubernetesTerraformCI/CDLinuxDevOpsSite Reliability EngineeringMachine LearningAirflowMLOpsRoboticsSensorsFirewallLeadershipCommunication

About the role

  • Our data pipeline, machine learning training platform, and web app could all benefit from further productionization.
  • Help us scale and harden the platform that schedules our pipelines, runs machine learning training, and hosts our web app.
  • We run Apache Airflow on Astronomer with DAGs that orchestrate high-volume processing across AWS and Kubernetes, including machine learning inference inside pipeline tasks.

Responsibilities

  • Partner with the data platform and engineering teams to understand how changes propagate across pipeline execution (Astronomer-hosted Airflow DAGs), containerized workers (Kubernetes), and AWS services (S3, SQS, Lambda, Step Functions, ECS).
  • Own CI/CD guardrails for production changes: build/deploy validation and safe rollout mechanics for Astronomer deployments (image builds pushed to ECR, and Airflow configuration updates via Astronomer CLI variable updates)
  • That said, you'll still help lead reliability improvements and operational readiness-so the team has faster diagnosis, better alerts, and safer releases when issues do occur.

Requirements

  • You are an experienced software engineer where the last 7-10 years required significant time on observability, systems/infrastructure engineering, SRE, or DevOps (ideally in a cloud environment).
  • Ability to reason about architecture end-to-end and articulate your thoughts with product impact in mind (data movement, execution, failure handling, and operational visibility).
  • Hands-on experience with infrastructure-as-code (Terraform and similar) and using it to deliver reliable environments.
  • Experience with container orchestration and debugging in practice (Kubernetes and/or ECS/container-based deployments).
  • Strong Linux debugging skills and demonstrated ability to investigate production issues with logs/metrics and clear hypotheses.
  • Experience working in early-stage or fast-moving environments where ownership and processes evolve quickly.
  • Experience with Apache Airflow and/or Astronomer.
  • Experience with AWS, although other cloud providers are fine. (DuploCloud experience is also helpful.)
  • Experience with geospatial/imagery/lidar/point-cloud style domains.

Compensation

  • The estimated salary range for this position is $160,000 - 220,000 USD.
  • Total compensation for this position is determined by skills, qualifications, relevant work experience, location, and other factors.
  • This salary estimate excludes the value of any potential bonuses; the value of any benefits offered; and the potential future value of any long-term incentives.
  • This information is provided per the New York City Human Rights Law.
  • Please note that the range provided is applicable only to New York City-based applicants.
  • Base compensation may vary if the work location is outside of New York City.

Benefits

  • dashboards and SLO/SLI definitions focused on correctness, throughput, and pipeline health
  • Make machine learning inference operations more reliable and observable:
  • Create operational tooling and continuously improve systems ('leave it better than you found it'), including:

Company info

  • We take pride in managing complexity and providing high-fidelity data that our customers can use to make better-informed decisions.
  • There is not currently an established on-call rotation for this platform, and the pipelines do not require real-time processing.

Equal opportunity

  • Treeswift is proud to be an equal opportunity employer.

This listing is sourced directly from Treeswift's careers page and normalized into a canonical job model.