Treeswift
Senior Site Reliability and Infrastructure Engineer
New York Office · Senior · Full-time
Sponsorship not specified$160k-$220kDetected 40 days ago
MongoDBAWSCloud PlatformsKubernetesTerraformCI/CDLinuxDevOpsSite Reliability EngineeringMachine LearningAirflowMLOpsRoboticsSensorsFirewallLeadershipCommunication
About the role
- Our data pipeline, machine learning training platform, and web app could all benefit from further productionization.
- Help us scale and harden the platform that schedules our pipelines, runs machine learning training, and hosts our web app.
- We run Apache Airflow on Astronomer with DAGs that orchestrate high-volume processing across AWS and Kubernetes, including machine learning inference inside pipeline tasks.
Responsibilities
- Partner with the data platform and engineering teams to understand how changes propagate across pipeline execution (Astronomer-hosted Airflow DAGs), containerized workers (Kubernetes), and AWS services (S3, SQS, Lambda, Step Functions, ECS).
- Own CI/CD guardrails for production changes: build/deploy validation and safe rollout mechanics for Astronomer deployments (image builds pushed to ECR, and Airflow configuration updates via Astronomer CLI variable updates)
- That said, you'll still help lead reliability improvements and operational readiness-so the team has faster diagnosis, better alerts, and safer releases when issues do occur.
Requirements
- You are an experienced software engineer where the last 7-10 years required significant time on observability, systems/infrastructure engineering, SRE, or DevOps (ideally in a cloud environment).
- Ability to reason about architecture end-to-end and articulate your thoughts with product impact in mind (data movement, execution, failure handling, and operational visibility).
- Hands-on experience with infrastructure-as-code (Terraform and similar) and using it to deliver reliable environments.
- Experience with container orchestration and debugging in practice (Kubernetes and/or ECS/container-based deployments).
- Strong Linux debugging skills and demonstrated ability to investigate production issues with logs/metrics and clear hypotheses.
- Experience working in early-stage or fast-moving environments where ownership and processes evolve quickly.
- Experience with Apache Airflow and/or Astronomer.
- Experience with AWS, although other cloud providers are fine. (DuploCloud experience is also helpful.)
- Experience with geospatial/imagery/lidar/point-cloud style domains.
Compensation
- The estimated salary range for this position is $160,000 - 220,000 USD.
- Total compensation for this position is determined by skills, qualifications, relevant work experience, location, and other factors.
- This salary estimate excludes the value of any potential bonuses; the value of any benefits offered; and the potential future value of any long-term incentives.
- This information is provided per the New York City Human Rights Law.
- Please note that the range provided is applicable only to New York City-based applicants.
- Base compensation may vary if the work location is outside of New York City.
Benefits
- dashboards and SLO/SLI definitions focused on correctness, throughput, and pipeline health
- Make machine learning inference operations more reliable and observable:
- Create operational tooling and continuously improve systems ('leave it better than you found it'), including:
Company info
- We take pride in managing complexity and providing high-fidelity data that our customers can use to make better-informed decisions.
- There is not currently an established on-call rotation for this platform, and the pipelines do not require real-time processing.
Equal opportunity
- Treeswift is proud to be an equal opportunity employer.
Apply directly at Treeswift →Create a free account for alerts like thisView Treeswift immigration profile
This listing is sourced directly from Treeswift's careers page and normalized into a canonical job model.