Altruist

Altruist

Senior Software Engineer, Cloud Infrastructure

San Francisco, CA · Senior

Sponsorship not specified$200k-$250kDetected 54 days ago
PythonBashNode.jsCode ReviewGitPostgreSQLRedisElasticsearchDynamoDBVector DatabasesAWSCloud PlatformsKubernetesTerraformHelmCI/CDGitHub ActionsJenkinsLinuxPrometheusGrafanaDatadogDevOpsPlatform Engineering

About the role

  • We're hiring a Senior Cloud Infrastructure Engineer to join the Cloud Infrastructure & Platform (CIN) team at Altruist.
  • This is not a ticket-driven infrastructure role.
  • At the Senior-to-Staff level, we expect you to:

Responsibilities

  • Own and evolve the Infrastructure as Code (IaC) strategy using Terraform - define module standards, enforce code review practices, and drive adoption of reusable patterns across teams.
  • Lead Kubernetes (EKS) platform strategy, including cluster upgrades, node group architecture, Helm chart governance, service mesh evolution, and workload autoscaling policies.
  • Design and drive CI/CD platform improvements (GitHub Actions, ArgoCD, or similar) to enable safe, fast, and self-service deployments for application engineering teams.
  • Architect and validate disaster recovery (DR) strategies, including cross-region failover designs, backup automation, and leading DR simulation exercises.
  • Lead infrastructure design reviews and architectural discussions; ensure solutions meet scalability, security, and compliance requirements before implementation.
  • Define and drive the observability strategy across the platform (Datadog, Prometheus, Grafana, CloudWatch, OpenSearch) - including SLO/SLI frameworks, alerting standards, and distributed tracing.
  • Serve as a senior on-call escalation point; lead root-cause analysis on critical production incidents and drive systemic improvements through blameless post-mortems.
  • Own monthly resource saturation reviews and capacity planning processes; proactively identify scaling needs and present findings to engineering leadership.
  • Drive cloud cost optimization strategy: FinOps practices, Reserved Instances/Savings Plans analysis, vendor spend governance, and accountability frameworks across teams.
  • Define and enforce security architecture standards across AWS environments: IAM policy governance, VPC design patterns, encryption strategies, secrets management (Vault, AWS Secrets Manager), and vulnerability remediation workflows.

Requirements

  • 5+ years of hands-on experience in cloud infrastructure engineering, with deep, production-proven expertise in AWS.
  • Expert-level proficiency with Terraform, including module design, state management strategies, and establishing IaC standards for engineering teams.
  • Strong Linux systems engineering skills and advanced scripting proficiency (Python, Bash, or Go).
  • Deep expertise in CI/CD platforms (GitHub Actions, ArgoCD, Jenkins) with experience designing deployment strategies for multi-service architectures.
  • Proven experience designing and operating observability platforms (Datadog, Prometheus/Grafana, CloudWatch, OpenSearch/ELK) at organizational scale.
  • Strong experience with database infrastructure and data layer architecture (Aurora PostgreSQL, RDS, ElastiCache/Redis, OpenSearch, DynamoDB).
  • Excellent technical communication skills. ability to write clear ADRs, present to leadership, and translate infrastructure complexity for non-technical stakeholders.
  • 7+ years of infrastructure or platform engineering experience, including 3+ years operating at a senior or staff level.
  • Experience in financial services, fintech, broker-dealer, or other heavily regulated industries with FINRA/SEC compliance requirements.
  • Experience with event streaming platforms at scale (Amazon MSK / Apache Kafka) including cluster operations, partition strategy, and consumer group management.
  • Track record of owning and driving infrastructure initiatives end-to-end - from design and architecture through implementation, rollout, and operational excellence.
  • Expert-level understanding of cloud networking (VPC architecture, Transit Gateway, peering, DNS, load balancing) and security (IAM, KMS, WAF, GuardDuty, Secrets Manager).
  • Demonstrated ability to lead disaster recovery planning, execute DR simulations, and design HA architecture patterns for mission-critical systems.
  • Proven track record of mentoring engineers and elevating team capabilities through knowledge sharing, design reviews, and tooling improvements.
  • Bonus points

Nice to have

  • Extensive production experience operating Kubernetes (EKS strongly preferred) at scale - cluster lifecycle management, multi-tenancy patterns, Helm governance, and GitOps workflows.

Skills

  • Cloud Infrastructure Architecture & Platform Engineering
  • Reliability, Observability & Operational Excellence

Compensation

  • San Francisco, CA salary range

Company info

  • We're looking for exceptional talent to help us achieve our mission of making financial advice better, more affordable, and accessible to all.
  • But first, our values
  • Grit - When challenges arise, we stay laser focused on achieving our mission and finding a way forward, even when it's hard.

This listing is sourced directly from Altruist's careers page and normalized into a canonical job model.