Altruist
Senior Software Engineer, Cloud Infrastructure
San Francisco, CA · Senior
Sponsorship not specified$200k-$250kDetected 54 days ago
PythonBashNode.jsCode ReviewGitPostgreSQLRedisElasticsearchDynamoDBVector DatabasesAWSCloud PlatformsKubernetesTerraformHelmCI/CDGitHub ActionsJenkinsLinuxPrometheusGrafanaDatadogDevOpsPlatform Engineering
About the role
- We're hiring a Senior Cloud Infrastructure Engineer to join the Cloud Infrastructure & Platform (CIN) team at Altruist.
- This is not a ticket-driven infrastructure role.
- At the Senior-to-Staff level, we expect you to:
Responsibilities
- Own and evolve the Infrastructure as Code (IaC) strategy using Terraform - define module standards, enforce code review practices, and drive adoption of reusable patterns across teams.
- Lead Kubernetes (EKS) platform strategy, including cluster upgrades, node group architecture, Helm chart governance, service mesh evolution, and workload autoscaling policies.
- Design and drive CI/CD platform improvements (GitHub Actions, ArgoCD, or similar) to enable safe, fast, and self-service deployments for application engineering teams.
- Architect and validate disaster recovery (DR) strategies, including cross-region failover designs, backup automation, and leading DR simulation exercises.
- Lead infrastructure design reviews and architectural discussions; ensure solutions meet scalability, security, and compliance requirements before implementation.
- Define and drive the observability strategy across the platform (Datadog, Prometheus, Grafana, CloudWatch, OpenSearch) - including SLO/SLI frameworks, alerting standards, and distributed tracing.
- Serve as a senior on-call escalation point; lead root-cause analysis on critical production incidents and drive systemic improvements through blameless post-mortems.
- Own monthly resource saturation reviews and capacity planning processes; proactively identify scaling needs and present findings to engineering leadership.
- Drive cloud cost optimization strategy: FinOps practices, Reserved Instances/Savings Plans analysis, vendor spend governance, and accountability frameworks across teams.
- Define and enforce security architecture standards across AWS environments: IAM policy governance, VPC design patterns, encryption strategies, secrets management (Vault, AWS Secrets Manager), and vulnerability remediation workflows.
Requirements
- 5+ years of hands-on experience in cloud infrastructure engineering, with deep, production-proven expertise in AWS.
- Expert-level proficiency with Terraform, including module design, state management strategies, and establishing IaC standards for engineering teams.
- Strong Linux systems engineering skills and advanced scripting proficiency (Python, Bash, or Go).
- Deep expertise in CI/CD platforms (GitHub Actions, ArgoCD, Jenkins) with experience designing deployment strategies for multi-service architectures.
- Proven experience designing and operating observability platforms (Datadog, Prometheus/Grafana, CloudWatch, OpenSearch/ELK) at organizational scale.
- Strong experience with database infrastructure and data layer architecture (Aurora PostgreSQL, RDS, ElastiCache/Redis, OpenSearch, DynamoDB).
- Excellent technical communication skills. ability to write clear ADRs, present to leadership, and translate infrastructure complexity for non-technical stakeholders.
- 7+ years of infrastructure or platform engineering experience, including 3+ years operating at a senior or staff level.
- Experience in financial services, fintech, broker-dealer, or other heavily regulated industries with FINRA/SEC compliance requirements.
- Experience with event streaming platforms at scale (Amazon MSK / Apache Kafka) including cluster operations, partition strategy, and consumer group management.
- Track record of owning and driving infrastructure initiatives end-to-end - from design and architecture through implementation, rollout, and operational excellence.
- Expert-level understanding of cloud networking (VPC architecture, Transit Gateway, peering, DNS, load balancing) and security (IAM, KMS, WAF, GuardDuty, Secrets Manager).
- Demonstrated ability to lead disaster recovery planning, execute DR simulations, and design HA architecture patterns for mission-critical systems.
- Proven track record of mentoring engineers and elevating team capabilities through knowledge sharing, design reviews, and tooling improvements.
- Bonus points
Nice to have
- Extensive production experience operating Kubernetes (EKS strongly preferred) at scale - cluster lifecycle management, multi-tenancy patterns, Helm governance, and GitOps workflows.
Skills
- Cloud Infrastructure Architecture & Platform Engineering
- Reliability, Observability & Operational Excellence
Compensation
- San Francisco, CA salary range
Company info
- We're looking for exceptional talent to help us achieve our mission of making financial advice better, more affordable, and accessible to all.
- But first, our values
- Grit - When challenges arise, we stay laser focused on achieving our mission and finding a way forward, even when it's hard.
Apply directly at Altruist →Create a free account for alerts like thisView Altruist immigration profile
This listing is sourced directly from Altruist's careers page and normalized into a canonical job model.