iCapital

iCapital

Site Reliability Engineer - Vice President

Salt Lake City, Utah, United States · Exec · Full-time

Sponsorship not specified$130k-$160kDetected 13 days ago
Distributed SystemsPostgreSQLMongoDBDynamoDBAWSKubernetesTerraformPrometheusGrafanaSite Reliability EngineeringIncident ResponseCommunication

About the role

  • As a Site Reliability Engineer, you'll work at the intersection of software engineering and operations, applying engineering principles to infrastructure challenges.

Responsibilities

  • Define, implement, and iterate service level objectives (SLOs) and service level indicators (SLIs) that reflect customer and business expectations.
  • Develop observability standards across metrics, logs, and traces, including instrumentation and dependency mapping patterns (OpenTelemetry where applicable).
  • Lead technical evaluations and PoCs for observability platforms and integrations
  • Drive automation to eliminate toil, improve repeatability, and accelerate recovery (incident workflows, runbooks, and remediation where appropriate).
  • Serve as Incident Commander for high-severity incidents, lead postmortems, and drive systemic improvements through action items and measurable follow-through using established tooling workflows.
  • Lead technical evaluations and PoCs for observability platforms and integrations; define success criteria and migration approach for adoption.

Requirements

  • 7+ years in SRE or related roles, with evidence of technical seniority across multiple services and teams.
  • Strong experience with AWS and container orchestration (Kubernetes) in production environments.
  • Familiarity with common data stores and managed services (e.g., Postgres, MongoDB, DynamoDB) and how they fail in distributed systems.
  • Experience with at least two observability stacks (Prometheus/Grafana, New Relic, Splunk, CloudWatch, ELK, etc.) and driving standardization across them.
  • Clear written and verbal communication skills with the ability to influence engineering teams through standards, tooling, and practical guidance.
  • Demonstrated experience defining SLOs/SLIs and using them to drive operational and engineering decisions.
  • Proven ability to design and implement observability solutions that produce actionable insights while reducing alert fatigue and operational noise.
  • Strong IaC skills (Terraform preferred) and the ability to build reusable automation and standards (monitoring as code, configuration patterns).
  • Strong incident response skills, including leading retrospectives/postmortems and improving reliability through systematic follow-up.
  • Strong debugging skills across distributed systems and production environments, including performance and reliability investigations.

Compensation

  • The base salary range for this role is $130,000 to $160,000 depending on level. iCapital offers a compensation package which includes salary, equity for all full-time employees, and an annual performance bonus.

Benefits

  • The base salary range for this role is $130,000 to $160,000 depending on level. iCapital offers a compensation package which includes salary, equity for all full-time employees, and an annual performance bonus.

Company info

  • https://www.linkedin.com/company/icapital-network-inc |
  • Define and implement reliability and operability standards for Kubernetes-based services, including scaling patterns, resource constraints, rollout safety, and baseline dashboards and alerts as part of service onboarding.
  • We believe the best ideas and innovation happen when we are together.

Equal opportunity

  • https://www.icapitalnetwork.com/about-us/recognition/
  • iCapital is proud to be an Equal Employment Opportunity and Affirmative Action employer.
  • We do not discriminate based upon race, religion, color, national origin, gender, sexual orientation, gender identity, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics.

This listing is sourced directly from iCapital's careers page and normalized into a canonical job model.