SHEIN

SHEIN

Staff Site Reliability Engineer

San Diego · Staff+

Sponsorship not specified$108k-$180kDetected 96 days ago
PythonGoDistributed SystemsRedisElasticsearchKubernetesLinuxNginxSite Reliability EngineeringKafkaLLMsIncident ResponseLeadership

About the role

  • We are seeking a Staff Site Reliability Engineer (Official Title: Staff Site Reliability Engineer I) with deep experience operating and evolving large-scale, mission-critical systems where availability and reliability are non-negotiable.
  • At SHEIN, Site Reliability Engineers are hybrid software and systems engineers responsible for keeping production services always on while enabling the platform to scale rapidly and safely.
  • At the Staff level, you will also provide technical leadership, influencing platform architecture, reliability strategy, and operational standards across the organization.

Responsibilities

  • In this role, you will own and support complex services and infrastructure, ensuring they consistently meet reliability and performance expectations.
  • You will work closely with global, cross-functional teams to design, build, and evolve observability and operational tooling-including metrics, logs, traces, alerting, and automation-providing deep visibility into system behavior.
  • Triage and resolve production incidents, leveraging AI-assisted log analysis and anomaly detection to accelerate root cause identification; drive continuous improvements that reduce MTTR and prevent recurrence.
  • Monitor and manage capacity planning and resource utilization, partnering with cross-functional teams to ensure systems scale safely while remaining cost-effective.
  • Own and operate core open-source infrastructure such as APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, Zookeeper and other large-scale distributed systems.
  • Develop and maintain technical documentation, including runbooks, architecture diagrams, operational procedures, and on-call playbooks.
  • Work closely with global engineering teams to improve infrastructure reliability and performance through better system design and operational discipline.

Requirements

  • Hands-on experience with incident response, troubleshooting, and performance optimization in distributed systems.

Nice to have

  • Keep SHEIN's mission-critical production systems running 24/7/365, participating in on-call rotations and acting decisively during incidents.
  • Triage and resolve production incidents, leveraging AI-assisted log analysis and anomaly detection to accelerate root cause identification
  • Mentor Senior and mid-level SREs, raising the overall technical bar and operational maturity of the team.
  • Bachelor's degree in Computer Science, Information Systems, or a related technical discipline, or equivalent practical experience.
  • 6+ years of experience owning and operating large-scale, high-traffic, 24/7 production systems, ideally in cloud or cloud-native environments.
  • Solid foundations in Linux, networking, and distributed systems, with the ability to debug complex production issues end to end.

Compensation

  • $108k-$180k

Company info

  • We are accountable for driving platform operability forward by reducing incident frequency, minimizing MTTR, and improving system resilience, efficiency, and resource utilization.
  • Your work will directly contribute to delivering a stable, scalable, and high-performing experience for customers worldwide.

This listing is sourced directly from SHEIN's careers page and normalized into a canonical job model.