SHEIN
Staff Site Reliability Engineer
San Diego · Staff+
Sponsorship not specified$108k-$180kDetected 96 days ago
PythonGoDistributed SystemsRedisElasticsearchKubernetesLinuxNginxSite Reliability EngineeringKafkaLLMsIncident ResponseLeadership
About the role
- We are seeking a Staff Site Reliability Engineer (Official Title: Staff Site Reliability Engineer I) with deep experience operating and evolving large-scale, mission-critical systems where availability and reliability are non-negotiable.
- At SHEIN, Site Reliability Engineers are hybrid software and systems engineers responsible for keeping production services always on while enabling the platform to scale rapidly and safely.
- At the Staff level, you will also provide technical leadership, influencing platform architecture, reliability strategy, and operational standards across the organization.
Responsibilities
- In this role, you will own and support complex services and infrastructure, ensuring they consistently meet reliability and performance expectations.
- You will work closely with global, cross-functional teams to design, build, and evolve observability and operational tooling-including metrics, logs, traces, alerting, and automation-providing deep visibility into system behavior.
- Triage and resolve production incidents, leveraging AI-assisted log analysis and anomaly detection to accelerate root cause identification; drive continuous improvements that reduce MTTR and prevent recurrence.
- Monitor and manage capacity planning and resource utilization, partnering with cross-functional teams to ensure systems scale safely while remaining cost-effective.
- Own and operate core open-source infrastructure such as APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, Zookeeper and other large-scale distributed systems.
- Develop and maintain technical documentation, including runbooks, architecture diagrams, operational procedures, and on-call playbooks.
- Work closely with global engineering teams to improve infrastructure reliability and performance through better system design and operational discipline.
Requirements
- Hands-on experience with incident response, troubleshooting, and performance optimization in distributed systems.
Nice to have
- Keep SHEIN's mission-critical production systems running 24/7/365, participating in on-call rotations and acting decisively during incidents.
- Triage and resolve production incidents, leveraging AI-assisted log analysis and anomaly detection to accelerate root cause identification
- Mentor Senior and mid-level SREs, raising the overall technical bar and operational maturity of the team.
- Bachelor's degree in Computer Science, Information Systems, or a related technical discipline, or equivalent practical experience.
- 6+ years of experience owning and operating large-scale, high-traffic, 24/7 production systems, ideally in cloud or cloud-native environments.
- Solid foundations in Linux, networking, and distributed systems, with the ability to debug complex production issues end to end.
Compensation
- $108k-$180k
Company info
- We are accountable for driving platform operability forward by reducing incident frequency, minimizing MTTR, and improving system resilience, efficiency, and resource utilization.
- Your work will directly contribute to delivering a stable, scalable, and high-performing experience for customers worldwide.
This listing is sourced directly from SHEIN's careers page and normalized into a canonical job model.