SHEIN
Senior Site Reliability Engineer
San Diego · Senior
Sponsorship not specified$92k-$149kDetected 97 days ago
PythonGoDistributed SystemsGitRedisElasticsearchKubernetesCI/CDLinuxNginxPrometheusGrafanaSite Reliability EngineeringKafkaLLMsIncident ResponseNetwork Monitoring
About the role
- We are seeking a Senior Site Reliability Engineer (Official Title: Senior Site Reliability Engineer I) with deep experience operating and evolving large-scale, mission-critical systems where availability and reliability are non-negotiable.
- At SHEIN, Site Reliability Engineers are hybrid software and systems engineers responsible for keeping production services always on while enabling the platform to scale rapidly and safely.
- We are accountable for driving platform operability forward by reducing incident frequency, minimizing MTTR, and improving system resilience, efficiency, and resource utilization.
Responsibilities
- In this role, you will own and support complex services and infrastructure, ensuring they consistently meet reliability and performance expectations.
- You will work closely with global, cross-functional teams to design, build, and evolve observability and operational tooling-including metrics, logs, traces, alerting, and automation-providing deep visibility into system behavior.
- Triage and resolve production incidents, leveraging AI-assisted log analysis and anomaly detection to accelerate root cause identification; drive continuous improvements that reduce MTTR and prevent recurrence.
- Monitor and manage capacity planning and resource utilization, partnering with cross-functional teams to ensure systems scale safely while remaining cost-effective.
- Own and operate core open-source infrastructure such as APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, Zookeeper and other large-scale distributed systems.
- Develop and maintain technical documentation, including runbooks, architecture diagrams, operational procedures, and on-call playbooks.
- Work closely with global engineering teams to improve infrastructure reliability and performance through better system design and operational discipline.
Requirements
- Experience with observability and monitoring systems (Prometheus, Grafana, Zabbix, etc.) and performance analysis.
- Familiarity with Git, CI/CD pipelines, and configu
Nice to have
- Keep SHEIN's mission-critical production systems running 24/7/365, participating in on-call rotations and acting decisively during incidents.
- Triage and resolve production incidents, leveraging AI-assisted log analysis and anomaly detection to accelerate root cause identification
- Bachelor's degree in Computer Science, Information Systems, or a related technical discipline, or equivalent practical experience.
- 3+ years of experience owning and operating large-scale, high-traffic, 24/7 production systems, ideally in cloud or cloud-native environments.
- Solid foundations in Linux, networking, and distributed systems, with the ability to debug complex production issues end to end.
- Hands-on experience with incident response, troubleshooting, and performance optimization in distributed systems.
Compensation
- $92k-$149k
Company info
- Your work will directly contribute to delivering a stable, scalable, and high-performing experience for customers worldwide.
This listing is sourced directly from SHEIN's careers page and normalized into a canonical job model.