SHEIN

SHEIN

Senior Site Reliability Engineer

San Diego · Senior

Sponsorship not specified$92k-$149kDetected 97 days ago
PythonGoDistributed SystemsGitRedisElasticsearchKubernetesCI/CDLinuxNginxPrometheusGrafanaSite Reliability EngineeringKafkaLLMsIncident ResponseNetwork Monitoring

About the role

  • We are seeking a Senior Site Reliability Engineer (Official Title: Senior Site Reliability Engineer I) with deep experience operating and evolving large-scale, mission-critical systems where availability and reliability are non-negotiable.
  • At SHEIN, Site Reliability Engineers are hybrid software and systems engineers responsible for keeping production services always on while enabling the platform to scale rapidly and safely.
  • We are accountable for driving platform operability forward by reducing incident frequency, minimizing MTTR, and improving system resilience, efficiency, and resource utilization.

Responsibilities

  • In this role, you will own and support complex services and infrastructure, ensuring they consistently meet reliability and performance expectations.
  • You will work closely with global, cross-functional teams to design, build, and evolve observability and operational tooling-including metrics, logs, traces, alerting, and automation-providing deep visibility into system behavior.
  • Triage and resolve production incidents, leveraging AI-assisted log analysis and anomaly detection to accelerate root cause identification; drive continuous improvements that reduce MTTR and prevent recurrence.
  • Monitor and manage capacity planning and resource utilization, partnering with cross-functional teams to ensure systems scale safely while remaining cost-effective.
  • Own and operate core open-source infrastructure such as APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, Zookeeper and other large-scale distributed systems.
  • Develop and maintain technical documentation, including runbooks, architecture diagrams, operational procedures, and on-call playbooks.
  • Work closely with global engineering teams to improve infrastructure reliability and performance through better system design and operational discipline.

Requirements

  • Experience with observability and monitoring systems (Prometheus, Grafana, Zabbix, etc.) and performance analysis.
  • Familiarity with Git, CI/CD pipelines, and configu

Nice to have

  • Keep SHEIN's mission-critical production systems running 24/7/365, participating in on-call rotations and acting decisively during incidents.
  • Triage and resolve production incidents, leveraging AI-assisted log analysis and anomaly detection to accelerate root cause identification
  • Bachelor's degree in Computer Science, Information Systems, or a related technical discipline, or equivalent practical experience.
  • 3+ years of experience owning and operating large-scale, high-traffic, 24/7 production systems, ideally in cloud or cloud-native environments.
  • Solid foundations in Linux, networking, and distributed systems, with the ability to debug complex production issues end to end.
  • Hands-on experience with incident response, troubleshooting, and performance optimization in distributed systems.

Compensation

  • $92k-$149k

Company info

  • Your work will directly contribute to delivering a stable, scalable, and high-performing experience for customers worldwide.

This listing is sourced directly from SHEIN's careers page and normalized into a canonical job model.