OutSystems
Senior Site Reliability Engineer
US - San Francisco Bay Area · Senior
Sponsorship not specifiedDetected 49 days ago
PythonGoBashReactDistributed SystemsGitAWSKubernetesTerraformLinuxPrometheusGrafanaSite Reliability EngineeringLLMsAgentic AIIncident ResponseNegotiationCustomer SuccessCommunicationCollaborationProblem SolvingCritical Thinking
About the role
- Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software engineering and applies them to infrastructure and operations problems.
- Our SREs ensure our production systems' reliability, performance, and scalability while enabling rapid development and deployment of new features and services.
- Qualifications and Skills To illustrate the desired profile for a Site Reliability Engineer.
Responsibilities
- Lead and onboard services and teams to the reliability tenets;
- Establish and maintain Service Level Objectives (SLOs) and Service Level Agreements (SLAs);
- Design and implement scalable, reliable, and secure infrastructure, while ensuring cloud-native best practices;
- Collaborate with software development teams to ensure systems are resilient (observable, fault-tolerant, recoverable, scalable) and performant;
- Implement monitoring, alerting, logging, and tracing solutions to detect and respond to incidents;
- Lead incident response efforts, ensuring quick resolution and minimal downtime, and conduct RCA/post-mortems;
- Participate in on-call rotation to provide 24/7 support for production systems.
Requirements
- 6+ years of experience in Site Reliability Engineering, managing infrastructure and services at scale
- Experience managing Hadoop and Kubernetes infrastructure and related services, or equivalent experience
- Advanced knowledge of Linux, Networking, and Containers
- Proficiency in at least one high-level programming language (Python, GoLang etc.).
- An understanding or hands-on experience with Prompt engineering in software development;
- Familiarity with AI Native IDEs or AI Assistants such as Cursor, GitHub CoPilot, and Claude.
- Nevertheless, the selection of candidates will always vary depending on specific knowledge of the field and prior experience.
- Ability to establish, monitor, and improve Service Level Objectives (SLOs), Indicators (SLIs), and Agreements (SLAs) in line with business needs.
- BS/MS in Computer Science or Equivalent
- History of end-to-end project delivery
- Strong troubleshooting and debugging skills.
- Fluency in English and excellent communication skills.
- Soft Skills
- Communication - able to communicate effectively (in English) both orally and written showing empathy for the other person
- Collaboration - Proactive collaboration and presentation skills to effectively communicate ideas and represent the deliverables and needs of the SRE team with leadership.
- Humbleness - accepts mistakes and acts accordingly, with a humble attitude, apologizing for them and mitigating them ASAP to avoid higher impact.
Skills
- Experience in any of the following is valued, but not fully required:
- Containerization technologies and orchestration platforms, mainly Kubernetes and EKS
- (CKA, CKAD, CKS certifications are valued)
- Experience with automation and Infrastructure as Code (IaC) tools, such as AWS CloudFormation, Terraform, Puppet, Chef, Spacelift, etc
- Experience with Python, Go, Bash/Shell scripting, or other automation tools/languages
- Familiarity with AWS services like EC2, RDS, ELB, CloudFront, Lambda, etc
- Proficiency in monitoring and troubleshooting complex distributed systems
- Experience with Grafana, ELK stack, Prometheus, or others
- Strong understanding of designing resilient and fault-tolerant systems
- Expertise in debugging complex distributed systems.
- More about OutSystems
Company info
- A company at the vanguard of the agentic revolution, where we don't just react to AI innovation-we architect it. Joining OutSystems means stepping onto a high-growth rocket ship that combines the fearless agility of a startup with the sophisticated, global foundation of an enterprise powerhouse.
- Real growth opportunities. We don't just talk about development; we invest in it through structured programs designed to scale your expertise. Whether you are aiming for vertical progression, exploring lateral moves into new domains, or mastering specialized AI skills through our Professional Development Fund and Internal Mobility Program, we provide the resources to get you there.
- A global collective of world-class talent, where you'll collaborate with enterprise software legends and sought-after thought leaders. At OutSystems, our industry experts aren't just visionaries-they are accessible, approachable mentors who are deeply invested in your growth as we architect the agentic future together.
Apply directly at OutSystems →Create a free account for alerts like thisView OutSystems immigration profile
This listing is sourced directly from OutSystems's careers page and normalized into a canonical job model.