Tenex
Staff Site Reliability Engineer
Remote, USA · Staff+
Sponsorship not specifiedDetected 8 days ago
Distributed SystemsAWSGCPAzureCloud PlatformsKubernetesTerraformPrometheusGrafanaDatadogDevOpsSite Reliability EngineeringTemporalMachine LearningLLMsCybersecuritySIEMSOARSOC OperationsDetection EngineeringIncident ResponseSystems EngineeringLeadershipCommunication
About the role
- Company Overview TENEX is an AI-native, automation-first, built-for-scale Managed Detection and Response (MDR) provider.
- Seed round led by Andreessen Horowitz (a16z).
- We're a small but well-funded team that just raised a substantial round - joining now comes with limited risk and unlimited upside.
Responsibilities
- You will play a crucial role in designing resilient infrastructure, automating operational workflows, and shaping the future of our production environments while collaborating across engineering teams to drive technical excellence.
- Design, build, and maintain highly available, scalable, and secure infrastructure to support our AI-native cybersecurity platform.
- Develop internal tooling and automation to streamline deployment processes, incident response, and capacity planning.
- Lead incident response efforts, conduct post-mortems, and implement long-term solutions to prevent recurring reliability issues.
- Partner with sibling Engineering teams, Product, and Security teams to ensure reliability is baked into our development lifecycle from concept to production.
Requirements
- 10+ years of experience in SRE, DevOps, or Software/Systems Engineering, particularly in managing production systems at scale.
- Extensive experience with tools like Terraform, Pulumi, or similar technologies to manage complex infrastructure deployments.
- Hands-on experience with monitoring, logging, and tracing stacks (e.g., Prometheus, Grafana, ELK, Datadog) to drive data-informed reliability decisions.
Nice to have
- Domain Background: Prior work in cybersecurity, specifically regarding SIEM, EDR, or SOAR infrastructure.
- AI/ML Infrastructure: Experience supporting infrastructure for large-scale AI/ML workloads (e.g., GPU scheduling, LLM serving optimization).
- Startup Mentality: Background driving high-impact engineering initiatives in high-growth startups or enterprise SaaS.
- Strong familiarity with Agentic Workflows such as Agno, Temporal, etc..
- Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
- Relevant certifications (CKA/CKAD, AWS/GCP Professional Cloud Architect, etc.) are a plus.
- Opportunity to work with cutting-edge AI-driven cybersecurity technologies and Google SecOps solutions.
- A culture of growth and development, with opportunities to expand your knowledge in AI, cybersecurity, and emerging technologies.
Skills
- Solid understanding of microservices architecture, distributed databases, and event-driven systems.
- Clear, concise communication skills and a bias for collaborative problem-solving.
- Strong problem-solving, debugging, and analytical skills, especially in high-pressure environments.
Compensation
- Competitive salary and benefits package.
Company info
- We are a force multiplier for defenders, helping organizations enhance their cybersecurity posture through advanced threat detection, rapid response, and continuous protection.
- Our team is composed of industry experts with deep experience in cybersecurity, automation, and AI-driven solutions.
- Backed by leading investors, we are rapidly growing and seeking top talent to join our mission of revolutionizing the AI-Native MDR landscape.
- As an early employee, you'll play a meaningful role in defining and building our culture.
- Culture is one of the most important things at TENEX.AI http://TENEX.AI-explore our culture deck at culture.tenex.ai http://culture.tenex.ai to witness how we embody it, prioritizing the irreplaceable collaboration and community of in-person work.
This listing is sourced directly from Tenex's careers page and normalized into a canonical job model.