Ciroos
Senior Forward Deployed Engineer
Remote (USA) · Senior
Sponsorship not specifiedDetected 14 days ago
PythonJavaGoBashNode.jsDistributed SystemsGitDatabricksAWSGCPAzureCloud PlatformsKubernetesTerraformAnsibleCI/CDJenkinsPrometheusDatadogDevOpsSite Reliability EngineeringMachine LearningLLMsAgentic AI
About the role
- If you thrive in fast-moving environments where deep engineering meets real-world customer impact, this role awaits you!
- This is a customer-facing engineering role where you directly integrate Ciroos into complex, heterogeneous enterprise environments.
- Your work materially influences our roadmap, accelerates customer adoption, and ensures successful production outcomes.
Responsibilities
- Plan, design, build, and maintain highly scalable, reliable, and efficient infrastructure for our AI SRE Teammate.
- Conduct post-incident reviews to identify root causes and implement preventative measures.
- Lead enterprise migrations onto Ciroos (e.g., replacing incumbent alerting/AIOps platforms such as BigPanda, Moogsoft, PagerDuty, Elastic, or Splunk), including replicating hundreds to thousands of existing correlation rules, running phased cutovers, and hitting hard customer go-live dates.
- Design, build, and tune alert normalization and correlation policies-writing conditions and operators (IN / CONTAINS / REGEX), extracting fields from noisy sources (email, Jenkins, CI/CD), and choosing tailored playbooks over catch-all rules.
- Validate and continuously improve investigation quality-diagnosing query-conversion and enrichment errors, unblocking failed investigations, and tuning RCA accuracy so the AI SRE Teammate earns customer trust.
- Own customer-facing communications and delivery artifacts-weekly executive status updates, SLA documents, escalation paths, and migration/scope-change tracking-and coordinate daily syncs with customer project leads.
- Build long-term technical relationships with senior engineering leaders at customer organizations.
- Own customer implementations end-to-end-from technical discovery and solution design through production cutover, UAT, go-live, hypercare, and post-launch stabilization-driving execution independently across customer, product, and engineering threads, and demonstrating MTTR reduction and elimination of manual toil.
- Translate messy, ambiguous customer requirements into clear technical designs, milestones, trade-offs, and acceptance criteria, and keep execution truth legible by making blockers, risks, dependencies, and scope changes explicit early with clear owners and next steps.
- Design and ship AI-powered investigation and remediation workflows with clear tool boundaries, guardrails, deterministic fallbacks, and human-in-the-loop controls-and define quality controls (versioning, grounding, evals, feedback loops) so the AI SRE Teammate stays trustworthy in production.
Requirements
- Hands-on experience designing and delivering production-grade AI/automation workflows, ideally LLM-powered, with practical guardrails and evaluation approaches.
- Comfort with enterprise security and architecture requirements-RBAC, encryption, auditability, private networking, identity, and data-handling expectations.
- Experience with operating/supporting customers running AI infrastructure or AI tools.
- Customer obsessed: Deep empathy for customers in operational roles such as yours (SREs, ITOps, and DevOps) with a passion to reduce their toil and innovate on their behalf, and always willing to go the extra mile to delight customers.
- At least 6 years of experience in Site Reliability Engineering, DevOps, or a similar operational role, including technical leadership or end-to-end ownership of customer-facing delivery. Prior experience in a technical customer success or forward-deployed/solutions engineering role is a strong plus.
- Intimately familiar with Generative AI, ML concepts and technologies. You must be actively using AI tools to drive your own personal productivity.
- Strong hands-on experience with observability tools (e.g., Dynatrace, Datadog, Prometheus), ITSM/ticketing systems (e.g., ServiceNow), and incident management/response systems-including integrating and mapping data between them. Experience migrating a customer off an incumbent AIOps, alerting, or incident-management platform (e.g., BigPanda, Moogsoft, PagerDuty, Elastic, Splunk) is a strong plus.
- Excellent problem-solving and analytical skills to diagnose complex systems systematically, including comfort writing and debugging alert-correlation logic, normalization rules, and regular expressions against messy real-world alert data.
- Track record of independently owning enterprise implementations end-to-end-not just contributing tasks within a larger delivery team-including managing timelines and deliverables, driving daily/weekly customer syncs, and producing executive-facing status updates and SLA documentation, with strong customer outcomes and minimal execution drift.
- Strong integration and data experience: APIs, webhooks, event-driven systems, schema alignment, transformations, retries/idempotency, reliable error handling, and end-to-end integration configuration (field mapping, bi-directional sync, SSO/SAML)-including, where required, taking apps through marketplace certification (e.g., ServiceNow Store, Microsoft Teams).
Nice to have
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- At least 6 years of experience in Site Reliability Engineering, DevOps, or a similar operational role, including technical leadership or end-to-end ownership of customer-facing delivery.
- Prior experience in a technical customer success or forward-deployed/solutions engineering role is a strong plus.
- Strong proficiency in at least one programming language (e.g., Python, Go, Java).
- Extensive experience with cloud platforms (AWS, GCP, or Azure).
- Solid understanding of Kubernetes, infrastructure as code tools (e.g., Terraform, CloudFormation, Ansible), and CI/CD pipelines.
- Intimately familiar with Generative AI, ML concepts and technologies.
- You must be actively using AI tools to drive your own personal productivity.
Benefits
- Comprehensive medical, vision, and dental benefits.
- 401k plans and commuter benefits.
- Equity: Equity that could change your life, not just look nice on paper.
Company info
- Implement and optimize our AI SRE Teammate to meet the needs of our customers in production and pre-production environments.
- Proactively monitor customer deployments (with Ciroos!) to ensure that our customers get the best out of the product.
- Recommend best practices to customers for implementing the Ciroos AI SRE Teammate in their environment.
- We're looking for an experienced and curious Senior Forward Deployed Engineer (FDE) to join our team.
- You will serve as the technical vanguard for our largest and most sophisticated customers - deploying Ciroos in enterprise ecosystems, solving real-world production engineering bottlenecks, and translating operational pain into product leverage.
Equal opportunity
- Collaborative coworkers with high IQ and high EQ.
- No politics.
- No bureaucracy.
- No permission-seeking
This listing is sourced directly from Ciroos's careers page and normalized into a canonical job model.