Zillow
Senior Manager, Site Reliability Engineering, Follow Up Boss
Remote-USA · Senior
Sponsorship not specified$207k-$330kDetected 27 days ago
GitRedisAWSCloud PlatformsKubernetesTerraformAnsibleCI/CDSite Reliability EngineeringAgentic AICybersecurityIncident ResponseComplianceCRMCustomer SupportLeadershipCommunicationMentoring
About the role
- Move the infra team into a strong posture for proactive, roadmap-driven investments in reliability, scalability, security, and developer productivity.
- FUB application infrastructure across dev, QA, and production AWS accounts
- Observability, monitoring, and cost management for FUB workloads
Responsibilities
- Own execution for the FUB infra & security roadmap, turning strategic goals (e.g., DB scalability, ZGCP adoption, infra cost and reliability targets) into a sequenced, realistic plan with clear milestones and measures of success.
- Ensure the team hits commitments with rare surprises, and when risk emerges, proactively engage partners to adjust scope, resources, or timeline with clear communication and tradeoffs.
- These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation.
- Own team execution, quality, and innovation across multiple systems and workloads that support the entire FUB+ org.
- This is an M4 scope role: you will be expected to consistently deliver through others (including senior ICs who do not report to you), shape technical strategy beyond your immediate team, and operate with limited oversight.
- Partner with principal/staff engineers to refine FUB's service scaling strategy, ensuring clear guidance on when teams build in the monolith vs. new services, and how infra supports these choices.
- Drive faster, safer deployments by improving CI/CD (GitLab, pipelines, AMI replacements, canary/progressive delivery) and aligning with ZG best practices for trunk-based development and feature flags.
- Partner with product SDMs and tech leads to lower operational friction for dev teams (e.g., better runbooks, improved observability, easier infra integrations, automated guardrails and guardrails-powered AI tooling).
Requirements
- Proven track record as an Senior Engineering Manager or equivalent leading SRE, platform, or infrastructure teams supporting high-availability SaaS products.
- Experience scaling production systems and databases in a cloud environment (ideally AWS) and leading meaningful improvements to reliability, performance, and cost.
- Demonstrated ability to shift a team from reactive to proactive roadmap-driven execution, including setting strategy, defining metrics, and driving sustained progress across multiple quarters.
- Experience partnering with security, database, networking, and central platform teams in a multi-org environment
Nice to have
- Comfortable experimenting with and operationalizing AI tools in engineering workflows
- SaaS / Sales CRM experience is a plus
- Strong experience with scaling large LAMP / web applications.
Skills
- audits, compliance, app sec, tooling, and policy in partnership with ZG security teams
Compensation
- In California, Connecticut, Maryland, Massachusetts, New Jersey, New York, Washington state, and Washington DC the standard base pay range for this role is $206,700.00 - $330,300.00 annually.
Company info
- The FUB Infrastructure (i12e) owns the core platform that powers
- Follow Up Boss, including:
- Legacy Partner Development infrastructure and critical shared services
- Incident handling and communication, including on-call, response, and follow-through
- Drive urgent, sustained progress on database scaling and performance, including capacity management, query and schema optimization, and modernization of data infrastructure.
- Lead the FUB modernization strategy and execution for prioritized workloads (e.g., workers, supporting services), balancing devex wins, reliability, and risk while coordinating with central teams.
- Raise the bar on developer environments and onboarding, reducing friction from dev boxes, tooling setup, and infra access; ensure new engineers can be productive quickly with reliable, self-service workflows.
- Lead and grow a high-performing, inclusive SRE/infrastructure/security team, set clear expectations, provide candid feedback, and manage performance.
- Develop technical leaders within and adjacent to the team (SREs, SDEs, security engineers, P5 ICs) through sponsorship, delegation, and stretch opportunities that expand impact beyond the immediate team.
- Hire, retain, and onboard talent across SRE, infra SDE, ensuring skills match the breadth of FUB infra (AWS, Terraform/Ansible, Kubernetes/ZGCP, observability, security, databases).
- Be the primary technical and operational interface for FUB infra with FUB+ leadership and central Zillow platform orgs, driving alignment on priorities, tradeoffs, and architectural decisions.
- Contribute materially to FUB+ tech vision and infra strategy, especially around service scaling, platform adoption, and our long-term operations model (e.g., SRE ownership boundaries, infra/security shared services, cost posture).
- Help identify and resolve cross-org misalignment (e.g., ownership boundaries, duplicated infra work, conflicting platform choices) and advocate for solutions that maximize Zillow-wide value, not just local optimization.
- Champion innovation that improves reliability, scalability, cost, and devex for multiple teams, including adoption of ZG-standard tooling and patterns and infra-focused AI agents for automation, diagnostics, and operations.
- Normalize AI usage within the infra team (e.g., code generation, runbook drafting, incident summarization, capacity modeling) and share successful patterns more broadly across FUB+ and platform partners.
- Partner with security (ZG and FUB) to ensure infra and application environments meet audit, SOC2, SOX, privacy, and app-sec requirements, with clear ownership for remediation work and sustainable controls.
- Forecast and manage runtime and infra costs (compute, storage, observability, networking), using tagging, dashboards, and guardrails to keep costs within budget while supporting growth.
- Strong background in developer experience and CI/CD, with hands-on familiarity with tools such as Terraform/Ansible, GitLab, Kubernetes/ZGCP, and modern observability stacks.
- Experience partnering with security, database, networking, and central platform teams in a multi-org environment; able to navigate ambiguity and complex stakeholder landscapes.
Visa & Work Authorization
- Develop technical leaders within and adjacent to the team (SREs, SDEs, security engineers, P5 ICs) through sponsorship, delegation, and stretch opportunities that expand impact beyond the immediate team.
This listing is sourced directly from Zillow's careers page and normalized into a canonical job model.