The Allen Institute for Artificial Intelligence

The Allen Institute for Artificial Intelligence

Senior Engineering Manager, AI Infrastructure

Seattle, WA · Senior

Sponsorship not specified$147k-$220kDetected 4 days ago
GoAWSGCPKubernetesLinuxAI OrchestrationAgileSystems EngineeringRoboticsResearchLeadership

About the role

  • Persons in these roles are expected to work from our offices in Seattle.
  • If you have questions about on-site work arrangements for this role, please ask your recruiter.
  • Our ideal candidate is a: - Systems Expert: You have a deep, hands-on understanding of the Linux kernel, container runtimes, and distributed systems.

Responsibilities

  • your mandate is to keep the platform fast, reliable, and well-utilized, and to deliver against the roadmap set with your PM counterpart.
  • You plan and deliver against near-term operational goals, keep reliability and researcher velocity high, and turn priorities set with leadership into shipped, dependable systems.
  • You will manage some of the most dense and high-performance compute environments currently in operation.
  • Manage and grow a team of systems engineers, SREs, and software developers.
  • Proficient in designing and managing SDLC processes including sprint planning and technical design reviews.
  • We develop foundational AI research and innovation to deliver real-world impact through large-scale open models, data, robotics, conservation, and beyond.
  • We value diversity - We seek to hire, support, and promote people from all genders, ethnicities, and all levels of experience regardless of age.

Requirements

  • Our ideal candidate is a:
  • Systems Expert: You have a deep, hands-on understanding of the Linux kernel, container runtimes, and distributed systems.
  • Pragmatic Operator: You are comfortable making trade-offs between technical elegance and operational necessity.
  • Bachelor's degree in a related field: a relevant advanced degree may substitute for equivalent years of technical work experience.
  • Storage: Hands-on experience with distributed filesystems (e.g., WEKA, Ceph, Lustre) and cloud storage integration at scale.
  • Team members will be required to follow any other job-related instructions and to perform any other job-related duties requested by any person authorized to give instructions or assignments.

Skills

  • Manage GPU compute allocation against budget.
  • Strong background in Kubernetes, Slurm, or similar orchestration frameworks, particularly in hybrid-cloud configurations.

Compensation

  • Team members will be able to receive annual bonuses and can participate in the long-term incentive plan.

Benefits

  • Team members and their families are covered by medical, dental, vision, and an employee assistance program.
  • Team members are able to enroll in our health savings account plan, our healthcare reimbursement arrangement plan, and our health care and dependent care flexible spending account plans.
  • Team members will receive $125 per month to assist with commuting or internet expenses and will also receive $200 per month for fitness and wellbeing expenses.
  • Team members will also receive up to ten sick days per year, up to seven personal days per year, up to 20 vacation days per year and twelve paid holidays throughout the calendar year.
  • Our base salary range is $146,880 - $220,320, and in addition we have generous bonus plans to provide a competitive compensation package.
  • Manage the availability, performance, and health of our dense on-prem GPU clusters.

Company info

  • Ai2 is a non-profit research institute at the forefront of open-source AI development. Unlike industry peers, our goal is to share our findings, data, code, and models with the global scientific community.
  • Why Ai2:
  • Open Science: Your work directly enables the release of open models like OLMo, providing the broader research community with tools they can't get elsewhere.
  • Mission-Driven: We prioritize scientific impact over profit margins. This allows us to focus on building the "right" infrastructure for long-term research goals.
  • Ai2 is a non-profit research institute at the forefront of open-source AI development.
  • Unlike industry peers, our goal is to share our findings, data, code, and models with the global scientific community.
  • Complexity at Scale: You will manage some of the most dense and high-performance compute environments currently in operation.
  • Your Next Challenge:
  • Cluster Operations: Manage the availability, performance, and health of our dense on-prem GPU clusters. Coordinate with hardware vendors and internal teams to keep physical infrastructure meeting the demands of frontier model training.
  • Orchestration & Scheduling: Operate and improve Beaker, our internal orchestration platform by optimizing resource allocation and driving high utilization across on-prem assets and elastic cloud resources (AWS/GCP).
  • Storage Operations: Execute and continuously improve our storage environment, balancing high-throughput performance for active training against cost-effective durability for petascale research data. Contribute to the longer-term storage roadmap.
  • Resource Management: Manage GPU compute allocation against budget. Track utilization, surface the data, and recommend when to burst to the cloud versus investing in on-prem capacity, escalating larger trade-offs as needed.
  • User Support & Velocity: Serve as the technical bridge to our research teams. Ensure infrastructure is an accelerator, not a bottleneck, for a diverse set of research objectives.
  • Team Leadership: Manage and grow a team of systems engineers, SREs, and software developers. Set the bar for operational rigor, engineering quality, and a collaborative culture, and keep the team unblocked and delivering.

Visa & Work Authorization

  • This employer participates in E-Verify and will provide the federal government with your Form I-9 information to confirm that you are authorized to work in the U.S. If E-Verify cannot confirm that you are authorized to work, this employer i

This listing is sourced directly from The Allen Institute for Artificial Intelligence's careers page and normalized into a canonical job model.