xAI

xAI

Sr. Software Engineer (Data Center Automation)

Memphis, TN · Senior

Sponsorship not specifiedDetected 36 days ago
PythonGoRustC++DockerKubernetesLinuxMachine LearningSystems EngineeringElectrical EngineeringCommunicationMentoringAdaptability

About the role

  • Software Engineer to join our team in managing and enhancing reliability across a multi-data center environment.
  • In an era where AI workloads demand near-zero downtime, this position plays a pivotal role in bridging software engineering principles with physical data center realities.
  • The primary objective of this team is to mitigate downtime and minimize impact to end-users from both scheduled and unscheduled maintenance, as well as events affecting onsite data centers.

Responsibilities

  • Design, develop, and deploy scalable code and services (primarily in Python and Rust, with flexibility for emerging languages) to automate reliability workflows, including monitoring, alerting, incident response, and infrastructure provisioning.
  • Optimize Linux-based systems for performance, security, and reliability, including kernel tuning, container orchestration (e.g., Kubernetes or emerging alternatives), and scripting for automation.
  • This attracts candidates from varied networking and systems backgrounds to drive forward-thinking solutions.
  • Mentor junior team members and document processes to foster a culture of automation, knowledge sharing, and adaptability to new technologies.
  • If you thrive in lightning-fast, distributed environments and are passionate about leveraging automation to drive efficiency, this is an opportunity to make a significant impact on our infrastructure's resilience and scalability.

Requirements

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a closely related technical field (or equivalent professional experience).
  • 3+ years of hands-on experience in site reliability engineering (SRE), infrastructure engineering, DevOps, or systems engineering, preferably supporting large-scale, distributed, or production environments.
  • Solid experience with Linux systems administration, performance tuning, kernel-level understanding, and scripting/automation in production environments.
  • Practical knowledge of containerization and orchestration technologies, such as Docker and Kubernetes (or similar systems).
  • 5+ years of experience in SRE or infrastructure roles, ideally in hyperscale, cloud, or AI / ML training infrastructure environments with multi-data center setups.

Nice to have

  • experience with Rust or willingness to work in Rust is a plus, but strong coding fundamentals in at least one systems-level language (e.g., Python, Go, C++) are essential.
  • Strong programming skills with proven production experience in Python (required for automation and tooling)

Skills

  • Proficiency in Rust for systems programming and performance-critical components.
  • Background in optimizing Linux-based systems for AI workloads, GPU clusters, or high-throughput compute environments.
  • Prior work with bare-metal provisioning, data center interconnects, or hybrid/multi-site failover mechanisms.
  • Mentoring experience, strong documentation skills, and a track record of fostering knowledge sharing and automation culture.
  • Comfort with rapid technology adaptation in fast-evolving domains like AI infrastructure.
  • This organization is for individuals who appreciate challenging themselves and thrive on curiosity.
  • All employees are expected to be hands-on and to contribute directly to the company's mission.
  • Work ethic and strong prioritization skills are important.
  • All employees are expected to have strong communication skills.
  • They should be able to concisely and accurately share knowledge with their teammates.

Company info

  • We are seeking a highly skilled Sr.

This listing is sourced directly from xAI's careers page and normalized into a canonical job model.