xAI
Sr. Software Engineer (Data Center Automation)
Memphis, TN · Senior
Sponsorship not specifiedDetected 36 days ago
PythonGoRustC++DockerKubernetesLinuxMachine LearningSystems EngineeringElectrical EngineeringCommunicationMentoringAdaptability
About the role
- Software Engineer to join our team in managing and enhancing reliability across a multi-data center environment.
- In an era where AI workloads demand near-zero downtime, this position plays a pivotal role in bridging software engineering principles with physical data center realities.
- The primary objective of this team is to mitigate downtime and minimize impact to end-users from both scheduled and unscheduled maintenance, as well as events affecting onsite data centers.
Responsibilities
- Design, develop, and deploy scalable code and services (primarily in Python and Rust, with flexibility for emerging languages) to automate reliability workflows, including monitoring, alerting, incident response, and infrastructure provisioning.
- Optimize Linux-based systems for performance, security, and reliability, including kernel tuning, container orchestration (e.g., Kubernetes or emerging alternatives), and scripting for automation.
- This attracts candidates from varied networking and systems backgrounds to drive forward-thinking solutions.
- Mentor junior team members and document processes to foster a culture of automation, knowledge sharing, and adaptability to new technologies.
- If you thrive in lightning-fast, distributed environments and are passionate about leveraging automation to drive efficiency, this is an opportunity to make a significant impact on our infrastructure's resilience and scalability.
Requirements
- Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a closely related technical field (or equivalent professional experience).
- 3+ years of hands-on experience in site reliability engineering (SRE), infrastructure engineering, DevOps, or systems engineering, preferably supporting large-scale, distributed, or production environments.
- Solid experience with Linux systems administration, performance tuning, kernel-level understanding, and scripting/automation in production environments.
- Practical knowledge of containerization and orchestration technologies, such as Docker and Kubernetes (or similar systems).
- 5+ years of experience in SRE or infrastructure roles, ideally in hyperscale, cloud, or AI / ML training infrastructure environments with multi-data center setups.
Nice to have
- experience with Rust or willingness to work in Rust is a plus, but strong coding fundamentals in at least one systems-level language (e.g., Python, Go, C++) are essential.
- Strong programming skills with proven production experience in Python (required for automation and tooling)
Skills
- Proficiency in Rust for systems programming and performance-critical components.
- Background in optimizing Linux-based systems for AI workloads, GPU clusters, or high-throughput compute environments.
- Prior work with bare-metal provisioning, data center interconnects, or hybrid/multi-site failover mechanisms.
- Mentoring experience, strong documentation skills, and a track record of fostering knowledge sharing and automation culture.
- Comfort with rapid technology adaptation in fast-evolving domains like AI infrastructure.
- This organization is for individuals who appreciate challenging themselves and thrive on curiosity.
- All employees are expected to be hands-on and to contribute directly to the company's mission.
- Work ethic and strong prioritization skills are important.
- All employees are expected to have strong communication skills.
- They should be able to concisely and accurately share knowledge with their teammates.
Company info
- We are seeking a highly skilled Sr.
This listing is sourced directly from xAI's careers page and normalized into a canonical job model.