xAI
Data Engineer
Palo Alto, CA · Internship
Sponsorship not specified$150k-$210kDetected 2 days ago
PythonKubernetesMachine LearningData EngineeringData ScienceStatisticsForecastingResearchLeadershipCommunication
About the role
- High-quality data is fundamental to every stage of that mission.
- Our Data team is responsible for ensuring that the models are trained on the right data, in the right form, at the right quality, across every phase of the training lifecycle.
- We work at the intersection of data, infrastructure, and machine learning to ensure our models train effectively and reliably.
Responsibilities
- Investigate anomalous model behavior and rigorously identify the data issues that drive poor downstream performance
- Research, evaluate, and develop frontier methods for improving data quality and effectiveness in AI model development
- Partner across teams to identify where data needs exist and define the highest-impact opportunities for new data acquisition and improvement
- Build and maintain production-grade data pipelines, tooling, and software systems that ingest, process, validate, and deliver data for training
- Develop metrics, evaluation frameworks, and monitoring systems to assess how data quality influences model behavior at scale
- Create shared datasets, tooling, and internal data products that enable other teams to analyze, debug, and improve model performance
- You will work closely with acquisition teams, ML engineers, and software engineers to identify data needs, build scalable data pipelines, and continuously improve the quality of the data that shapes model behavior.
Requirements
- Bachelor's degree in computer science, data science, physics, mathematics, or a STEM discipline
- 1+ years of data/software engineering experience (internship experience is applicable)
- Experience in implementing or analyzing language models or neural networks
- Ability to operate effectively in a dynamic environment with evolving priorities, changing requirements, and fast-moving technical challenges
- Design, build, and improve the data cleaning, transformation, and quality-control steps required to produce high-quality training data
Skills
- Professional experience in analytics, data science, machine learning, or data engineering
- Strong experience with Python and the broader ecosystem of libraries and tools used in modern machine learning and data development
- Experience working with Parquet or similar columnar storage formats in large-scale data systems
- Familiarity with Kubernetes and distributed production environments
- Experience working with very large-scale datasets, including terabyte- to petabyte-scale data systems
- This organization is for individuals who appreciate challenging themselves and thrive on curiosity.
- All employees are expected to be hands-on and to contribute directly to the company's mission.
- Work ethic and strong prioritization skills are important.
- All employees are expected to have strong communication skills.
- They should be able to concisely and accurately share knowledge with their teammates.
Compensation
- $150,000 - $210,000 USD
Company info
- At SpaceXAI, we are building AI systems that push the frontier of human knowledge and scientific discovery.
Equal opportunity
- equal opportunity employer.
This listing is sourced directly from xAI's careers page and normalized into a canonical job model.