CrowdStrike
Sr. AI Infrastructure Engineer, LLM/AI Platforms (Remote)
USA - Remote · Senior
Sponsorship not specified$140k-$215kDetected 5 days ago
PythonDistributed SystemsCode ReviewAWSGCPCloud PlatformsDockerKubernetesTerraformMachine LearningPyTorchSparkAirflowData EngineeringLLMsRAGAgentic AIMLOpsA/B TestingResearchLeadershipMentoring
About the role
- CrowdStrike is looking for a Senior AI Infrastructure Engineer with expertise in Large Language Models (LLMs) Infrastructure and data platforms to join our growing AI Infrastructure Team.
- You will champion engineering best practices, write high-quality code, and actively mentor and strengthen the team's technical knowledge and capabilities.
- CrowdStrike is a computer security company, but we do not require candidates for this role to have prior security industry experience.
Responsibilities
- Develop and optimize LLM model-serving infrastructure, including deployment and optimization of various inference frameworks.
- Lead model lifecycle management including versioning, checkpointing and reproducibility across training and inference deployments.
- Design and champion robust evaluation frameworks to assess model performance, accuracy, and reliability, ensuring AI systems are consistently at production-ready standards.
- Architect and maintain data platforms and pipelines specifically designed to support LLMs, Retrieval-Augmented Generation (RAG), and AI Agentic Systems at scale.
- Deliver production-ready code with a focus on performance, maintainability, and testing rigor, ensuring the ability to ship fast without compromising quality.
- Document architectural designs thoroughly and communicate technical decisions clearly to stakeholders
- Collaborate across the organization with Data Scientists, Product Managers, and other engineering teams to transform research prototypes into robust, production-grade services.
- Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections
Requirements
- 6+ years of experience in Infrastructure/Data Engineering, with at least 2 years focused on building and maintaining platforms/pipelines that support LLM-based systems and applications
- Demonstrable hands-on experience in LLM infrastructure engineering including cluster provisioning, optimizing training workloads, and maintaining inference pipelines
- Exceptional ability to write clean, elegant, performant, and well-tested code, coupled with a strong focus on action and delivering results quickly.
Nice to have
- Prior experience in the cybersecurity, intelligence, or high-compliance industries.
- Experience with common agentic workflow frameworks (e.g., LangChain, LlamaIndex).
- Experience with distributed data processing frameworks (e.g., Spark, Dask, Flink).
- Bachelor's degree in Computer Science, Data Engineering, or a related STEM field
- Master's degree preferred
Skills
- Hands-on experience with MLOps Tools (MLflow, Sagemaker, Vertex AI).
- Strong understanding of CUDA, NVIDIA drivers, GPU, and TPU compute fundamentals.
- Experience with inference serving frameworks such as vLLM and Triton Inference Server.
- Proficiency with distributed training frameworks including Pytorch, Ray, Megatron, and JAX.
- Expert-level proficiency in a high-level coding language (Python).
- Deep knowledge of containerization and orchestration (Docker, Kubernetes, Slurm, Airflow).
- Proficiency with Infrastructure as Code tooling like Terraform and Ansible.
- Experience with cloud platforms (AWS, GCP, or OCI) and related data services.
Compensation
- Market leader in compensation and equity awards
Benefits
- Market leader in compensation and equity awards
- Competitive vacation and holidays for recharge
- Paid parental and adoption leaves
Apply directly at CrowdStrike →Create a free account for alerts like thisView CrowdStrike immigration profile
This listing is sourced directly from CrowdStrike's careers page and normalized into a canonical job model.