CrowdStrike

CrowdStrike

Sr. AI Infrastructure Engineer, LLM/AI Platforms (Remote)

USA - Remote · Senior

Sponsorship not specified$140k-$215kDetected 5 days ago
PythonDistributed SystemsCode ReviewAWSGCPCloud PlatformsDockerKubernetesTerraformMachine LearningPyTorchSparkAirflowData EngineeringLLMsRAGAgentic AIMLOpsA/B TestingResearchLeadershipMentoring

About the role

  • CrowdStrike is looking for a Senior AI Infrastructure Engineer with expertise in Large Language Models (LLMs) Infrastructure and data platforms to join our growing AI Infrastructure Team.
  • You will champion engineering best practices, write high-quality code, and actively mentor and strengthen the team's technical knowledge and capabilities.
  • CrowdStrike is a computer security company, but we do not require candidates for this role to have prior security industry experience.

Responsibilities

  • Develop and optimize LLM model-serving infrastructure, including deployment and optimization of various inference frameworks.
  • Lead model lifecycle management including versioning, checkpointing and reproducibility across training and inference deployments.
  • Design and champion robust evaluation frameworks to assess model performance, accuracy, and reliability, ensuring AI systems are consistently at production-ready standards.
  • Architect and maintain data platforms and pipelines specifically designed to support LLMs, Retrieval-Augmented Generation (RAG), and AI Agentic Systems at scale.
  • Deliver production-ready code with a focus on performance, maintainability, and testing rigor, ensuring the ability to ship fast without compromising quality.
  • Document architectural designs thoroughly and communicate technical decisions clearly to stakeholders
  • Collaborate across the organization with Data Scientists, Product Managers, and other engineering teams to transform research prototypes into robust, production-grade services.
  • Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections

Requirements

  • 6+ years of experience in Infrastructure/Data Engineering, with at least 2 years focused on building and maintaining platforms/pipelines that support LLM-based systems and applications
  • Demonstrable hands-on experience in LLM infrastructure engineering including cluster provisioning, optimizing training workloads, and maintaining inference pipelines
  • Exceptional ability to write clean, elegant, performant, and well-tested code, coupled with a strong focus on action and delivering results quickly.

Nice to have

  • Prior experience in the cybersecurity, intelligence, or high-compliance industries.
  • Experience with common agentic workflow frameworks (e.g., LangChain, LlamaIndex).
  • Experience with distributed data processing frameworks (e.g., Spark, Dask, Flink).
  • Bachelor's degree in Computer Science, Data Engineering, or a related STEM field
  • Master's degree preferred

Skills

  • Hands-on experience with MLOps Tools (MLflow, Sagemaker, Vertex AI).
  • Strong understanding of CUDA, NVIDIA drivers, GPU, and TPU compute fundamentals.
  • Experience with inference serving frameworks such as vLLM and Triton Inference Server.
  • Proficiency with distributed training frameworks including Pytorch, Ray, Megatron, and JAX.
  • Expert-level proficiency in a high-level coding language (Python).
  • Deep knowledge of containerization and orchestration (Docker, Kubernetes, Slurm, Airflow).
  • Proficiency with Infrastructure as Code tooling like Terraform and Ansible.
  • Experience with cloud platforms (AWS, GCP, or OCI) and related data services.

Compensation

  • Market leader in compensation and equity awards

Benefits

  • Market leader in compensation and equity awards
  • Competitive vacation and holidays for recharge
  • Paid parental and adoption leaves

This listing is sourced directly from CrowdStrike's careers page and normalized into a canonical job model.