Windsurf

Windsurf

Research Engineer, Infrastructure

San Francisco

Sponsorship not specifiedDetected 106 days ago
PythonC++Distributed SystemsMachine LearningDeep LearningPyTorchData EngineeringSystems EngineeringResearchCommunicationPipeline Integrity

About the role

  • We're the makers of Devin, the first AI software engineer, and Windsurf, the AI-native IDE.
  • Together, they represent our vision for collaborative AI teammates that enable engineers to focus on more interesting problems and empower teams to strive for more ambitious goals.
  • Among our founding team, we have world-class competitive programmers, former founders, and leaders from companies at the cutting edge of AI including Scale AI, Palantir, Cursor, Waymo, Tesla, Lunchclub, Modal, Google DeepMind, and Nuro.

Responsibilities

  • This is not a support role.
  • You will work directly alongside researchers, understand the science deeply enough to anticipate what they need next, and build systems that hold up under the pressure of training jobs running across thousands of GPUs.
  • Build and own the systems that run large-scale training jobs reliably across GPU clusters.
  • Own the infrastructure that runs hundreds of thousands of concurrent coding agent rollouts in VM sandboxes, from high-fidelity environment design to the distributed systems that hold up at our largest RL training scales.
  • Implement solutions that meaningfully improve step time and MFU at scale.
  • Design and maintain the systems researchers use to launch, track, and analyze experiments.
  • Build high-throughput, reliable data pipelines for training and evaluation.
  • Maintain detailed understanding of failure modes and build systems that fail gracefully and recover fast.
  • Implement and optimize parallelism strategies: data, tensor, pipeline, and sequence parallelism.
  • Deep experience building and operating distributed training systems for large models; comfortable owning infrastructure end to end from the cluster level down to the training loop

Requirements

  • Proficiency in Python and C++
  • Hands-on experience with GPU performance profiling, memory optimization, and compute efficiency
  • Experience implementing or optimizing parallelism strategies (data, tensor, pipeline, sequence) for large model training

Equal opportunity

  • Cognition is an equal opportunity employer.
  • We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic under applicable law.
  • We are committed to providing reasonable accommodations for candidates with disabilities throughout the hiring process - please let us know if you need any.

This listing is sourced directly from Windsurf's careers page and normalized into a canonical job model.