Seekr

Seekr

Research Engineer, Foundation Model Training, SeekrGEO

Washington, District of Columbia

Sponsorship not specifiedDetected 13 days ago
PythonCode ReviewCI/CDMachine LearningPyTorchData EngineeringNLPAgentic AICadenceResearchCommunicationCollaborationRemote Sensing

About the role

  • We care deeply about transparency, auditability, and defensibility because high-stakes AI is only useful when people can understand and trust how it behaves.
  • About The Opportunity SeekrGEO is Seekr's geospatial AI product.
  • This role contributes to the foundation model program behind it: pretraining and post-training of large multi-modal models on geospatial data, together with the distributed training systems that make that work possible at scale.

Responsibilities

  • Own the parallelism strategy for our training workloads: FSDP, tensor / pipeline / sequence parallelism, ZeRO variants, activation and gradient checkpointing, mixed precision, and the memory and throughput tradeoffs that come with each.
  • Design and operate the data pipeline for large training corpora: sharded formats, streaming loaders, deduplication, mixture tuning, and the versioning discipline that makes runs reproducible.
  • Build the evaluation infrastructure that makes model comparisons trustworthy and reproducible, both during training and after deployment.
  • Support deployed models through their lifecycle: monitor systems behavior in production, diagnose regressions, and close the loop back into the next training cycle.
  • Partner with Research Scientists to pressure-test ideas: reproduce a paper's core claim, verify a proposed recipe scales, and turn research prototypes into production runs.
  • Partner with the SeekrGEO product team and customer-facing teams to align training infrastructure with the workflows the model needs to support.

Requirements

  • Hands-on experience with at least one large-scale training framework such as Megatron-LM, torchtitan, or DeepSpeed.
  • Ability to move comfortably between engineering and research.
  • You can read a paper, reproduce its core idea, and pressure-test whether it will hold up at scale.
  • Comfort designing experiments and evaluating ambiguous technical tradeoffs.
  • Experience deploying or distilling large models for inference under real latency and cost constraints.

Nice to have

  • Experience operating distributed training at scale across accelerator clusters, with comfort in collective communication and the failure modes specific to large-scale runs.
  • Hands-on experience with Megatron-LM, torchtitan, and other distributed training frameworks.
  • Performance work on accelerators: kernel-level profiling, mixed precision, activation and gradient checkpointing, attention kernels, memory layout optimization.
  • Experience with AMD ROCm is a strong plus
  • CUDA / NVIDIA experience translates directly and is welcome.
  • Experience with data infrastructure for large training corpora: sharded formats, deduplication, streaming pipelines, mixture tuning.
  • Experience with checkpoint management, fault-tolerant and elastic training, and the operational hygiene needed for multi-week runs.
  • Experience with experiment tracking, model and data versioning, evaluation pipelines, and diagnosing production issues in trained models.

Skills

  • Seekr's Mission
  • Seekr builds trusted AI for mission-critical decisions.
  • About The Opportunity
  • SeekrGEO is Seekr's geospatial AI product.

Benefits

  • Equity Ownership - RSUs that let you share directly in Seekr's long‑term success and growth.
  • Time Off That Respects Real Life - Unlimited PTO plus 14 paid company holidays to truly recharge.
  • Competitive Total Rewards - A role‑appropriate compensation structure that supports long‑term growth, including base salary, bonuses, or commission plans depending on role.
  • 401(k) with Company Match - Build your future with a retirement plan that includes employer matching.
  • Comprehensive Health & Wellness - Medical, dental, vision, and life insurance coverage starting day one-for you and your family.
  • Parental Leave - Paid parental leave to support employees as they welcome a new child through birth, adoption, or foster placement.
  • Work Your Way - A flexible hybrid work environment with offices in Reston, VA and Austin, TX.

Company info

  • Seekr is a leader in explainable and trustworthy artificial intelligence designed to power mission-critical decisions in enterprises, government, and regulated industries.
  • SeekrFlow™, our end-to-end AI platform, provides secure, auditable AI solutions tailored to sectors where transparency, accuracy, and compliance are paramount.
  • Available across cloud, on-premises, and edge environments, SeekrFlow reduces bias, strengthens data integrity, and simplifies model oversight so organizations can rely on trusted AI decisions in high-stakes settings that impact society's most sensitive and vital systems.
  • Trusted by leading enterprises and government agencies, we partner with defense, finance, telecom, and critical infrastructure leaders to enable AI solutions that drive real-world results with unmatched transparency and control.
  • We are a team of strategic thinkers and problem-solvers tackling the toughest challenges facing critical infrastructure and global enterprises through best-in-class AI models and customer deployment.
  • Our team operates with unwavering commitment to our core values and mission:
  • We are driven by outcomes-our customers' success is what we strive for every day.
  • We believe trust is earned, which is why we build explainability and transparency into the entire AI lifecycle.
  • We take our responsibility to deliver secure AI seriously.
  • We believe innovation drives progress-we are building the technologies that power the systems our society depends on.
  • Company Benefits:
  • Meaningful Mission & Impact - Work with a deeply talented, collaborative team solving some of the toughest AI challenges that matter.

This listing is sourced directly from Seekr's careers page and normalized into a canonical job model.