Vmax
Research Fellowship - Mechanistic Interpretability
San Francisco · Staff+
Sponsorship not specifiedDetected 62 days ago
PythonMachine LearningPyTorchLLMsA/B TestingResearchCommunicationPublic Speaking
About the role
- LLMs are fantastically powerful and there is a rapidly growing corpus of work devoted to understanding their internal representations and computations.
- We use the tools of mechanistic interpretability to enhance reinforcement learning by generating intrinsic rewards as a supplement or alternative to downstream human-generated verifiers.
- This 3 to 6 month fellowship is for PhD students or equivalent early-career researchers who want to work at the intersection of mechanistic interpretability and reinforcement learning.
Responsibilities
- Develop mechanistic interpretability methods for understanding internal representations, features, circuits, and computations in language models and agents.
- Build research code, evaluation harnesses, and experimental infrastructure that make results reproducible and useful to the broader team.
Requirements
- Strong programming ability in Python and experience with at least one major ML framework such as PyTorch or JAX.
Nice to have
- Experience with LLM post-training methods
- Experience with scalable ML experimentation, distributed training, experiment tracking, or reproducible research infrastructure.
- Interest in turning mechanistic understanding into practical training methods, rather than only analyzing models after training.
- This role is based in our San Francisco office
Benefits
- Investigate how model internals can be used to generate intrinsic rewards, auxiliary objectives, diagnostics, or training signals for reinforcement learning.
- Design and run experiments that test whether interpretability-derived signals improve learning, exploration, generalization, robustness, or sample efficiency.
- Currently enrolled in a PhD program in machine learning, computer science, artificial intelligence, computational neuroscience, mathematics, or a related technical field.
- Working understanding of reinforcement learning.
Company info
- V max is an applied research lab developing AI capable of open-ended learning.
- We are building systems to exceed humans in all capacities by optimizing beyond the local maxima of learning from human expertise.
This listing is sourced directly from Vmax's careers page and normalized into a canonical job model.