Modal
Member of Technical Staff - Research, Inference
New York · Staff+
Sponsorship not specifiedDetected 16 days ago
PythonNode.jsRESTLLMsValuationResearch
About the role
- Most of the value of owning a model shows up at serving time.
- We already run elastic inference, sandboxes, distributed volumes, and multi-node training, and we control the infrastructure underneath, so the serving stack is ours to shape rather than something we resell.
- The bets that matter most are the ones that move cost per token and tail latency on the workloads our customers actually run.
Responsibilities
- Own end-to-end inference research bets: speculative decoding, disaggregated prefill/decode, quantization (FP8, INT4), KV-cache and memory management, autoscaling for spiky serverless traffic, and whatever else the research agenda calls for.
- our work with ZLab on DFlash https://modal.com/blog/spec-is-all-u-need, a speculator design built on KV injection and blockwise parallel drafting
Requirements
- A research-leaning or systems background in LLM inference, with work you can point to.
- Ability to work in-person, in our NYC or San Francisco office.
Company info
- AI needs a new infrastructure layer. We're building it at Modal.
- Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now.
- Our customers include category-defining companies like Lovable https://modal.com/blog/lovable-case-study, Ramp https://modal.com/blog/how-ramp-built-a-full-context-background-coding-agent-on-modal, Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale.
- We recently raised a $355M Series C https://modal.com/blog/modal-series-c at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September.
- AI needs a new infrastructure layer.
- We're building it at Modal.
- Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud.
- Each time, the company that rebuilt the layer underneath defined the decade.
- AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now.
- Our customers include category-defining companies like Lovable https://modal.com/blog/lovable-case-study, Ramp https://modal.com/blog/how-ramp-built-a-full-context-background-coding-agent-on-modal, Cognition, DoorDash, and Suno.
- They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale.
- We recently raised a $355M Series C https://modal.com/blog/modal-series-c at a $4.65B valuation, led by General Catalyst and Redpoint Ventures.
- We've crossed $300M+ ARR and grown fivefold since September.
- Our team includes creators of popular open-source projects (e.g.,Seaborn https://github.com/mwaskom/seaborn,Luigi https://github.com/spotify/luigi), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience.
This listing is sourced directly from Modal's careers page and normalized into a canonical job model.