Netskope

Netskope

Senior Staff / Principal Machine Learning Scientist, AI Inference & Optimization

Santa Clara, California, United States · Principal

Sponsorship not specified$183k-$261kDetected 6 days ago
PythonC++Machine LearningNLPLLMsCybersecurityElectrical EngineeringResearchCommunicationCollaboration

About the role

  • The hard, interesting inference problems live here: quantization, KV-cache and memory management, sparsity, fine-tuning, and hardware acceleration under real-world resource constraints.

Responsibilities

  • Real scale to build against. Netskope's customer footprint gives you production signals most teams never see, so you deploy, validate, and iterate fast.
  • High-impact ownership. You own the model layer of a net-new product that changes the performance and economics of agentic AI.
  • You own the model layer of a net-new product that changes the performance and economics of agentic AI.
  • Real scale to build against.
  • Netskope's customer footprint gives you production signals most teams never see, so you deploy, validate, and iterate fast.
  • Build and optimize the model inference path: quantization, KV-cache optimization, batching, and latency/memory/throughput tuning on constrained, commodity hardware.
  • Fine-tune and evaluate models for bounded tasks; build eval harnesses that gate a capability to release on real accuracy, latency, and security relevance.
  • Design and grow the task execution runtime (bounded sub-agents), pushing toward dynamic task generation and context compaction.
  • Drive hardware acceleration / sparsity and support for larger models as the platform matures.
  • Partner with the systems and backend engineers to ship capabilities end-to-end and iterate on real production signals.

Requirements

  • Required skills and experience

Nice to have

  • Fine-tune and evaluate models for bounded tasks
  • 10+ years of overall industry experience, with 4+ years hands-on in ML/AI (model development, fine-tuning, and inference optimization).
  • Hands-on with fine-tuning (e.g. LoRA/QLoRA), quantization (GGUF/AWQ/GPTQ), and inference runtimes (vLLM/SGLang, TensorRT-LLM, ONNX Runtime, llama.cpp, or MLX/CoreML).
  • On-device or edge inference experience is a strong plus.
  • comfort reaching into C++ for low-level interop is a plus.
  • Solid grasp of transformer internals and the levers that move real inference performance and cost: KV cache, attention, batching, memory footprint.
  • PhD in a related field strongly preferred.

Skills

  • About Netskope
  • Louis, Bangalore, London, Paris, Melbourne, Taipei, and Tokyo.
  • Visit us at Netskope Careers.
  • Please follow us on LinkedIn and Twitter @Netskope.
  • Positions are available at Senior Staff and above.
  • Candidates are assessed individually and leveled according to their specific skills and background.

Compensation

  • At Netskope, salary is one component of our competitive total rewards package.
  • The salary range for this position is as listed below.
  • The successful candidate's starting pay will also be determined based on job-related skills, experience, qualifications, location, and market conditions.
  • For all sales roles, the posted salary range is the On Target Earnings (OTE) range for the role, which is the sum of base salary and target commission amount at 100% goal achievement.
  • In addition to salary, candidates may be eligible for other forms of compensation such as participation in a bonus plan (for non-sales roles) and a stock award program.
  • $182,500 - $260,500 USD

This listing is sourced directly from Netskope's careers page and normalized into a canonical job model.