Enigma

Enigma

Machine Learning Engineer

San Jose, California

Sponsorship not specifiedDetected 3 days ago
PythonSQLNoSQLVector DatabasesCI/CDMachine LearningDeep LearningPyTorchData EngineeringMLOpsLoad BalancingResearchCollaboration

About the role

  • Machine Learning Engineer | Python | Pytorch | Distributed Training | Optimisation | GPU | Hybrid, San Jose, CA
  • • Productize and optimize models from Research into reliable, performant, and cost-efficient services with clear SLOs (latency, availability, cost).

Responsibilities

  • Productize and optimize models from Research into reliable, performant, and cost-efficient services with clear SLOs (latency, availability, cost).
  • Scale training across nodes/GPUs (DDP/FSDP/ZeRO, pipeline/tensor parallelism) and own throughput/time-to-train using profiling and optimization.
  • Implement model-efficiency techniques (quantization, distillation, pruning, KV-cache, Flash Attention) for training and inference without materially degrading quality.
  • Build and maintain model-serving systems (vLLM/Triton/TGI/ONNX/TensorRT/AITemplate) with batching, streaming, caching, and memory management.
  • Partner with ML Ops on CI/CD, telemetry/observability, model registries; partner with Scientists on reproducible handoffs and evaluations.

Requirements

  • 3-5 years in ML/AI engineering roles owning training and/or serving in production at scale.
  • Experience collaborating across Research, Platform/Infra, Data, and Product functions.

Nice to have

  • Exposure to large model training techniques (DDP, FSDP, ZeRO, pipeline/tensor parallelism)
  • distributed training experience a plus
  • Bachelors in computer science, Electrical/Computer Engineering, or a related field required
  • Master's preferred (or equivalent industry experience).

Skills

  • Familiarity with deep learning frameworks: PyTorch (primary), TensorFlow.
  • Exposure to large model training techniques (DDP, FSDP, ZeRO, pipeline/tensor parallelism); distributed training experience a plus
  • Scalable serving: autoscaling, load balancing, streaming, batching, caching; collaboration with platform engineers.
  • Data & storage: SQL/NoSQL, vector stores (FAISS/Milvus/Pinecone/pgvector), Parquet/Delta, object stores.
  • Write performant, maintainable code
  • Understanding of the full ML lifecycle: data collection, model training, deployment, inference, optimization, and evaluation.
  • Machine Learning Engineer | Python | Pytorch | Distributed Training | Optimisation | GPU | Hybrid, San Jose, CA

Benefits

  • Educational Qualifications:

This listing is sourced directly from Enigma's careers page and normalized into a canonical job model.