Zenotis Technologies INC

Zenotis Technologies INC

AI/ML Engineer

Burlingame, California, USA · full-time

Sponsorship not specifiedDetected 72 days ago
PythonGoRustDistributed SystemsVector DatabasesAWSGCPAzureCloud PlatformsDockerKubernetesTerraformCI/CDDevOpsMachine LearningLLMsRAGAgentic AILLMOpsLangGraphAI OrchestrationSystems EngineeringResearch

About the role

  • In this role, you will act as the bridge between our AI research/development teams and our production environments.
  • Act as the first line of defense for production AI systems, diagnosing and resolving issues related to memory limits, inference queues, and cluster failures.

Responsibilities

  • Optimize inference latency, throughput, and cost using modern serving frameworks (e.g., vLLM, Triton Inference Server, Ray Serve).
  • Manage and orchestrate GPU/TPU clusters, ensuring high utilization and efficient resource allocation.
  • Building and Scaling Agentic Operations (AgentOps) Architect and deploy infrastructure to support autonomous AI agents and multi-agent systems.
  • Integrate and maintain agent orchestration frameworks (e.g., LangGraph, CrewAI) within production environments.
  • Observability, Evaluation, and Reliability Implement comprehensive observability stacks tailored for LLMs and agents (tracing, prompt logging, cost tracking) using tools like Langfuse, Arize, or Datadog.
  • Design automated evaluation pipelines to monitor agent performance, safety, and reliability in real-time (LLMOps/AgentOps).
  • Developer Platform & CI/CD for AI Build internal developer platforms and tooling that allow AI engineers and data scientists to easily deploy models and agents to production.

Requirements

  • you will be designing the high-performance, distributed systems required to serve Large Language Models (LLMs), orchestrate multi-agent workflows, and optimize GPU compute at scale.
  • If you are passionate about turning complex AI capabilities into highly reliable, scalable, and cost-efficient production systems, this is the role for you.
  • Build robust state management and memory systems (vector databases, graph databases) required for agentic workflows.
  • Proficiency in Python (essential for the AI ecosystem) and systems languages like Go or Rust.
  • Hands-on experience with LLM serving engines (vLLM, TGI, Triton) and distributed computing frameworks (Ray).
  • Familiarity with modern agentic development frameworks like LangChain, LangGraph, or CrewAI.

Skills

  • Systems Engineering: Strong background in distributed systems, backend engineering, or DevOps/SRE.
  • Programming: Proficiency in Python (essential for the AI ecosystem) and systems languages like Go or Rust.
  • Containerization & Orchestration: Deep expertise in Kubernetes (K8s), Docker, and infrastructure-as-code (Terraform, Pulumi).
  • AI/ML Tooling: Hands-on experience with LLM serving engines (vLLM, TGI, Triton) and distributed computing frameworks (Ray).
  • Agent Frameworks: Familiarity with modern agentic development frameworks like LangChain, LangGraph, or CrewAI.
  • Experience with vector databases (Pinecone, Milvus, Qdrant) and retrieval-augmented generation (RAG) pipelines.
  • Understanding of model optimization techniques (quantization, LoRA, KV caching).
  • Strong background in distributed systems, backend engineering, or DevOps/SRE.
  • Experience managing high-performance compute (GPTPUs) on major cloud providers (AWS,

Benefits

  • Machine Learning Infrastructure & Serving Design, build, and manage scalable infrastructure for training, fine-tuning, and serving LLMs and multimodal models.

This listing is sourced directly from Zenotis Technologies INC's careers page and normalized into a canonical job model.