DDN
Platform Support Architect
Remote - California
Sponsorship not specifiedDetected 6 days ago
ElasticsearchVector DatabasesDockerKubernetesHelmCI/CDLinuxPrometheusGrafanaSite Reliability EngineeringPlatform EngineeringMachine LearningLLMsRAGAgentic AIMLOpsComplianceHIPAALeadershipCommunication
About the role
- DDN is expanding our Enterprise and Sovereign AI Solution offerings, for example Hyperpod - a turnkey NVIDIA AI Data Platform built on DDN Infinia storage, NVIDIA AI Enterprise (NVAIE), and Supermicro reference hardware, optimized for inference and RAG workloads. Our support organization is deep on storage (Infinia, EXAScaler); we are now hiring an AI
- platform specialist to lead supportability and enablement for the AI side of the stack – NVIDIA AI Enterprise services (NIMs, NeMo, Triton, GPU Operator, licensing), vector databases (initially Milvus), RAG/agentic workflows, and the high‑performance storage and networking fabric that underpins them.
Responsibilities
- Act as the primary NVIDIA AI Enterprise and vector database solutions expert for HyperPOD customer environments, bringing deep knowledge of NVAIE services (e.g., NIMs, NeMo, Triton, TensorRT/TensorRT‑LLM, GPU Operator, licensing/NLS) and vector databases (e.g., Milvus) to guide diagnosis, optimization, and solution design.
- Own complex end‑to‑end triage across GPU, NVAIE services, vector DB, Kubernetes, Docker, high‑speed networking, and Infinia storage, distinguishing product defects from environmental and integration issues.
- build minimal repros and high‑quality defect reports for escalation to NVIDIA, vector‑DB vendors, OEMs, and internal engineering.
- Author and maintain support triage runbooks and checklists for HyperPOD covering NVAIE services, Milvus/vector DB, GPU stack, Docker, Kubernetes resources, and their interaction with Infinia and the network fabric.
- Define and validate unified diagnostics bundles that capture the right logs/configs/metrics from all relevant layers (Infinia, GPUs, NVAIE, Milvus, Kubernetes, network) to enable fast problem isolation and high‑signal escalations.
- Build hands‑on labs and PoCs that mirror customer RAG and agentic AI use cases on HyperPOD, validating supportability and capturing "known good" configurations and troubleshooting patterns.
- DESIGN FEEDBACK, READINESS, AND CROSS‑FUNCTIONAL LEADERSHIP
- Collect and interpret logs and telemetry across Linux, containers, Kubernetes, GPU stack, vector DB, and storage/networking; build minimal repros and high‑quality defect reports for escalation to NVIDIA, vector‑DB vendors, OEMs, and internal engineering.
- Collaborate closely with NVIDIA solutions architects, OEM architects, PS, and Support Innovation to align reference architectures and best practices with real‑world support experience.
- Solid understanding of RAG and Generative AI workflows: embeddings, retrieval, reranking, prompt design, context management, and how these interplay with vector search and GPU inference at scale.
Requirements
- Strong hands‑on experience with containers and Kubernetes (Docker/containerd, Helm, Operators
- Experience with one or more vector databases (Milvus, Qdrant, Pinecone, pgVector, OpenSearch/Elasticsearch vectors, etc.), including schema design, ingestion, and operations.
- Practical experience with AI storage and networking for HPC/AI clusters:
Nice to have
- Prior experience with scale‑out storage in GPU/AI environments.
- Hands‑on work with NVIDIA reference blueprints (Enterprise RAG, VSS, AIQ, industry‑specific blueprints) or similar enterprise AI architectures.
- Familiarity with AI observability and responsible AI practices (guardrails, monitoring for drift/toxicity, basic understanding of regulatory considerations like GDPR/HIPAA in the context of AI systems).
- Experience with observability stacks (Prometheus, Grafana, Loki/ELK, NetQ, etc.) tuned for AI workloads, including service‑level dashboards and SLOs.
- Within 6-12 months, a successful AI Data Platform Solutions Architect will have:
- NVAIE, vector DB, and AI‑workflow issues through high‑quality diagnostics, architecture insight, and well‑defined "golden stack" patterns.
- Established clear, repeatable triage and escalation patterns for AI‑side incidents that L1/L2 storage engineers can follow with confidence.
- 8+ years total technical experience preferred.
Skills
- Demonstrated experience operating GPU‑accelerated workloads in production:
- NVIDIA GPUs, drivers, CUDA concepts, GPU utilization/perf triage
- NVIDIA GPU Operator and Kubernetes‑based GPU lifecycle management
- Familiarity with DGX / HGX or similar GPU cluster platforms.
Company info
- Develop reusable technical assets - implementation guides, best‑practice playbooks, tuning checklists, example architectures - to accelerate time‑to‑value for customers, PS, and Support.
This listing is sourced directly from DDN's careers page and normalized into a canonical job model.