Oracle
Lead Principal Machine Learning Engineer
San Francisco, California, USA · Principal · full-time
Stay score
odds of building a lasting career here
Sponsors, but it's cap-subject — you still face the weighted lottery (~61% per draw at Level IV). Good if you win; have a cap-exempt backup on your list.
Lottery odds assume a STEM candidate.
Personalize to your clock →H-1B wage level
the lottery is wage-weighted — each level is one more entry
$23,349 more — $193,149 — moves this role to Level III and 3 lottery entries. That figure is inside the range the employer already advertised.
Based on the DOL prevailing wage for this occupation and worksite, a base salary of $193,149 would place this position at wage Level III. That figure is within the posted range, and I'd like to target it. This role classifies under "Software Developers" for prevailing-wage purposes.
DOL prevailing wage, 2026-27 wage year · Software Developers (15-1252) · San Francisco-Oakland-Fremont, CA. Wage level is derived by USCIS from the offered wage, occupation and worksite; the occupation shown is inferred from the job title.
Community outcomes
No reports yet — be the first to help the next applicant.
About the role
- This person will set architecture and engineering direction for production-grade agentic AI platforms, autonomous workflows, scalable inference infrastructure, and enterprise AI applications used in large-scale, business-critical environments.
- The expectation is to ship, scale, and operate reliable, secure, observable, and cost-aware AI platform systems while raising the technical bar for engineers across the organization.
- Responsibilities Serve as a senior technical owner for OCI AI platform capabilities, including agent execution, inference systems, model serving, AI workflow orchestration, evaluation, and observability.
Responsibilities
- Design, architect, and deliver scalable agentic AI systems capable of reasoning, planning, tool use, workflow execution, multi-step task orchestration, and safe human-in-the-loop escalation.
- Build production-grade services for tool calling, agent memory, context management, Model Context Protocol (MCP) integration, vector retrieval, multi-agent coordination, policy enforcement, and evaluation.
- Drive technical strategy across infrastructure, platform, security, data, and application engineering teams, converting broad goals into executable multi-quarter plans and measurable milestones.
- Own critical production outcomes, including reliability, performance, security posture, cost efficiency, and supportability for the systems delivered.
Requirements
- Bachelor's, Master's, or Ph.D. in Computer Science, AI/ML, Engineering, or a related field, or equivalent practical experience.
- 12+ years of professional software engineering experience, including significant ownership of production systems
- Proven track record as a Staff, Senior Staff, Principal, or equivalent technical leader influencing architecture and execution across multiple teams.
- Hands-on experience with production AI systems, agentic AI applications, autonomous workflows, tool-using agents, multi-step orchestration, or multi-agent systems.
- Practical experience with orchestration frameworks such as LangGraph, LangChain, CrewAI, AutoGen, LlamaIndex, or similar ecosystems.
- Strong programming skills in Python and ability to contribute high-quality production code, reviews, tests, and debugging in complex distributed environments.
- Strong expertise with Kubernetes, Docker, cloud-native infrastructure, service-to-service communication, scalability, fault tolerance, observability, and performance analysis.
- Experience defining SLIs/SLOs, production readiness criteria, incident response practices, monitoring, tracing, experiments, and reliability programs for AI or distributed systems.
- Strong understanding of AI safety, governance, security, and operational risks for autonomous or semi-autonomous systems, including data handling, access control, auditability, and human accountability.
- The ideal candidate combines deep distributed systems experience with practical AI-native engineering, including orchestration of LLMs, tools, APIs, memory, retrieval, evaluation, guardrails, and cloud services.
Nice to have
- Experience optimizing large-scale GPU inference or training workloads for latency, throughput, utilization, availability, and cost.
- Experience integrating AI systems with enterprise APIs, databases, cloud services, vector databases, embeddings, retrieval systems, identity systems, and policy enforcement layers.
- Experience with LLM fine-tuning, long-context systems, reasoning models, model routing, caching, batching, quantization, or emerging generative AI research.
- Experience using AI-assisted software development tools such as Codex, Claude Code, Cursor, Copilot, or similar systems in large-scale engineering environments.
- Track record of defining architectural standards, platform capabilities, or engineering practices adopted across multiple teams or organizations.
- Experience in enterprise, cloud infrastructure, regulated, security-sensitive, or mission-critical environments.
- 401(k) Savings and Investment Plan with company match 8.
- 11 paid holidays 10.
Compensation
- $170k-$355k
Company info
- Career Level - IC5 About Us Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care.
- And with AI embedded across our products and services, we help customers turn that promise into a better future for all.
This listing is sourced directly from Oracle's careers page and normalized into a canonical job model.