Architect
Member of Technical Staff - Compilers
Palo Alto · Staff+
Sponsorship not specifiedDetected 96 days ago
PythonMachine LearningTensorFlowPyTorchFPGAResearchLeadership
About the role
- Born out of Stanford Research, our team blends AI with Silicon with a founding team from Anthropic, Google DeepMind, Meta SuperIntelligence, xAI, Apple and Intel.
- Ideal candidates have shipped production compilers at places like Apple, Google (XLA/TPU), Groq, Cerebras, Qualcomm, AMD, or similar.
Responsibilities
- As a Member of the Technical Staff on the Compilers team at Architect, you'll own the compiler stack targeting our SIMD/VLIW NPU - from graph ingestion through code generation on production silicon.
- You'll work directly with the NPU architect to co-design the ISA, closing the loop between compiler needs and hardware decisions.
- Own the compiler end-to-end: graph ingestion (ONNX, PyTorch) through IR optimization, AI-driven code generation, instruction scheduling, and register allocation for a SIMD/VLIW NPU.
- Implement and own the memory management layer
- Design and iterate on mid-end and backend optimization passes: operator fusion, loop transformations, vectorization, and software pipelining to close the gap between peak and achieved throughput.
- Co-design the ISA and instruction encoding with the architect and silicon team. Feed real workload performance data back into architectural decisions.
- Support quantization and mixed-precision lowering (32bit single-precision FP or INT, along with lower INT8/4, BF16, FP16/8/4 precisions) with correct numerics end-to-end.
- Benchmark compiler output against cycle-accurate models, RTL simulation, and FPGA prototypes. Own QoR tracking.
- Grow into a compiler team lead as the team scales.
- Domain Background: Hands-on experience on at least one of: Apple Neural Engine compiler, Google XLA / Edge TPU / TPU codegen, Groq TSP compiler (spatial scheduling, IR dialect design), Cerebras compiler stack, Qualcomm Hexagon NN / AI Engine, AMD AIE / Vitis AI, or similar/equivalent custom accelerator compiler(s).
Requirements
- Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, or a closely related field.
- Must have targeted ML/AI hardware compiler experience, as general-purpose (GCC/LLVM for CPUs) is not sufficient.
- Experience with tiling strategies, loop nest optimization, and operator fusion for ML workloads (such as convolution, attention, element-wise ops, reduction, transpositions, etc.).
- Experience with scratchpad type memory allocation, data layout, DMA orchestration, and multi-buffering.
- Python proficiency.
- Familiarity with MLIR or LLVM infrastructure.
- ML Optimizations: Experience with tiling strategies, loop nest optimization, and operator fusion for ML workloads (such as convolution, attention, element-wise ops, reduction, transpositions, etc.).
Skills
- Degree: Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, or a closely related field.
- SW-Managed Memory: Experience with scratchpad type memory allocation, data layout, DMA orchestration, and multi-buffering.
- Coding: Strong C++. Python proficiency. Familiarity with MLIR or LLVM infrastructure.
- ML framework experience (PyTorch, TensorFlow) and portable graph formats (ONNX).
- Experience benchmarking and profiling compiler output on real hardware, FPGA, or cycle-accurate simulators.
- Contributions to open-source ML compiler projects (TVM, MLIR, Triton, XLA).
- Domain-specific expertise: Track record on energy-efficient, high-performance HW accelerator bring-up.
- Competitive salary and meaningful equity stake
- Fast-paced startup with autonomy and visible impact
- Direct ownership of the compiler stack as we scale
Compensation
- Competitive salary and meaningful equity stake
Company info
- Architect is a frontier AI lab for chip design. We build AI models and tools for on-demand custom ASICs at scale. Our goal is to co-design custom ASICs alongside evolving ML workloads, and enable a new era of domain-specific chips that unlock capabilities impossible with current hardware paradigms. Born out of Stanford Research, our team blends AI with Silicon with a founding team from Anthropic, Google DeepMind, Meta SuperIntelligence, xAI, Apple and Intel.
- We're looking for staff/principal-level compiler engineers with deep experience building code generation toolchains for custom AI accelerators. Ideal candidates have shipped production compilers at places like Apple, Google (XLA/TPU), Groq, Cerebras, Qualcomm, AMD, or similar.
- Architect is a frontier AI lab for chip design.
- We build AI models and tools for on-demand custom ASICs at scale.
- Our goal is to co-design custom ASICs alongside evolving ML workloads, and enable a new era of domain-specific chips that unlock capabilities impossible with current hardware paradigms.
- We're looking for staff/principal-level compiler engineers with deep experience building code generation toolchains for custom AI accelerators.
Apply directly at Architect →Create a free account for alerts like thisView Architect immigration profile
This listing is sourced directly from Architect's careers page and normalized into a canonical job model.