Persimmons AI
Compiler Engineer (Mid and/or Backend)
San Jose, California, United States · Mid
Sponsorship not specifiedDetected 143 days ago
PythonC++ExpressAlgorithmsMachine LearningPyTorchLLMsResearchCollaboration
About the role
- Why join us: We're growing fast and looking for bold thinkers, builders, and curious problem-solvers who want to push the limits of AI hardware and software.
- If you're ready to join a world-class team and play a critical role in making a global impact - we want to talk to you.
- Summary of Role: This role focuses on transforming higher-level MLIR-based large language models by applying sophisticated mid- and backend compiler techniques to target Persimmons.ai's custom accelerator hardware.
Responsibilities
- Develop and enhance MLIR-based compiler pipelines targeting Persimmons' custom spatial accelerator hardware.
- Design and optimize the Persimmons Compiler mid- and backend techniques for efficient lowering, graph-to-resources mapping, and code generation.
- Implement transformations to convert Python, PyTorch, and similar kernel representations to LLVM IR and runtime-ready libraries.
- Architect and implement efficient support for SPMD-based, distributed collective operations and lower them through specialized MLIR compiler dialects (e.g., MESH, SHARDY).
- Drive advanced loop optimizations leveraging polyhedral analysis: loop tiling, fusion, interchange, skewing, and related techniques.
- Apply and optimize techniques such as bufferization, padding, inlining, and integration of custom operations and kernels within the compilation workflow.
- Collaborate across teams to deliver performant compilation flows from high-level ML representations to low-level executable artifacts.
- You will help design and optimize the Persimmons Compiler mid- and backend, integrate it with custom operations and kernels, as well as implement compiler passes that convert higher-level intermediate representations into runtime-oriented code and libraries.
Requirements
- strong candidates will demonstrate expertise in several key areas.
- Solid understanding and experience with underlying principles and methods of the MLIR framework (SSA representation, interfaces, rewriting, dialect hierarchy, etc.).
- Hands-on experience with developing MLIR-based compiler infrastructure, algorithms, and techniques for non-GPU/custom spatial hardware architectures.
- Working experience with lowering SIMD operations from PyTorch, Triton, xDSL, pyDSL, or similar Python-based frontends toward LLVM IR and, further, to SIMD kernel library.
- Experience and understanding of SPMD-based, distributed collective operations, specialized MLIR compiler dialects (e.g., MESH, SHARDY), and collective operation lowering in compilers for spatial hardware.
- Experience with techniques such as padding, bufferization, inlining, and other lowering techniques.
- Knowledge of register allocation and instruction scheduling in spatial architectures.
- Experience in lowering and integration of custom operations and kernels at the compiler mid- and backend.
- Familiarity with graph and tensor partitioning and mapping optimization algorithms and their integration in the compiler workflow.
- High level of understanding and 5+ years of experience with C++ and appreciation for writing clean and maintainable code.
Nice to have
- Good knowledge of Python is a big plus.
Company info
- Persimmons is building the infrastructure that will power the next decade of AI.
- Founded in 2023 by veteran technologists from the worlds of semiconductors, AI systems, and software innovation, We're on a mission to enable smarter devices, more sustainable data centers, and entirely new applications the world hasn't imagined yet.
- This position offers the opportunity to directly shape Persimmons.ai's innovative AI hardware and software stack through close collaboration with teams across hardware, systems, and software.
- Work on register allocation and instruction scheduling for Persimmons' spatial hardware, ensuring high resource utilization, throughput, and low latency.
- Contribute to graph and tensor partitioning logic for optimal hardware-targeted execution.
- What You Bring To The Table: We do not expect candidates to meet all of the requirements listed below; strong candidates will demonstrate expertise in several key areas.
- Extensive experience and understanding of loop optimization based on polyhedral principles.
- Demonstrated fluency with modern AI tools and workflows (e.g., leveraging AI assistants for research, analysis, or productivity).
- Competitive salary and benefits package.
Apply directly at Persimmons AI →Create a free account for alerts like thisView Persimmons AI immigration profile
This listing is sourced directly from Persimmons AI's careers page and normalized into a canonical job model.