Huawei Technologies Canada Co., Ltd.
Intern Engineer – RL Post-Training for LLMs
Vancouver, British Columbia, Canada · Intern · Internship
Sponsorship not specified$58k-$104kDetected 44 days ago
PythonFull-Stack DevelopmentAlgorithmsMachine LearningPyTorchData AnalysisNLPLLMsA/B TestingAR/VRCommunicationProblem Solving
About the role
- One of the goals of this lab are to enhance algorithm performance and training efficiency across industries, fostering long-term competitiveness.
- Conduct experiments to improve model performance, reasoning, and alignment.
- About the ideal candidate: Enrolled as Master or Ph.D. student in Computer Science, AI, or related field.
Responsibilities
- This team focuses on full-stack innovations, including software-hardware co-design and optimizing data efficiency at both the storage and runtime layers.
- About the job: Develop and optimize RL post-training pipelines for LLMs (e.g., GRPO, reward modeling).
- Build scalable training, evaluation, and data generation systems.
- Collaborate with researchers and engineers on cutting-edge LLM projects Stay current with advancements in RL, LLMs, and post-training research.
Nice to have
- Enrolled as Master or Ph.D. student in Computer Science, AI, or related field.
- Familiarity with Large Language Models, transformer architectures, and post-training methods.
- Proficiency in Python, PyTorch, and LLM frameworks.
- Hands-on experience with LLMs and RL training algorithms (e.g., GRPO) is an asset.
- Familiarity with RL frameworks, such as VeRL.
- Experience with open-source LLM frameworks such as Hugging Face, DeepSpeed, vLLM, or SGLang is an asset.
- Knowledge of domain-specific languages used with AI accelerators.
- Experience with distributed training frameworks, large-scale experimentation, or LLM infrastructure is an asset.
Compensation
- Develop and optimize RL post-training pipelines for LLMs (e.g., GRPO, reward modeling).
- Conduct experiments to improve model performance, reasoning, and alignment.
- Build scalable training, evaluation, and data generation systems.
- Collaborate with researchers and engineers on cutting-edge LLM projects Stay current with advancements in RL, LLMs, and post-training research.
- The total target annual compensation (based on 2,080 hours per year) ranges from $58,000 to $104,000 depending on education, experience, and demonstrated expertise.
Company info
- The Computing Data Application Acceleration Lab aims to create a leading global data analytics platform organized into three specialized teams using innovative programming technologies.
- This team also develops next-generation GPU architecture for gaming, cloud rendering, VR/AR, and Metaverse applications.
- The total target annual compensation (based on 2,080 hours per year) ranges from $58,000 to $104,000 depending on education, experience, and demonstrated expertise.
- Strong background in machine learning, reinforcement learning, and deep learning.
- All applications for this position are reviewed directly by our hiring team, we do not use artificial intelligence tools to screen or select candidates.
Apply directly at Huawei Technologies Canada Co., Ltd. →Create a free account for alerts like thisView Huawei Technologies Canada Co., Ltd. immigration profile
This listing is sourced directly from Huawei Technologies Canada Co., Ltd.'s careers page and normalized into a canonical job model.