Huawei Technologies Canada Co., Ltd.

Huawei Technologies Canada Co., Ltd.

Co-op Software Engineer - AI System & Infrastructure

Vancouver, British Columbia, Canada · Intern · Internship

Sponsorship not specified$58k-$104kDetected 71 days ago
PythonJavaGoRustC++C#Distributed SystemsAlgorithmsCloud PlatformsKubernetesPyTorchLLMsCommunication

About the role

  • The Intelligent Cloud Infrastructure Lab aims to innovate technologies, algorithms, systems, and platforms for next-generation cloud infrastructure.
  • The lab addresses scalability, performance, and resource utilization challenges in existing cloud services while preparing for future challenges with appropriate technologies and architectures.

Responsibilities

  • Collaborate with internal and external teams to deliver the project or project features that improve our overall system scalability and performance.

Nice to have

  • Bachelors, Master/PhD degree in Computer Science, Computer Engineering Experience in building large scale and high-performance distributed system Experience in Nvidia TensorRT and/or Triton servers.

Skills

  • vLLM, Ray, SGLang, Kubernetes, TensorRT-LLM, Pytorch framework, Cuda libraries, GPU technologies

Compensation

  • Collaborate with internal and external teams to deliver the project or project features that improve our overall system scalability and performance.
  • The target annual compensation (based on 2080 hours per year) ranges from $58,000 to $104,000 depending on education, experience and demonstrated expertise.

Company info

  • Additionally, the lab aims to understand industry dynamics and technology trends to create a robust ecosystem.
  • About the job: Understand AI System and Infrastructure technology landscape, and identify scalability/performance issues or challenges of current LLM/multi-modal LLM systems Initiate and charter innovation projects to build or re-architect AI infrastructure platform, and plan milestones accordingly Provide/contribute a scalable and high-performance architecture design or re-design for the infrastructure system that is optimized for AI training and inferencing, which includes but not limited to cluster management and scheduling, LLM model deployment, elastic LLM as well as AI container cold/warm start-up optimization, and so on.
  • The target annual compensation (based on 2080 hours per year) ranges from $58,000 to $104,000 depending on education, experience and demonstrated expertise.
  • About the ideal candidate: Bachelors, Master/PhD degree in Computer Science, Computer Engineering Experience in building large scale and high-performance distributed system Experience in Nvidia TensorRT and/or Triton servers.
  • Experience in container virtualization technologies Knowledge & experience in distributed system design & development, including serverless technologies Work experience in one or more of the following

This listing is sourced directly from Huawei Technologies Canada Co., Ltd.'s careers page and normalized into a canonical job model.