Overview

Huawei Canada has an immediate permanent opening for a Distinguished Engineer - AI Computing System.

About the team

The Advanced Computing and Storage Lab, part of the Vancouver Research Centre, aims to explore adaptive computing system architectures to address the challenges posed by flexible and variable application loads. It supports stability and quality of training clusters, develops dynamic cluster configuration strategy solvers, and establishes precision control systems to create stable and efficient computing power clusters. The lab focuses on industry AI application scenarios such as large model training and inference, based on technologies like low-precision training, multi-modal training, and reinforcement learning, with responsibility for bottleneck analysis and the design and development of optimization solutions to improve training and inference performance and usability.

About the job

  • As a leading expert in training cluster software frameworks and technologies, gain insights into the evolution of industry AI large model training frameworks and key features. Plan and layout AI frameworks and software features for scenarios such as large model pre-training, post-training, and integrated training and inference, building key capabilities for the company\'s training cluster software framework.
  • In the field of large model training optimization, lead the team to build key technologies such as low-precision training, parallel strategy tuning, and training resource optimization, promoting the commercial implementation of large model perception optimization-related technologies.
  • Lead the team to build large model AI training frameworks, operator libraries, acceleration libraries, and other software frameworks and acceleration features for training servers and super nodes, leveraging system engineering and software-hardware collaboration to enhance AI cluster computing efficiency.
  • Identify high-quality academic resources in large model training, collaborate with domain experts and scholars on projects, layout related standards and patents, and support ongoing innovation in the training cluster field to build long-term competitiveness.
  • Cultivate a team of technical experts and key technical staff in AI training cluster frameworks and software optimization.

The base salary for this position ranges from $172,000 to $230,000 depending on education, experience and demonstrated expertise.

About the ideal candidate

  • Major in artificial intelligence, computer science, software, automation, physics, mathematics, electronics, microelectronics, information technology, or related fields, with more than 5 years of R&D experience in large model training and optimization.
  • Proficient in common model structures of large models such as Deepseek and Llama, with deep technical expertise in large model training and inference optimization in fields like LLM, MoE, and multimodal learning.
  • Familiar with the hardware architecture and programming systems of AI accelerators such as GPU and NPU, with experience in optimizing AI systems with software-hardware-cores collaboration.
  • Familiar with cluster computing and cloud computing, with experience in software architecture design for cluster scheduling.
  • Enjoys research, has strong learning ability, good communication skills, and teamwork ability.

#J-18808-Ljbffr
Similar jobs

Distinguished Engineer - AI Computing System

Apply Now
Back to search page