2026 - Senior AI Operator Optimization Engineer - Contractor
We usually respond within three days
Location: Dublin, Ireland
About Huawei
Huawei is a leading global provider of information and communications technology (ICT) infrastructure and smart devices. With integrated solutions across four key domains – telecom networks, IT, smart devices, and cloud services – we are committed to bringing digital to every person, home and organization for a fully connected, intelligent world.
At Huawei, innovation focuses on customer needs. We invest heavily in basic research, concentrating on technological breakthroughs that drive the world forward. We have more than 180,000 employees, and we operate in more than 170 countries and regions.
About the IRC
Huawei Ireland Research Centre's (IRC) mission is to position Huawei as a recognized technology leader and global information and communications technology (ICT) solutions provider. We are building an industry-recognized multi-disciplinary Research Centre of experts focusing on medium-term to long-term issues to achieve this. The IRC will work closely with an open, innovative ecosystem with Huawei customers to address real-world issues. The IRC will also engage with key European universities to build a basic research capability to support Huawei technical projects.
Job Summary
We are seeking a highly motivated AI Operator/Kernel Optimization Engineer to join our AI Infrastructure Platform team in Huawei Ireland Research Center. In this role, you will focus on optimizing AI operators/kernels across diverse hardware architectures (including Huawei Ascend NPU architecture), driving high-performance execution for large-scale AI training and inference workloads. You will work closely with HQ computing software platform architects, compiler engineers, and AI researchers to maximize performance, scalability, and efficiency of AI computing systems.
We believe that open-source contributions are one of the strongest indicators of engineering excellence. We encourage applicants to include links to their GitHub/GitLab profiles, merged pull requests, technical blogs, publications, and evidence of their contributions to open-source AI communities.
Responsibilities
~ Design, optimize, and accelerate AI operators/kernels on modern computing architectures, including CPUs, NPUs, AI accelerators, and heterogeneous computing platforms.
~ Analyze hardware bottlenecks and optimize operator implementations to improve throughput, latency, memory efficiency, and power efficiency.
~ Collaborate with computing software platform architecture teams to co-design next-generation AI computing systems software stacks.
~ Develop architecture-aware optimization AI agents for deep learning operators, including compute-intensive and memory-intensive kernels.
~ Optimize communication operators for distributed AI training and inference across large-scale computing clusters.
~ Improve collective communication performance (e.g., AllReduce, AllGather, ReduceScatter, Broadcast) by optimizing communication algorithms, topology awareness, and overlap between communication and computation.
~ Collaborate with compiler and runtime teams to enable automatic operator optimization and hardware-specific code generation.
~ Benchmark, profile, and analyze AI workloads to identify optimization opportunities across the software-hardware stack.
~ Stay up to date with emerging AI hardware architectures, compiler technologies, and distributed AI systems.
Qualifications
Required
~ Master's degree or Ph.D. in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
~ Strong experience in computer architecture, hardware architecture design, or hardware performance optimization.
~ Hands-on experience optimizing AI operators or high-performance computing kernels on modern hardware platforms.
~ Strong understanding of processor architecture, memory hierarchy, cache systems, SIMD/SIMT execution, and parallel computing.
~ Experience with AI software frameworks such as vLLM, PyTorch, JAX, or ONNX Runtime.
~ Proficiency in Triton, C/C++ and Python.
~ Experience with GPU programming (CUDA, ROCm) or AI accelerator software stacks.
~ Strong performance analysis and profiling skills.
Preferred
~ Experience optimizing communication operators for distributed AI training.
~ Experience with collective communication libraries such as NCCL, RCCL, MPI, Gloo, UCX, or oneCCL.
~ Experience with AI compilers such as MLIR, TVM, XLA, Triton, LLVM, or Apache IREE.
~ Experience with AI accelerator architectures (Ascend, TPU, Graphcore IPU, Cerebras, Tenstorrent, etc.).
~ Familiarity with distributed systems, RDMA, InfiniBand, NVLink, PCIe, CXL, Ethernet fabrics, or high-performance networking.
~ Publications or open-source contributions related to AI systems, HPC, compilers, or distributed AI.
Open Source Community Experience
~ We highly value candidates with a strong track record of contributing to the AI open-source ecosystem.
~ Proven contributions to major AI open-source projects through merged pull requests (PRs), accepted code commits, feature implementations, performance optimizations, or bug fixes.
~ Demonstrable evidence of code contributions (e.g., GitHub/GitLab contribution history, merged PRs, commit records, technical proposals, or release notes).
~ Experience contributing to one or more leading AI open-source communities, Candidates with contributions to one or more of the following ecosystems are highly preferred: Triton, PyTorch, MLIR / LLVM, Hugging Face Transformers, NCCL / RCCL, JAX
~ Experience serving as a Maintainer, Core Contributor, Reviewer, Technical Steering Committee (TSC) member, Module Owner, or Community Leader in a major open-source project is highly preferred.
~ Active participation in open-source community governance, technical discussions, RFC reviews, or ecosystem development is considered a strong advantage.
Preferred Skills
~ Strong understanding of AI system software, compilers, runtime systems, and hardware-software co-design.
~ Experience in AI models performance modeling and bottleneck analysis.
~ Excellent problem-solving and analytical skills.
~ Strong communication and cross-functional collaboration abilities.
~ Passion for pushing the performance limits of next-generation AI infrastructure.
DUE TO THE HIGH VOLUME OF REPLIES, ONLY CANDIDATES WHO ARE SHORTLISTED FOR INTERVIEWS WILL BE CONTACTED.
Privacy Statement
Please read and understand our West European Recruitment Privacy Notice before submitting your personal data to Huawei so that you fully understand how we process and manage your personal data received.
http://career.huawei.com/reccampportal/portal/hrd/weu_rec_all.html
About Huawei Ireland Research Centre
Huawei Ireland Research Centre (IRC) mission is to position Huawei as a recognized technology leader and a global provider of information and communications technology (ICT) solutions. To achieve this we are building an industry-recognized multi-discipline Research Centre of experts with focus on medium-term to long-term issues. The IRC will work closely with an open innovative ecosystem with Huawei customers to address real-world issues. The IRC will also engage with key European universities to build a basic research capability to support Huawei technical projects.