Machine Learning Engineer — GPU Kernel
Quick Summary
About the Institute of Foundation Models We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research,
The GPU Kernel Engineer will play a role at the forefront of optimizing performance for the machine learning software stacks, especially at training and inference, and support the team to develop new and cutting-edge systems. The ideal candidate will have a strong background in parallel computing, and hands-on experience in system level coding, debug methodologies, and large-scale machine learning experience.
This role focuses on CUDA kernel development and optimization. Distributed training experience is a plus.
- Understand, analyze, profile, optimize, and provide guidance to the team on deep learning workloads on state-of-the-art hardware and software platforms to improve their efficiency with different levels of optimization
- Design and implement performance benchmarks and testing methodologies to evaluate application performance
- Build tools to automate workload analysis, workload optimization, and other critical workflows
- Triage system issues and identify bottleneck and inefficiencies by analyzing the sources of issues and the impact on hardware, network and propose solutions to enhance GPU utilization
- Support the team to develop appropriate kernels and systems for new model architectures and algorithms
- Participate in, or lead design reviews with peers and stakeholders to decide amongst available technologies.
- Review code developed by other developers and provide feedback to ensure best practices (e.g., style guidelines, checking code in, accuracy, testability, and efficiency).
- Contribute to existing documentation or educational content and adapt content based on product/program updates and user feedback.
- Represent MBZUAI at industry conferences and events, showcasing the institution’s cutting-edge HPC and deep learning capabilities and establishing MBZUAI as a global leader in AI research and innovation.
- Perform all other duties as reasonably directed by the line manager that are commensurate with these functional objectives.
- Validate CUDA kernel outputs and gradients against reference implementations, and benchmark representative shapes, dtypes, and model workloads.
- Strong C++ skills and hands-on CUDA kernel development and optimization for deep-learning workloads.
- Understanding of GPU memory hierarchy, warp/block execution, and compute-memory trade-offs, with demonstrated profiling-driven optimization.
- Strong Python skills and experience integrating kernels with PyTorch or an equivalent framework, including numerical and gradient validation where needed.
- Experience with Triton, CUTLASS, or PTX/SASS analysis.
- Experience with multi-node distributed training or inference systems.
- Experience validating mixed-precision computations, such as BF16 or FP8.
Location & Eligibility
Listing Details
- Posted
- October 5, 2026
- First seen
- October 5, 2026
- Last seen
- October 5, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 79%
- Scored at
- October 5, 2026
Signal breakdown
Similar Machine Learning Engineer jobs
View all →Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.