alex zhang (@a1zhang) on X

2 min read Original article ↗

Post

Post

  • user avatar

    Excited to announce the $100K competition hosted by

    @AMD

    and

    @GPU_MODE

    as part of the first round of our GPU kernel leaderboard!😁 The theme is writing LLM inference kernels🍿on AMD MI300s, provided for FREE through the

    @GPU_MODE

    Discord Registration / competition details👇🧵:

  • user avatar

    You (+ anyone around the 🌎) can participate for FREE by signing up here: datamonsters.com/amd-developer-… The competition will begin on April 15th! The format will consist of 3 kernels that are core to DeepSeek’s LLM inference 🐋: FP8 GEMM, Multi-head Latent Attention, & Fused MOE!

    user avatar

    Competitors will target the AMD MI300 in primarily Triton, but other DSLs and languages that target AMD hardware are permitted! Participants can form up to teams of 3️⃣ when competing and prizes will be awarded based on team rankings averaged over the three kernels.

    user avatar

    The challenge will be hosted on the

    @GPU_MODE

    Discord, where we previously hosted a practice round based on classic PMPP textbook problems 🧩! Unlike before, we will release the reference kernels and problem specifications one-by-one (every 2-3 weeks).

    user avatar

    Special thanks to

    @AMD@indianspeedster

    and the other amazing

    @GPU_MODE

    Project Popcorn core devs,

    @m_sirovatka

    ,

    @marksaroufim

    , ngc92 (Erik S.), and Ben Horowitz for making this competition possible. We’re very excited to accelerate AI research by building on kernels, and we’re