CODA: Letting LLMs and Novices Write Speed-of-Light GPU Kernels

·Toolin Editorial Team

An open-source project from MIT, Princeton and others rewrites the scattered computations in Transformer training into the GEMM-Epilogue pattern, speeding up backpropagation by 1.6-1.8x

CODA: Letting LLMs and Novices Write Speed-of-Light GPU Kernels

GPU kernel optimization has always been a field with an extremely high barrier to entry, typically requiring senior CUDA engineers to hand-tune everything. CODA, a project from researchers at MIT, Princeton, Together AI, and Meta, tries to change that -- with a set of programming abstractions that let LLMs, and even novices, write high-performance GPU kernels for Transformers.

Tri Dao, lead author of FlashAttention, put it bluntly while resharing the news: "LLMs and novices can just go ahead and write speed-of-light kernels."

Paper: arxiv.org/abs/2605.19269 Code: github.com/HanGuo97/coda-kernels

The Problem: The "Laziness Tax" of Training Large Models

Train a LLaMA-3-style 1B-parameter model on a single H100, and matrix multiplication (GEMM) and attention do dominate the compute. But a profiler reveals a quiet gang of "time killers": RMSNorm, SwiGLU, RoPE, residual adds, cross-layer reductions.

Training time breakdown

Individually, these operations don't compute much, yet they constantly shuttle large intermediate tensors in and out of GPU memory. This is the memory bandwidth bottleneck -- like a master chef who has to haul ingredients back and forth from a distant warehouse for every single dish.

As low-precision formats like FP8 and FP4 make matrix math faster and faster, the relative cost of these "hauling" operations keeps rising. PyTorch expresses a Transformer as a sequence of operators, and the boundaries between operators are precisely what blocks cross-operator fusion optimizations.

The Core Insight: Treasure Hidden in the "Epilogue"

A high-performance matrix multiplication (GEMM) kernel on the GPU has two parts:

  • Mainloop: the core blocked matrix multiply-accumulate computation
  • Epilogue: the finishing touches before results are written back to memory (adding bias, type conversion, and so on)

GEMM-Epilogue structure

The point of the epilogue: the matrix multiplication's output is still "alive" in on-chip registers and hasn't yet landed in global memory. This is a brief golden window -- squeeze more computation into this moment and you save a full round trip of writing to memory and reading it back.

CODA's core insight: the memory-intensive operations in a Transformer can be algebraically reparameterized and stuffed into the "epilogue" window for execution.

Take the most common GEMM-RMSNorm-GEMM pattern: the row scaling factor r in RMS normalization commutes with the matrix multiplication that follows, so applying r can be deferred to the second GEMM's epilogue. That way the full RMSNorm computation simply disappears.

Computation fusion diagram

Five Kinds of Composable "Building Blocks"

CODA is not one specific fused kernel but a set of programming abstractions. It pins down the expert-optimized GEMM mainloop, then exposes five kinds of composable primitives at the epilogue position:

Primitive typePurpose
Elementwise transformsresidual adds, activation functions, RoPE
Vector load and storebroadcasting RMSNorm weights
Blocked matrix load and storesaving intermediate activations for backpropagation
Block reductionslocal root-mean-square, blocked log-sum-exp
Stateful transformsthe max and sum-exp statistics needed for online normalization

With these five kinds of building blocks, nearly every operation in a standard Transformer's forward and backward passes — apart from attention — can be covered.

Can LLMs Write GPU Kernels?

The paper evaluated two implementation modes:

  • CODA (LLM): generated by Claude Code, with the researchers providing primitive documentation, examples, and implementation logs; the AI wrote the bulk of the code under light human supervision
  • CODA (Human): written independently by human programmers using the same reparameterization approach

The LLM-generated kernels are on par with the hand-written versions on most benchmarks, and even slightly ahead under some configurations. In GPU kernel optimization — historically a field with an extremely high barrier — that is a rare conclusion.

Experimental Results

The benchmarks picked demanding opponents: cuBLAS + torch.compile, Liger Kernel, and FlashInfer.

Performance comparison data

Key numbers:

  • GEMM-RMSNorm-GEMM: beats the cuBLAS + PyTorch baseline at hidden dimensions across three model scales — 1B, 7B, and 70B
  • Backpropagation gains stand out: the GEMM-Residual-PartialRMS-GEMM backward kernel is 1.6 to 1.8x faster than the baseline
  • SwiGLU backward: roughly 1.4 to 1.6x improvement
  • The gap between LLM and human implementations is tiny — nearly identical in the backward direction

Backpropagation performance

Who It's For

  • Teams training large models who want to squeeze out the last drop of GPU performance
  • GPU kernel optimization newcomers who want to write high-performance code through high-level abstractions
  • Researchers who want to test the proposition "can AI write GPU kernels"

References:

Related articles

XtraGPT: An AI Paper Revision Tool Built on Full-Paper Context
AI Products

XtraGPT: An AI Paper Revision Tool Built on Full-Paper Context

An ACL 2026 paper that uses 20 academic writing criteria and full-paper context modeling to turn AI paper revision from generic polishing into controlled, targeted edits

Toolin Editorial Team
Reviving Fable 5 with One Line of Code: System Prompt Injection in Practice, and the Principle Behind It
AI Tutorials

Reviving Fable 5 with One Line of Code: System Prompt Injection in Practice, and the Principle Behind It

Using the leaked Fable 5 system-level prompt and the --system-prompt-file flag to inject Fable 5's 'personality blueprint' into Opus 4.8 and get closely matching output

Toolin Editorial Team
FuseSearch: How a 4-Billion-Parameter Small Model Outguns Commercial LLMs at Code Localization
AI Products

FuseSearch: How a 4-Billion-Parameter Small Model Outguns Commercial LLMs at Code Localization

Ant Group's ACL 2026 work FuseSearch-4B uses an adaptive parallel search strategy to match Claude Haiku 4.5 on code localization — 93.6% faster with 68.9% fewer tokens

Toolin Editorial Team
GLM-5.2's Million-Token Context, Tested: An 85-Page World Cup Preview in One Click
AI Tutorials

GLM-5.2's Million-Token Context, Tested: An 85-Page World Cup Preview in One Click

Zhipu's GLM-5.2 supports a 1 million token context window and will be open-sourced under MIT next week; in testing it produced an 85-page World Cup preview deck, with multi-agent parallelism that beat expectations

Toolin Editorial Team
OpenRouter Fusion Tutorial: A Three-Model Combo That Matches Fable 5 at Half the Cost
AI Tutorials

OpenRouter Fusion Tutorial: A Three-Model Combo That Matches Fable 5 at Half the Cost

OpenRouter's Fusion multi-model blending scheme: a Kimi K2.6 + DeepSeek V4 Pro + Gemini 3 Flash combo matches Fable 5 on the DRACO benchmark at only 50% of the cost

Toolin Editorial Team
VeraRetouch: A 0.6B-Parameter AI Retouching Model and a New Path to On-Device Mobile Deployment
AI Products

VeraRetouch: A 0.6B-Parameter AI Retouching Model and a New Path to On-Device Mobile Deployment

vivo and Zhejiang University release VeraRetouch, a lightweight retouching framework built on a 0.6B vision-language model that supports auto, style, and param retouching, processing on an iPhone in about 13 seconds

Toolin Editorial Team