CODA: Letting LLMs and Novices Write Speed-of-Light GPU Kernels
An open-source project from MIT, Princeton and others rewrites the scattered computations in Transformer training into the GEMM-Epilogue pattern, speeding up backpropagation by 1.6-1.8x


CODA: Letting LLMs and Novices Write Speed-of-Light GPU Kernels
An open-source project from MIT, Princeton and others rewrites the scattered computations in Transformer training into the GEMM-Epilogue pattern, speeding up backpropagation by 1.6-1.8x
GPU kernel optimization has always been a field with an extremely high barrier to entry, typically requiring senior CUDA engineers to hand-tune everything. CODA, a project from researchers at MIT, Princeton, Together AI, and Meta, tries to change that -- with a set of programming abstractions that let LLMs, and even novices, write high-performance GPU kernels for Transformers.
Tri Dao, lead author of FlashAttention, put it bluntly while resharing the news: "LLMs and novices can just go ahead and write speed-of-light kernels."
Paper: arxiv.org/abs/2605.19269 Code: github.com/HanGuo97/coda-kernels
The Problem: The "Laziness Tax" of Training Large Models
Train a LLaMA-3-style 1B-parameter model on a single H100, and matrix multiplication (GEMM) and attention do dominate the compute. But a profiler reveals a quiet gang of "time killers": RMSNorm, SwiGLU, RoPE, residual adds, cross-layer reductions.

Individually, these operations don't compute much, yet they constantly shuttle large intermediate tensors in and out of GPU memory. This is the memory bandwidth bottleneck -- like a master chef who has to haul ingredients back and forth from a distant warehouse for every single dish.
As low-precision formats like FP8 and FP4 make matrix math faster and faster, the relative cost of these "hauling" operations keeps rising. PyTorch expresses a Transformer as a sequence of operators, and the boundaries between operators are precisely what blocks cross-operator fusion optimizations.
The Core Insight: Treasure Hidden in the "Epilogue"
A high-performance matrix multiplication (GEMM) kernel on the GPU has two parts:
- Mainloop: the core blocked matrix multiply-accumulate computation
- Epilogue: the finishing touches before results are written back to memory (adding bias, type conversion, and so on)

The point of the epilogue: the matrix multiplication's output is still "alive" in on-chip registers and hasn't yet landed in global memory. This is a brief golden window -- squeeze more computation into this moment and you save a full round trip of writing to memory and reading it back.
CODA's core insight: the memory-intensive operations in a Transformer can be algebraically reparameterized and stuffed into the "epilogue" window for execution.
Take the most common GEMM-RMSNorm-GEMM pattern: the row scaling factor r in RMS normalization commutes with the matrix multiplication that follows, so applying r can be deferred to the second GEMM's epilogue. That way the full RMSNorm computation simply disappears.

Five Kinds of Composable "Building Blocks"
CODA is not one specific fused kernel but a set of programming abstractions. It pins down the expert-optimized GEMM mainloop, then exposes five kinds of composable primitives at the epilogue position:
| Primitive type | Purpose |
|---|---|
| Elementwise transforms | residual adds, activation functions, RoPE |
| Vector load and store | broadcasting RMSNorm weights |
| Blocked matrix load and store | saving intermediate activations for backpropagation |
| Block reductions | local root-mean-square, blocked log-sum-exp |
| Stateful transforms | the max and sum-exp statistics needed for online normalization |
With these five kinds of building blocks, nearly every operation in a standard Transformer's forward and backward passes — apart from attention — can be covered.
Can LLMs Write GPU Kernels?
The paper evaluated two implementation modes:
- CODA (LLM): generated by Claude Code, with the researchers providing primitive documentation, examples, and implementation logs; the AI wrote the bulk of the code under light human supervision
- CODA (Human): written independently by human programmers using the same reparameterization approach
The LLM-generated kernels are on par with the hand-written versions on most benchmarks, and even slightly ahead under some configurations. In GPU kernel optimization — historically a field with an extremely high barrier — that is a rare conclusion.
Experimental Results
The benchmarks picked demanding opponents: cuBLAS + torch.compile, Liger Kernel, and FlashInfer.

Key numbers:
- GEMM-RMSNorm-GEMM: beats the cuBLAS + PyTorch baseline at hidden dimensions across three model scales — 1B, 7B, and 70B
- Backpropagation gains stand out: the GEMM-Residual-PartialRMS-GEMM backward kernel is 1.6 to 1.8x faster than the baseline
- SwiGLU backward: roughly 1.4 to 1.6x improvement
- The gap between LLM and human implementations is tiny — nearly identical in the backward direction

Who It's For
- Teams training large models who want to squeeze out the last drop of GPU performance
- GPU kernel optimization newcomers who want to write high-performance code through high-level abstractions
- Researchers who want to test the proposition "can AI write GPU kernels"
References:
Related articles

XtraGPT: An AI Paper Revision Tool Built on Full-Paper Context
An ACL 2026 paper that uses 20 academic writing criteria and full-paper context modeling to turn AI paper revision from generic polishing into controlled, targeted edits

Reviving Fable 5 with One Line of Code: System Prompt Injection in Practice, and the Principle Behind It
Using the leaked Fable 5 system-level prompt and the --system-prompt-file flag to inject Fable 5's 'personality blueprint' into Opus 4.8 and get closely matching output

FuseSearch: How a 4-Billion-Parameter Small Model Outguns Commercial LLMs at Code Localization
Ant Group's ACL 2026 work FuseSearch-4B uses an adaptive parallel search strategy to match Claude Haiku 4.5 on code localization — 93.6% faster with 68.9% fewer tokens

GLM-5.2's Million-Token Context, Tested: An 85-Page World Cup Preview in One Click
Zhipu's GLM-5.2 supports a 1 million token context window and will be open-sourced under MIT next week; in testing it produced an 85-page World Cup preview deck, with multi-agent parallelism that beat expectations

OpenRouter Fusion Tutorial: A Three-Model Combo That Matches Fable 5 at Half the Cost
OpenRouter's Fusion multi-model blending scheme: a Kimi K2.6 + DeepSeek V4 Pro + Gemini 3 Flash combo matches Fable 5 on the DRACO benchmark at only 50% of the cost

VeraRetouch: A 0.6B-Parameter AI Retouching Model and a New Path to On-Device Mobile Deployment
vivo and Zhejiang University release VeraRetouch, a lightweight retouching framework built on a 0.6B vision-language model that supports auto, style, and param retouching, processing on an iPhone in about 13 seconds