Build1 publisher3 min readPublished Updated
1,500 submissions in 14 days: what a 12th-place GPU kernel says about agent loops
Sankalp placed 12th of 183 on GPU Mode's B200 QR benchmark using OpenAI Codex. The transferable part is the checker, not the ranking.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Sankalp finished 12th among 183 entrants in GPU Mode's qr_v2 competition with a time of 1,804.779 microseconds on an NVIDIA B200.
- He reached that result after more than 1,500 submissions over 14 days using OpenAI Codex.
- The project was part of GPU Mode's Linear Algebra Kernels in the Age of Research series and its qr_v2 competition.
- Sankalp's blog reports a 232x improvement over an approximate PyTorch baseline of 419,000 microseconds; that comparison is specific to this benchmark workload and was not independently reproduced in the reports.
- The competition ended on June 29, about six and a half weeks before the article.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A developer named Sankalp finished 12th out of 183 entrants in GPU Mode's qr_v2 kernel competition, with a geometric-mean time of 1,804.779 microseconds on an NVIDIA B200 [1]. He got there with more than 1,500 leaderboard submissions over 14 days driven by OpenAI Codex [2], and that submission count, not the placing, is the part worth studying.
The contest, run by GPU Mode with Core Automation as part its Linear Algebra Kernels in the Age of Research series, asked for batched square compact-Householder QR factorization [3][10]. Submissions had to take FP32 CUDA matrices and return the same compact form as PyTorch's torch.geqrf: an H matrix holding upper-triangular R plus stored Householder vectors, and a tau vector of reflector coefficients [11]. The checker rebuilt Q, tested orthogonality and residual error, then ranked only correct entries by geometric-mean runtime across shapes and input conditions [12], with matrices up to 4,096 by 4,096 and deliberately awkward conditioning [13]. Internal use of lower precision was allowed as long as the output still passed FP32-style checks [14].
That is the whole story in one paragraph. The objective was machine-checkable, cheap to evaluate, and adversarial about correctness, so an agent could not win by sounding plausible. GPU Mode's popcorn command-line tool let Codex test, benchmark and submit candidates itself [15], and shape-level timings plus profiling told it which specific change helped or hurt [16]. More than 1,500 submissions in 14 days works out to an average above 107 per day [28]; that cadence only means anything because every one of them was scored.
The scaffolding was ordinary file discipline. Sankalp kept an AGENTS.md with operating instructions, the problem statement, experiment records and timestamped submission logs [17], so later Codex sessions could read what had already failed rather than rediscover it [18]. He set numerical targets, let some runs go overnight, and checked in every two or three hours to ask what had changed and which bottleneck was being chased [19]. The working directory ended up with 560 named submission variants, 119 Modal B200 probe and comparison scripts, and 68 per-experiment documents [20]. Modal supplied GPU credits, according to his account [21].
The architecture came from him. He used Claude and course material to learn Householder QR [22], then chose a blocked design with a trailing WY update [23]. The real constraint is serial dependency: a conventional implementation walks columns in order, each reflector depending on the previous matrix, which strands work in matrix-vector operations while the tensor cores idle [24]. Blocking confines the serial part to a narrow panel and turns the trailing update into matrix multiplication [25]. He reported about 5,000 microseconds on the heavily weighted 512 by 512 case within a day of switching [26]. Everything after that was stack-wide grinding: custom Triton panels, grouped WY updates, CUDA graph replay, fused layout assembly, fixed-shape specialization [27].
Context on the ceiling: his entry was roughly 48% slower than the winning 1,220.774 microseconds, a gap of 584.005 microseconds [7][30], and he entered with about a year of self-taught kernel work, mostly Triton, and no professional experience in the field [8]. His blog also claims a 232x improvement over an approximate 419,000-microsecond PyTorch baseline [5]; that figure is workload-specific and was not independently reproduced in the reports [5]. The competition closed on June 29 [6].
What to watch: whether the same loop produces anything when the scoring function is slower, noisier or written by the same person being graded. Kernel leaderboards are the easy case. The interesting test is a domain where the check costs hours instead of seconds.