import matmul
# Verify your IR computes A @ B correctly and return its read-cost.
cost = matmul.score_1x1("1,2;mul 3,1,2;3") # 5
ir = matmul.generate_baseline_4x4() # naive triple loop, 4×4
cost = matmul.score_4x4(ir)
ir = matmul.generate_baseline_16x16() # naive triple loop, 16×16
cost = matmul.score_16x16(ir)
ir = matmul.generate_tiled_16x16() # 4×4 scratchpad-cached tiles
cost = matmul.score_16x16(ir)
Correctness is checked symbolically, so an IR must compute A @ B for
arbitrary inputs, not just one sample pair. Intermediates above degree two are
rejected, which admits the usual bilinear matmul algorithms.
| Date | Cost | Submission | Contributors | Description |
|---|---|---|---|---|
| 2026-04-29 | 1,316 | ir, report | @yaroslavvb | generate_baseline_4x4 (naive) |
| 2026-04-30 | 800 | ir, report | @sjbaebae | generate_outer_product_4x4 (size-1 sA) |
| 2026-09-01 | 689 | ir, report, py | @jurajselep | row-wise lifetime fusion + dead-A output reuse + exact tier allocation |
| 2026-09-01 | 683 | ir, report, py | @npow | HKKK + one deferred output add + exact fixed-schedule allocation |
| 2026-09-01 | 681 | ir, report, py | @jurajselep | hybrid row-0 schedule + JIT A staging + interleaved final adds |
| 2026-09-04 | 675 | ir, report, py | @jurajselep | selective A/B lifetime splits + dependency-safe scheduling + exact lifetime coloring |
| 2026-09-20 | 667 | ir, report, py | @jurajselep | contraction-order search + input captures + storage allocation ★ best |
| Date | Cost | Submission | Contributors | Description |
|---|---|---|---|---|
| 2026-04-29 | 340,704 | ir, report | @yaroslavvb | generate_baseline_16x16 (naive) |
| 2026-05-08 | 237,456 | ir, report | @yaroslavvb | generate_recursive_16x16 (1×1-leaf D&C, Z-order) |
| 2026-04-29 | 133,783 | ir, report | @yaroslavvb | generate_tiled_16x16 (4×4 tiles) |
| 2026-04-30 | 110,743 | ir, report | @SethTS | generate_tiled_16x16_opt1 (tmp@1) |
| 2026-04-30 | 80,217 | ir, report | @sjbaebae | generate_hierarchical_16x16 (asym. reload) |
| 2026-04-30 | 73,602 | ir, report | @adotzh | sA-cache + sB scratchpad (rank2) |
| 2026-05-01 | 72,642 | ir, report | @sjbaebae | + redirect last-mul to addr 1 |
| 2026-05-01 | 71,724 | ir, report | @sjbaebae | + last-super-block outputs in sC |
| 2026-05-01 | 70,053 | ir, report | @sjbaebae | + dead-input output reuse + B packing |
| 2026-05-06 | 69,697 | ir, report | @yaroslavvb | C↔A address aliasing + final-add fusion |
| 2026-05-05 | 68,452 | ir, report | @zh4ngx | + column-major order + fused final copy-out |
| 2026-05-13 | 68,390 | ir, report | @cosminscn | + liveness order + output-read-aware packing + five-output scratch tail |
| 2026-05-13 | 67,834 | ir, report | @cosminscn | + live-B evacuation + output deferral + A staging + value-lifetime coloring |
| 2026-05-14 | 67,821 | ir, report | @cosminscn | + live-B evacuation + output deferral + tiny A-staging mask + staged-reload endpoint lift + value-lifetime coloring |
| 2026-05-08 | 66,707 | ir, report, py | @sjbaebae | weighted-lifetime pressure search + copy elimination |
| 2026-05-25 | 66,633 | ir, report, py | @cosminscn | macro B-staging + row-7 later-panel prestaging from addr 1 |
| 2026-05-25 | 66,524 | ir, report, py | @cosminscn | late B-block cheap capture from addr 1 + value-lifetime coloring |
| 2026-05-26 | 66,400 | ir, report, py | @cosminscn | late copy-schedule motif bundle + value-lifetime coloring |
| 2026-05-28 | 66,300 | ir, report, py | @cosminscn | Claude-assisted simulated annealing over a leaderboard physical-address IR |
| 2026-08-29 | 66,199 | ir, report, py | @sigkillme0 | dependency-safe rescheduling + exact tier allocation |
| 2026-08-30 | 66,178 | ir, report, py | @sigkillme0 | exact LP-optimal address assignment (provably optimal for this operation order) |
| 2026-09-04 | 65,084 | ir, report, py | @jurajselep | newest-first snake passes + JIT A staging + local exact tier allocation |
| 2026-09-04 | 64,431 | ir, report, py | @cosminscn | asymmetric panel schedule + persistent B captures + dependency-safe order search |
| 2026-09-07 | 64,074 | ir, report, py | @SecurityQQ | temporary input captures + cheapest surviving replica reads + redundant-copy deletion |
| 2026-09-07 | 63,819 | ir, report, py | @SecurityQQ | 6+10 asymmetric panels + dependency-safe scheduling + input captures and address allocation (frozen-artifact verifier) |
| 2026-09-15 | 63,639 | ir, report, py | @jurajselep | input copy-chain and neutral-plan optimization + whole-program allocation with an exact rational certificate |
| 2026-09-15 | 63,354 | ir, py | @jurajselep | exact block reordering + joint reduction-tree and storage allocation |
| 2026-09-18 | 63,350 | ir, report, py | @jurajselep | joint reduction-tree/storage repair + redundant-copy elimination |
| 2026-09-18 | 63,290 | ir, report, py | @jurajselep | structural schedule search + six local sum reassociations (frozen-artifact verifier) ★ best |
access_distance — read-distance histograms for the plotted submission set.