sutro-problems

View this project on GitHub ↗

Matmul

naive_4x4_matmul

API

import matmul

# Verify your IR computes A @ B correctly and return its read-cost.
cost = matmul.score_1x1("1,2;mul 3,1,2;3")    # 5

ir = matmul.generate_baseline_4x4()      # naive triple loop, 4×4
cost = matmul.score_4x4(ir)

ir = matmul.generate_baseline_16x16()    # naive triple loop, 16×16
cost = matmul.score_16x16(ir)

ir = matmul.generate_tiled_16x16()       # 4×4 scratchpad-cached tiles
cost = matmul.score_16x16(ir)

Correctness is checked symbolically, so an IR must compute A @ B for arbitrary inputs, not just one sample pair. Intermediates above degree two are rejected, which admits the usual bilinear matmul algorithms.

4×4 Record History

Date Cost Submission Contributors Description
2026-04-29 1,316 ir, report @yaroslavvb generate_baseline_4x4 (naive)
2026-04-30 800 ir, report @sjbaebae generate_outer_product_4x4 (size-1 sA)
2026-09-01 689 ir, report, py @jurajselep row-wise lifetime fusion + dead-A output reuse + exact tier allocation
2026-09-01 683 ir, report, py @npow HKKK + one deferred output add + exact fixed-schedule allocation
2026-09-01 681 ir, report, py @jurajselep hybrid row-0 schedule + JIT A staging + interleaved final adds
2026-09-04 675 ir, report, py @jurajselep selective A/B lifetime splits + dependency-safe scheduling + exact lifetime coloring
2026-09-20 667 ir, report, py @jurajselep contraction-order search + input captures + storage allocation ★ best

16×16 Record History

Date Cost Submission Contributors Description
2026-04-29 340,704 ir, report @yaroslavvb generate_baseline_16x16 (naive)
2026-05-08 237,456 ir, report @yaroslavvb generate_recursive_16x16 (1×1-leaf D&C, Z-order)
2026-04-29 133,783 ir, report @yaroslavvb generate_tiled_16x16 (4×4 tiles)
2026-04-30 110,743 ir, report @SethTS generate_tiled_16x16_opt1 (tmp@1)
2026-04-30 80,217 ir, report @sjbaebae generate_hierarchical_16x16 (asym. reload)
2026-04-30 73,602 ir, report @adotzh sA-cache + sB scratchpad (rank2)
2026-05-01 72,642 ir, report @sjbaebae + redirect last-mul to addr 1
2026-05-01 71,724 ir, report @sjbaebae + last-super-block outputs in sC
2026-05-01 70,053 ir, report @sjbaebae + dead-input output reuse + B packing
2026-05-06 69,697 ir, report @yaroslavvb C↔A address aliasing + final-add fusion
2026-05-05 68,452 ir, report @zh4ngx + column-major order + fused final copy-out
2026-05-13 68,390 ir, report @cosminscn + liveness order + output-read-aware packing + five-output scratch tail
2026-05-13 67,834 ir, report @cosminscn + live-B evacuation + output deferral + A staging + value-lifetime coloring
2026-05-14 67,821 ir, report @cosminscn + live-B evacuation + output deferral + tiny A-staging mask + staged-reload endpoint lift + value-lifetime coloring
2026-05-08 66,707 ir, report, py @sjbaebae weighted-lifetime pressure search + copy elimination
2026-05-25 66,633 ir, report, py @cosminscn macro B-staging + row-7 later-panel prestaging from addr 1
2026-05-25 66,524 ir, report, py @cosminscn late B-block cheap capture from addr 1 + value-lifetime coloring
2026-05-26 66,400 ir, report, py @cosminscn late copy-schedule motif bundle + value-lifetime coloring
2026-05-28 66,300 ir, report, py @cosminscn Claude-assisted simulated annealing over a leaderboard physical-address IR
2026-08-29 66,199 ir, report, py @sigkillme0 dependency-safe rescheduling + exact tier allocation
2026-08-30 66,178 ir, report, py @sigkillme0 exact LP-optimal address assignment (provably optimal for this operation order)
2026-09-04 65,084 ir, report, py @jurajselep newest-first snake passes + JIT A staging + local exact tier allocation
2026-09-04 64,431 ir, report, py @cosminscn asymmetric panel schedule + persistent B captures + dependency-safe order search
2026-09-07 64,074 ir, report, py @SecurityQQ temporary input captures + cheapest surviving replica reads + redundant-copy deletion
2026-09-07 63,819 ir, report, py @SecurityQQ 6+10 asymmetric panels + dependency-safe scheduling + input captures and address allocation (frozen-artifact verifier)
2026-09-15 63,639 ir, report, py @jurajselep input copy-chain and neutral-plan optimization + whole-program allocation with an exact rational certificate
2026-09-15 63,354 ir, py @jurajselep exact block reordering + joint reduction-tree and storage allocation
2026-09-18 63,350 ir, report, py @jurajselep joint reduction-tree/storage repair + redundant-copy elimination
2026-09-18 63,290 ir, report, py @jurajselep structural schedule search + six local sum reassociations (frozen-artifact verifier) ★ best

access_distance — read-distance histograms for the plotted submission set.