GitHub ↗

Grid-model redesign: decision dossier

Every source bearing on the four open decisions — the instruction set, the grid size, the physical dimensions, and what competitors submit — with what is decided, what is proposed, what contradicts what, and a 58-item fork index. Companion to The Grid VM: competition design, The expressivity–scorability ladder and Multiprocessors on the Grid VM. September 8, 2026.

Assembled 2026-09-08 (week 37). Everything bearing on the four decisions you named: the instruction set, the grid size, the physical dimensions, and what competitors actually submit. This is a reference, not a proposal — it says what is written down, where, and what is still open. Nothing here is a recommendation unless a source made one, in which case the source is named.

Do not push this file as-is. It links private ChatGPT/Gemini/Claude sessions and internal Google Docs, and the repo root is served publicly by GitHub Pages. It is deliberately left uncommitted.

Read-order if you only have ten minutes: §0 (state of each decision) → §9 (the complete fork index — 58 choices with status) → §7 (contradictions) → §8 (gaps). If you have thirty: add §6½ (the forks nobody has named yet) — that is where the answer to "what am I forgetting?" mostly lives.


0. State of each decision

# Decision Status Where the current answer is written
1 ISA v0/v3 live and frozen per-problem; the redesign ISA (13 ops + recv/send + mode) is proposed, not merged simplified-dally-model/instruction-sets/; docs/grid-vm-competition-design.html §4
2 Grid size Three MNIST tiers decided (3×3/600, 9×9/6k, 28×28/60k). Cell count/address-space caps proposed (1 MB / 16 MB / 256 MB per tier). Pitch-to-capacity mapping open internal log 03sep26 + 06sep26; docs/grid-vm-competition-design.html §8
3 Physical dimensions Decided as a mnemonic triple: 1 µm pitch, 1 fJ/byte·µm, c/160. Their derivation from Dally is partly unsourced and the per-mm constants contradict each other by 3–9× Sutro meeting #30 agenda; RFC #48 §8.2; docs/grid-vm-multiprocessor.html §2
4 Submission format Two-track, proposed and coherent: trace for episodic/tier-1, schedule language for at-scale/tiers 2–3. Awaiting your decision — Andy has built it and you have not reviewed PR #48 PR #48; docs/expressivity-scorability-ladder.html §2

The single most consequential thing in the dossier, and it is not on your list of four:

"the one decision that actually shapes the competition is not in the language at all, but in where the episodes come from." — docs/grid-vm-competition-design.html, §11 closing sentence

That is decision D1 (embedded weights). If the training pool is public, a submitter pretrains offline and ships int8 weights as immediates — ~100 B for a tier-1 linear head, ~80 KB for a tier-3 MLP at ~98%. Embedding costs a ~10⁻⁵ tax and skips roughly 15/16 of tier-3 energy, so the "learning" axis becomes dead code and the benchmark measures inference circuits. See §6.3.


1. Where everything lives

1.1 The 106-commit gap — read this first

Your local working tree at ~/Library/CloudStorage/Dropbox/git0/sutro-problems is 106 commits behind origin/main. The three design reports that settle most of this exist only on the remote and are not in your checkout:

cd ~/Library/CloudStorage/Dropbox/git0/sutro-problems && git log --oneline HEAD..origin/main | wc -l
Doc Landed What it settles
docs/expressivity-scorability-ladder.html (00c8cb7) Sep 4 rungs (a)–(f); where the scoring cliff is; the fetch fork
docs/grid-vm-competition-design.html (4f4d3e2) Sep 4 the whole machine: language, scoring function, adjudication, judge cost, targets, decisions D1–D6
docs/grid-vm-multiprocessor.html (5706824) Sep 7 four multiprocessor scenarios + open decisions P1–P8; also silently corrects the metric and the c/160 figure
docs/dataflow-io-order.html Sep 3 ports, determinacy, closed/open division, the resident-vs-restream break-even
docs/rfc-scheduled-dally-language.md branch rfc-scheduled-dally only (PR #48) SDL: grammar, closed-form scoring, sandbox

Live at: - https://cybertronai.github.io/sutro-problems/docs/grid-vm-competition-design.html - https://cybertronai.github.io/sutro-problems/docs/expressivity-scorability-ladder.html - https://cybertronai.github.io/sutro-problems/docs/grid-vm-multiprocessor.html - https://cybertronai.github.io/sutro-problems/docs/dataflow-io-order.html - https://cybertronai.github.io/sutro-problems/docs/spatial-model-analysis.html - https://cybertronai.github.io/sutro-problems/docs/

The only origin-only directory is a100-grid-energy-report/. dally-eval is an external repo (https://github.com/cybertronai/dally-eval); this repo carries a 74-line shim dally_eval.py. There is no scoring-at-scale/, no mnist/, no grid/ on any of the 22 remote branches, despite references to them.

1.2 Repos

Repo Role
cybertronai/sutro-problems the benchmark suite, leaderboards, design reports
cybertronai/simplified-dally-model the cost model + the versioned ISA specs (v0–v3)
cybertronai/dally-eval the fast Rust scorer (wired in by PR #44)
~/Library/CloudStorage/Dropbox/git0/spatial-computer/ earlier long-form: dally-spatial-report.html, pram-tutorial.html, slides, paper.pdf
~/Library/CloudStorage/Dropbox/git0/SutroYaro/ Yad's repo; telegram.db (Feb 9 – Mar 28 only), mirrored group docs under docs/google-docs/, docs/catchups/
~/Library/CloudStorage/Dropbox/git0/sutro/ sutro.org, docs/findings/, docs/research/

1.3 Google Docs

Full text of 35 of these is mirrored nightly to ~/drive/gdocs-archive/<docId>/<date>.txt (with meta.json), which is faster and more reliable than opening them.

Doc id
Sutro internal Log (week 36 onward) 1IjNHiv…
Sutro Group: top level 1B9867EN…
sutro meeting #30 — the authoritative "design choices" list 11Xj14ug…
sutro #29 1RYSonuC…
MNIST on grid model? 1uDZ1Nh…
Numbers to know 1vUbTxn…
sutro fJ report — how fJ were chosen 1ZMx22vh…
Bjarke Roune, Designing AI Chip Software and Hardware 1EUVW1tB…
sutro group challenge #1: sparse parity 16eeltCa…
Yaroslav's sutro planning sprint #1 1oSTIM0h…
08sep26 — interlude show 1RvwoLX…

1.4 Chat

1.5 Dictated working notes

~/My Drive/hiq-transcribed/recording-<YYYYMMDD>-<HHMMSS>-<engine>.txt, engines soniox / scribe / parakeet / groq (prefer scribe and soniox; groq hallucinates). 392 clips in the Aug 26 – Sep 8 window. Signal is concentrated in ~20 clips; the densest by far is 20260903-180729, which you titled the design description.

1.6 Meeting cadence and people

Mondays 18:00, South Park Commons, 380 Brannan (meeting #29's doc says 18:15). #29 was Aug 31, #30 was Sep 7. Mission line as of Sep 8: "huge amount of energy waste due to legacy learning algorithms. Address this in a maximally open way."

Contributors on the current leaderboards: @zh4ngx (Andy Zhang), @cosminscn, @jurajselep, @npow, @b0nce, @sigkillme0, @SecurityQQ, @sjbaebae, @SethTS, @adotzh, @ab-10.


2. Decision 1 — the instruction set architecture

2.1 What exists today

Four versioned specs, each strictly extending the last (instruction-sets/):

Version Ops added Total Used by
v0 add, sub, mul, copy 4 matmul (4×4, 16×16)
v1 + and, or, not, xor 8 —
v2 + set (integer immediate; no read, so free) 9 —
v3 + div, cmp, select, abs 13 sparse-parity, symmetry

Three-address code, LLVM-flavoured: %d = add %a, %b ⇒ add d,a,b. Two-operand short form add d,s ≡ add d,d,s. cmp takes a predicate from {eq,ne,lt,le,gt,ge}. v3's stated purpose is Gaussian elimination with partial pivoting as straight-line code.

2.2 What the scorer actually implements — and where it differs from the spec

This matters because the code, not the README, is the ruler.

matmul/matmul.py — v0 only. unknown op: … (v0 supports add/sub/mul/copy). - No cell width at all. Values are unbounded Python ints; correctness is checked symbolically over a polynomial ring, and intermediates of degree > 2 are rejected (which admits the usual bilinear algorithms). There is no wrapping. - No address-space limit. Any integer ≥ 1; ≤ 0 raises. - Cost: _cost(addr) = math.isqrt(addr - 1) + 1 = ⌈√addr⌉, summed over every source operand of every instruction, plus one _cost per declared output address at exit. Writes, arithmetic and input placement are free. - Its own docstring already says: "the cell at linear index addr sits at Manhattan distance ⌈√addr⌉ from the core."

sparse-parity / symmetry — v3, and a genuinely different machine: - 8-bit signed cells, wrapped after every op: _to_signed_8bit(v) = (v & 0xFF) - 0x100 if ≥ 0x80. - Address cap bit_length() ≤ 64; IR length cap 2,000,000 lines (declarations included). - set is free (zero reads); select charges 3 reads.

So "the ISA" is already two incompatible machines. Any redesign has to say which one it is continuous with.

2.3 Your own prior constraints on changing it

These are the only decision-bearing review comments you have ever left in the repo, and they all point the same way:

"Changing the instruction set changes the ruler, so algorithms' performance in the old instruction set aren't directly comparable against algorithms with the new instruction set, I guess maybe I should put the eval in a different repo so agents don't touch it." — PR #1, 2026-04-30

"Each problem has it's own leaderboard. I think evaluation metric should be fixed for each problem. Over time we may discover that something is wrong with current IR, and have v4 or v5 IR of Dally model and start using it for future problems, but for now v3 seems sufficient" — PR #18, 2026-05-14

"Can you modify this to not modify matmul.py? matmul.py is where evaluation code sits." — PR #2, 2026-04-30

"Could you only include the first two files and not matmul/exp_madd.py. The reason is that it re-implements the scoring function, but that changes the metric" — PR #1

Implication for the redesign: the precedent you set is new ISA version → new problems, old problems keep their ruler. That is a clean migration story for MNIST and it costs nothing; it also means the matmul and sparse-parity leaderboards need not move.

Note: PR #58 (MNIST tier 1) both uses set — which v0's matmul.py does not accept — and re-implements static_cost(), which is exactly what you rejected in PR #1.

2.4 The proposed redesign ISA

From docs/grid-vm-competition-design.html §4 and RFC #48 §3:

That last prohibition is load-bearing, not stylistic: given an emptiness test, output stops being a function of input alone (Brock–Ackerman), two runs of the same submission score differently, and the leaderboard is noise.

2.5 What your own dictated notes add

The notes contain no mention of an instruction set, opcode, IR, or assembly in the whole Aug 26 – Sep 8 window. What they do contain are two design criteria:

"the real issue is that you want to be able to do algorithms with small memory footprint, because that's what ultimately it's about" — 20260907-164225

"what if we reinvent all of linear algebra, all of algorithms for a new abstraction? … So the issue with linear algebra is that it … discourages algorithms with small memory footprint. … this was the big surprise of vanilla attention. It's like, well, no abstraction, the cost is invisible." — 20260904-125521

and one governing philosophy for how coarse the model should be:

"the signal cannot go from one [end to the other] in a single clock cycle, so we no longer can pretend that— … Or the danger of modeling is too finely. You can have super accurate models, but if there's many, it splits the community. So one somehow needs a model of computation which is at least conceptually in the right direction, that we can reuse the lessons from, and then fields can specialize." — 20260830-175235

2.6 Open forks on the ISA

2.7 What each fork actually costs in code

The engineering price of every option, in the files that would have to change:

Wanted change What must change
2-D (x,y) addressing matmul.py:_cost, _check_addrs, _parse (operands are int(x)); mask_sparse_parity.py:_cost, parse_addrs, _compile_ir (addr_to_idx sorts scalar ints), _compile_vector; addr_renamer.py (assumes cost is monotone in a scalar address); energy-report/eightk_scaling.py:shell_prefix/shell_sum — its closed form F(x) = q(q+1)(4q−1)/6 + (x−q²)(q+1) is 1-D only
Manhattan lin(a) kept, 2-D merely exposed nothing. PR #61 shows the embedding is an exact isometry, so scores are bit-identical. This is the free option.
Ports / streaming I/O _parse (line 1 = input list, last line = output list) disappears; recv/send need entries in the op tables of both scorers, plus E_in/E_out; test_matmul.py:_assert_submission_invariants hard-asserts exactly 32 inputs / 16 outputs
Wider words _to_signed_8bit / _wrap8 in mask_sparse_parity.py; the int16 batch dtype; the set literal bound (−128..255). Note the matmul side is already unbounded-int, so it needs a width added, not widened
Priced writes both simulators charge reads only — 3 cost += sites in matmul.py, 1 loop in _compile_ir
Instruction-fetch pricing no code exists anywhere
Time / area metric nothing exists. eightk-audit.json states it outright: "No throughput or clock rate is used, so the model produces energy but does not produce time."
An op cap for matmul mask_sparse_parity has OP_CAP; matmul.py has none
Non-bilinear matmul algorithms _poly_mul — the degree-2 restriction blocks them by construction

Two observations worth carrying into the decision:


3. Decision 2 — the grid size

3.1 What is sized today

Problem Instance Caps
sparse-parity n = 32 bits, k = 5 secret positions, m = 18 strings; 32 output cells ≤ 2,000,000 IR lines; address bit_length() ≤ 64; 1,024-instance dev suite, fresh 2,048-instance adjudication suite
matmul 4×4 and 16×16 only no address cap at all
symmetry 6-bit palindrome (8 of 64 patterns) —

3.2 The MNIST tier ladder — decided

Stated twice within a minute on 2026-09-03 and again in the log on 06sep26:

Tier Image Train / test Grid cap (proposed) K episodes Fuel cap
1 3×3 600 / 600 1 MB 50 10⁹
2 9×9 6,000 / 6,000 16 MB 10 10¹¹
3 28×28 60,000 / 60,000 256 MB 3 3×10¹² laptop / 10¹³ CI

"These scales roughly two orders of magnitude" (20260903-153539). QMNIST supplies the extras. Memory stays under ~1.5 GB worst case, plus one shared ~110 MB corpus.

Area is defined as the origin-anchored bounding square of touched cells in µm² — deliberately not touched-cell count, which is gameable by scatter — unioned over mode arms.

3.3 Capacity anchors

The arithmetic a grid-size decision has to match, reverse-engineered from Dally in the internal log, 03sep26 (your own parenthetical "(where is the micron measurement from?)" is still unanswered):

⚠️ Two off-by-ones in this derivation, and the pitch inherits them. - The series is wrong: the log says "1+3+…+31 = 1024", but 1+3+…+31 = 256 (16 terms). The identity you want is Σ(2c−1) for c = 1..32 = 32² = 1024, whose last term is 63: "1+3+…+63". - The bounding box is wrong: a 32-shell upper-half-plane diamond spans x ∈ [−31, 31] and y ∈ [1, 32], i.e. 63 × 32, not 62 × 31. (62 × 31 = 1,922 cells, not 1,024.) - So 1.77 µm is derived from the wrong box. 110/62 = 55/31 = 1.774 exactly — which strongly suggests 110 × 55 µm was itself back-solved from 62 × 31 at 1.774. On the correct box the two axes stop agreeing: 110/63 = 1.746 and 55/32 = 1.719.

The conclusion (32 shells hold 1,024 word-sites) is right; the arithmetic under it is not. Worth fixing before it is quoted, because this is the same 2c−1 shell identity that settles the Manhattan-vs-Chebyshev question in §6.1. - ⇒ 181 × 181 grid of 64 Kb sub-arrays, ≈ 10 mm wide, sub-array ≈ 55 µm wide.

The parallel derivation in the multiprocessor note's calibration thread gets 55.2 µm square-equivalent per sub-array and 1.73 µm per 64-bit word site = 0.61 µm per byte, with the 8 KiB sub-array as 32 half-diamond shells of mean one-way distance 37.7 µm, and Dally's 15 mm needing a routing factor of ~2.24 over the 6.7 mm mean of a flat 100 mm² layout.

Note the tension: 0.61 µm/byte (area-derived) vs the adopted 1 µm/byte mnemonic.

Workload anchors: - 8192³ matmul ≈ 1 T ops; measured A100 int8 = 0.904 J board energy above idle = 1.64 pJ/MAC at 5.5×10¹¹ MACs. (fp32 15.757 J, fp16 1.426 J, b1 0.648 J — a100-grid-energy-report/results.csv.) - Judge-scored blocked 8192³ schedule = 0.119 J (216 fJ/MAC), 7.6× below the A100; naive = 23.401 J (matmul/energy-report/eightk-audit.json). - Ciresan MNIST 784→2500→2000→1500→1000→500→10, 99% achievable, ~0.7 T ops forward; A100 energy ~2.5 J. One INT8 inference epoch over 60,000 images measured at 12.82 J on an A100-SXM4-40GB. - Meeting #30 closes with the scale target: "1 quadrillion moves."

3.4 Open forks on grid size


4. Decision 3 — the physical dimensions

4.1 The adopted set

The authoritative statement is the meeting #30 agenda, verbatim:

Design choices: 1. Keeping dataset in main memory is weird, augment grid model with streaming 2. Use physically plausible dimensions: 1 micron spacing between bytes 1 fJ to transfer 1 byte by 1 micron in 0.5 ps (obtaining by matching latency, energy, area against 256 MiB in Dally CACM)

docs/grid-vm-multiprocessor.html §2 states the same machine as settled:

Quantity Value Reading
Cell one byte at a lattice site, 1 µm pitch multi-byte values are adjacent cells, little-endian
Movement 1 fJ per byte per micron of Manhattan path the only priced quantity; arithmetic free
Propagation c/160 = 1,874 µm/ns 0.53 ns/mm (Dally's global-wire figure is 0.4 ns/mm = c/120)
Issue 1 ns per instruction, no overlap with flight (time v1) D4
Tapes input at (−1,0), output at (1,0), instructions at (0,−1); memory above the origin serial, data-independent order, blocking receive only
Port crossing E_in = E_out ≈ 5,000 fJ per byte chiplet/HBM class (0.6 pJ/bit)
Fetch 4 fJ per static line + 1 fJ per issued instruction a loop buffer amortising fetch
Writes free in v1 (D2) the write of one op is the read of the next

The calibration check that comes with it (added to the MNIST-on-grid doc 2026-09-04): one byte per grid location, 1 µm pitch, 1 fJ per byte per 1 µm Manhattan edge one-way, c/160. Applied to Dally's own 256 MiB benchmark, that machine gives 268 mm², 175 pJ per 64-bit access, 12 ns — against Dally's 100 mm², 58 pJ, 12.3 ns. So the adopted constants are 3× his energy and 2.68× his area, with latency matching almost exactly. That is the real "fitted to Dally" claim, and stating it in the spec is much stronger than "matched to a 256 memory bank", because a reader can check it.

And the constant that quietly decides the leaderboard: the endpoint term. Under the current pure-distance rule (1 fJ × S) the best 8192³ schedule beats naive by 5.14×. Under 2.5 fJ × (S + 160R) — i.e. adding a ~400 fJ fixed cost per paid read — the same comparison collapses to 1.0034×. A fixed per-access endpoint charge (which is what every real SRAM has, and what Dally's own 50 + 0.022√S law encodes as its additive 50 fJ/bit) almost entirely erases the advantage the benchmark exists to measure. Whatever is chosen, it should be chosen deliberately and published with the task, not inherited.

Your stated design method, which is worth writing into the spec (20260903-180729):

"I took the numbers to match the energy … of Bill Dally's model … match the energy of a 256 memory bank. And I chose the other to be within the same order of magnitude for time and area, while being easy to remember."

That is: energy is fitted; time and area are not fitted, they are mnemonic. Say so in the spec rather than letting a reader assume all three are derived.

4.2 Dally's published numbers — the calibration targets

CACM 65(9), Sept 2022 (On the model of computation: point), verbatim: - 32-bit add: 20 fJ and 150 ps. Fetching its two operands from main memory: 1.3 nJ — 64,000×. - Moving two 32-bit words 1 mm: 1.9 pJ and 400 ps. - 64 bits 40 mm corner-to-corner on a 400 mm² chip: 77 pJ and 16 ns. - Off chip: 320 pJ per 64 b, 6 ns per metre. - 8 KB (64 Kb) sub-array, 64 b access: 0.64 pJ, 300 ps. - 256 MB (≈100 mm²) from those sub-arrays: 58 pJ and 12.3 ns, of which 57.4 pJ and 12 ns are communication — 15 mm each way, 30 mm total. - CPU overhead turns a 20 fJ add into an 80 pJ add instruction. - PECM: "Each processing element and each memory is assigned a location l ∈ L and a distance function D: L × L ⇒ R … Two-dimensional Manhattan distance models on-chip communication … Off-chip communication can be modeled by adding additional distance when either coordinate spans a chip boundary." — note this model is explicitly parallel; the settled Grid VM is its P = 1 case.

CACM 63(7), July 2020 (Domain-Specific Hardware Accelerators), 14 nm: - 8-bit add 10 fJ and 4 µm²; double-precision FP multiply 5 pJ and 3,600 µm². - 8 KB local memory 50 fJ/bit, SRAM 0.013 µm²/bit. - On-chip communication 100 fJ/bit·mm. 100 MB memory ≈ 0.7 pJ/bit. - LPDDR4 ~4 pJ/bit; SDDR4 ~20 pJ/bit; off-chip SerDes ~10 pJ/bit. - And, crucially, the closed-form √-address law itself: the cost of accessing a memory of size S bits is 50 + 0.022·√S fJ per bit. That constant decomposes exactly: 0.022 = 2 (round trip) × 100 fJ/bit·mm × √(0.013 µm²/bit). Check at S = 800 Mbit: 50 + 0.022·28,284 = 672 fJ/bit ≈ the paper's own "0.7 pJ/bit for a 100 MB memory". ✓

This is the literature statement of the benchmark's pricing rule, in Dally's own words, and it is not currently cited anywhere in the repo. ⌈√a⌉ is not an approximation invented here — it is his formula with the additive endpoint term dropped.

AHA 2023 retreat keynote, in Dally's own preferred unit: - "Cost of an add (1 fJ/bit) = Cost of going 10 µm"; "An add is worth 10 µm of movement." Our grid charges 1 fJ/byte/µm = 0.125 fJ/bit/µm = 1.25 fJ per bit per 10 µm — i.e. 1.25× Dally's own figure. That is the tightest agreement anywhere in the constant set, and it is stated in the unit he prefers. - Instruction overhead, the axis that separates CPUs from TPUs: fetch + decode + operand fetch = 30 pJ; HFMA 1.5 pJ (2,000% overhead), HDP4A 6.0 pJ (500%), HMMA 110 pJ (22%), IMMA 160 pJ (16%). An out-of-order ARM A-15 integer add instruction: 250 pJ, ~4,000× the energy of the add itself. - Memory hierarchy per 64-b word: local SRAM (KB) 5 pJ, on-chip SRAM (MB) 50 pJ, LPDDR (GB) 640 pJ.

4.2b Dally's own number has moved, twice

The single largest unresolved discrepancy in the physical constants is not between us and him — it is within his own published work:

Year On-chip wire energy Source
2020 100 fJ/bit·mm CACM 63(7)
2022 ~30 fJ/bit·mm (1.9 pJ / 64 b / mm = 29.69; cross-checks at 40 mm: 77 pJ/64 b/40 mm = 30.08) CACM 65(9)
2023 ~100 fJ/b·mm AHA retreat keynote, twice

A 3.4× swing, same author, and node cannot explain it — Jouppi reports wire energy per unit length improved less than 2× from 45 nm to 7 nm.

This matters concretely, because our two derivations use different branches. The internal-log constant table (30 fJ/bit·mm, 400 ps/mm, 10 fJ/bit endpoint, 5 pJ/bit off-chip, 6 ps/mm) is CACM 2022 renormalised per bit. But 1 fJ/byte/µm = 125 fJ/bit·mm, which is calibrated against the 100 branch with a ×2 round-trip factor at a 0.61 µm byte pitch — not against the 30 branch that the same document quotes as its justification. Anyone re-deriving it from CACM 2022 gets a number 3.4× smaller and concludes there is a bug. Pick a branch, write the arithmetic into the spec, and state the round-trip convention.

Two further quoting hazards inside CACM 2022 itself: it says 1.3 nJ for a main-memory operand fetch in one paragraph and 1.2 nJ three paragraphs later; and it reports the same locality example as both 354× and 230×. Its 320 pJ is the crossing; its 640 pJ is the full DRAM access — do not cite them for each other.

4.3 The renormalisation that produced the working constants

Every line of your 03sep26 constant table is CACM 2022 re-expressed per bit:

Your line Dally source Check
On-chip movement 30 fJ/bit·mm 1.9 pJ / 64 b / mm = 29.7 ✓
On-chip latency 400 ps/mm same move ✓ (= c/120 exactly)
8 KB SRAM endpoint 10 fJ/bit, 300 ps 0.64 pJ / 64 b ✓
Off-chip 5 pJ/bit per crossing 320 pJ / 64 b ✓
Off-chip propagation 6 ps/mm 6 ns/m ✓

Independently, CACTI 7.0.3 at 22 nm (tmp/cacti-sram8k-22nm.txt, 8 KB scratch RAM, itrs-hp, 300 K, delay-optimal) gives: access 0.129 ns, read 2.048 pJ / 64 b = 32 fJ/bit, sub-array 20.56 µm × 58.96 µm, data-array area 0.00683 mm² at 67.8% efficiency. Wire at 22 nm: 280 fJ/bit·mm repeated full-swing, 2.84 fJ/bit·mm low-swing.

Two consequences worth stating in the spec: 1. "fJ per bit per mm" is undefined until the signalling style is fixed — the two CACTI figures differ by ~100×, and Dally's own two papers differ by 3.4× (30 vs 100 fJ/bit·mm). 2. There is no 14 nm CACTI datapoint in tmp/ — the file named 8k-14nm.cfg actually sets -technology 0.032. CACTI 7 tops out at 22 nm.

4.4 Technology node — unpinned

Internal log 02sep26: "use 7nm because Jouppi gives numbers / Dally also uses 7nm." But Dally CACM 2020 states 14 nm; the AHA 2023 slides use 28 nm for area/add and 45 nm for the instruction table; the CACTI runs are 22 nm. Nothing in the dictated notes names a node at all — only "conceivably implementable on existing CMOS process."

4.5 The pitch question, and the best available defence of 1 µm

The grid needs exactly one number: how many µm² does one byte of addressable memory occupy? Then pitch = √(µm²/byte). Every candidate, side by side:

Basis µm²/byte ⇒ byte pitch Node
Dally's stated SRAM area (0.013 µm²/bit) 0.104 0.322 µm 14 nm
TSMC N7 HD SRAM bitcell 0.216 0.465 µm 7 nm
Dally's 256 MB in ~100 mm² (array level) 0.373 0.610 µm not stated
TSMC 16 nm bitcell 0.56 0.748 µm 16 nm
Intel 22 nm SRAM cell 0.736 0.858 µm 22 nm
TPUv4i CMEM macro (128 MB = 28% of a <400 mm² die) 0.834 0.913 µm 7 nm
the adopted grid 1.0 1.000 µm —

1 µm per byte is almost exactly the macro-level byte density of a shipping 7 nm inference chip's main scratchpad (TPUv4i's CMEM, 0.913 µm/byte). That is the strongest single justification for the choice and it is not cited anywhere in the repo — it is worth putting in the spec, because it converts "easy to remember" into "matches real silicon."

Two cautions that go with it: - Dally's 0.013 µm²/bit at "14 nm" is 5.4× denser than TSMC's 16 nm bitcell. It cannot be a bitcell figure; flag it as an idealisation. - Dally's 256 MB in 100 mm² is ~2.2× denser than TPUv4i achieves at 7 nm (2.56 vs 1.14 MB/mm²). Sizing the grid from his number is optimistic by that factor.

Sub-array cross-check: at 0.013 µm²/bit an 8 KB sub-array is 852 µm² of bitcell = 29.2 µm on a side, against the 55.2 µm square-equivalent the 100 mm² / 32,768 decomposition allocates it — ~28% array efficiency, which is plausible for SRAM with decoders, sense amps and an H-tree (CACTI measures 67.8% for a delay-optimal scratch RAM, so ~28% is on the conservative side).

4.6 The two softest constants

The 8 KB SRAM endpoint is the softest number in the model:

Value Source Node
10 fJ/bit (adopted) Dally CACM 2022, 0.64 pJ / 64 b not stated
50 fJ/bit Dally CACM 2020 14 nm
117 fJ/bit Jouppi ISCA 2021, 7.5 pJ / 64 b 7 nm
156 fJ/bit Horowitz ISSCC 2014, 10 pJ / 64 b 45 nm
32 fJ/bit CACTI 7.0.3, 2.048 pJ / 64 b 22 nm

10 fJ/bit is the most optimistic figure in the literature, by 5–12×. The plausible reconciliation: 0.64 pJ is bitcell + local bitline/wordline only, while 7.5–10 pJ is a macro access including decode, sense amps, periphery and the local H-tree. Say which one the model prices — a grid whose endpoint cost is 12× below the measured macro number will systematically under-reward locality relative to real silicon.

E_in = 5,000 fJ/byte is an HBM/chiplet crossing, not Dally's off-chip number. The repo's own correction, verbatim: "320 pJ per 64 b is 5 pJ per bit, 40,000 fJ per byte, a board-level crossing; the 5,000 fJ per byte the competition uses for E_in is a chiplet- or HBM-class crossing (0.6 pJ/bit), and the two must not be cited for each other." Cite HBM2 (250–450 pJ/64 b = 3.9–7.0 pJ/bit) for it, not CACM's 320 pJ.

4.7 Numbers available for free, if time and area become scored

Nothing needs deriving — the literature already supplies them: - Time: 150 ps (32-b add), 400 ps/mm (on-chip), 300 ps (8 KB sub-array), 12.3 ns (256 MB), 16 ns (40 mm), 6 ns/m (off-chip). - Area: 4 µm² (8-b add), 3,600 µm² (DP FP multiply), 0.013 µm²/bit (SRAM). - Node scaling, 45 → 7 nm (the only source giving the same quantity at two nodes, so the only basis for a "same program, different node" column): logic 2.1–4.3× (geomean 2.6×), SRAM only 1.3–2.4×, DRAM 6.3×, wire <2×.

One flag: the current time model's "1 ns per instruction, no overlap with flight" is an invention, not a literature number, and is ~3–7× slower than the 700–1,050 MHz TPU clock it is implicitly modelling.

4.8 What the HMM adds, beyond justifying ⌈√a⌉

docs/spatial-model-analysis.html already establishes that charging f(x) = x^α with α = ½ is exactly [AACS 87]'s Hierarchical Memory Model, whose authors themselves motivate it as 2-D physical distance. Its results at α = ½: scanning N inputs is Θ(N^1.5); sorting and FFT are also Θ(N^1.5) (the log factors vanish); binary search collapses from log N to Θ(√N); matmul is Θ(n³) for α < ½, Θ(n³ log n) at α = ½, Θ(n^(2α+2)) above — i.e. α = ½ is precisely the critical exponent at which cubic linear algebra turns movement-bound. Any time-T space-S RAM computation costs at most T·√S.

The one result with a direct design consequence is Thm 5.2: Block-LRU dynamic placement is O(1)-competitive with the offline optimum. That says a compiler that relocates data between phases is provably within a constant of optimal, and static single placement is not — which is a real argument about whether release/stage should be the only relocation mechanism, and it is measurable: the repo's own compiler audit puts the current hand layouts at 1.41–2.79× over the provably optimal static assignment, where the worst assignment would be 3.70–8.58×.

⚠️ That "1.41–2.79×" is the report's own summary line, but its own table lists 2.79, 2.78, 3.96, 1.68, 1.41, 2.46 — so the true range is 1.41–3.96×. The worst/optimal range (3.70–8.58×) does match the table. Fix the summary before quoting it.

4.9 Open forks on physical dimensions


5. Decision 4 — what competitors submit

5.1 What is submitted today

sparse-parity README, verbatim:

"A submission is one straight-line program (an IR) in the v3 instruction set: 8-bit cells, no loops, no branches, no data-dependent addressing, at most 2,000,000 lines (every line counts, declarations included). / An IR is plain text. Line 1 declares the input cell addresses … the last line declares the 32 output addresses, and each line between is one op…"

Delivery: "Submit a PR adding your .ir and generator under submissions/." matmul adds: "the scorer (matmul.py) must not be modified."

So the de-facto practice is already the hybrid: .ir + a .md report + the Python generator, with the generator outside the trust boundary. Andy's own sparse-parity records are exactly this shape (optimize_layout(generate_sis_mask(1, 2))).

Current records (origin/main, 2026-09-08):

Problem Best Date By
matmul 4×4 675 2026-09-04 @jurajselep
matmul 16×16 63,819 2026-09-07 @SecurityQQ
sparse-parity 20% band 86,753 2026-09-02 @b0nce
40% / 60% / 80% 141,218 / 149,665 / 182,744 2026-09-01 @npow
100% 392,666 2026-09-02 @jurajselep + @b0nce
symmetry 6-bit / 8-bit 20 / 29 2026-05-11 baselines, no competitor entries

(matmul/agent_prompt.md and directions.json are badly stale — they still quote 73,602 / 69,697.)

5.2 The wall that forces the decision

"Submissions … are straight-line programs … That choice makes energy a syntactic property — the judge scores by reading the program, never by running it … But it has a hard consequence: program length is proportional to work. A 10¹² -operation computation is a 10¹² -line program." — docs/expressivity-scorability-ladder.html §1

1.7×10¹² ops as raw trace text is ~27 TB. Tier-2 MNIST is ~1.2×10⁷ trace lines; tier-3 is ~10⁹ lines, a 10–19 GB file.

Your own answer, internal log 04sep26:

"the submitted artifact has to be the generator of the straight-line program, not the expanded trace"

5.3 The two-track resolution — the actual proposal on the table

docs/expressivity-scorability-ladder.html §2 and RFC #48 §1:

Generators are unrestricted wherever the expansion is scored, and restricted exactly where the schedule is scored.

Made normative in the design report §10: "tiers 2 and 3 are schedule-track only, declared in the rules from day one; tier 1 runs dual-track and serves as the trace-vs-schedule cross-validation corpus."

5.4 The ladder — the vocabulary for arguing about this

Rung Generator may contain Scoring Unlocks
a nothing (straight-line trace) O(L) summation episodic tasks
b affine loop nests, static bounds closed form GEMM, conv, transformer fwd/bwd
c + static recursion (compile-time depth) recurrence relations Strassen, recursive FFT, fractal tiling
d + bijective index maps from a fixed library free (permutation invariance) FFT butterflies, bitonic networks
e + mode: data-dependent selector among k static schedules worst-case over arms stays static MoE with capacity factor, early exit
f + while / data-dependent trip counts undecidable — simulation only convergence-based iteration

"The cliff is between (e) and (f), and it is a theorem, not a taste: once trip counts depend on data, boundedness — and with it cost — is undecidable (Buck)."

Rung (c) is not optional: "an energy leaderboard for matrix multiplication that cannot express the sub-cubic algorithms would prejudge the very question it poses." Rung (d) is free because Σ⌈√σ(i)⌉ = Σ⌈√i⌉ for a bijection σ — permuting which iteration touches which address cannot change the address multiset.

Three routes were compared:

Route Verdict
A — GridASM, flat opcode text, superset of today's IR keep as the SIR level — the address is the cost, visible to writer and reviewer, no compiler in the trust boundary
B — structured DSL (place W : i8[10][9] at (4,0), affine for, recursive def with a kind system, mode, recv/send) adopt surface ideas; lowers to a one-page Schedule IR which is the scored artifact
C — reuse survey decisive negative result: no existing system computes this scoring function. MLIR affine is the right vocabulary but a multi-GB LLVM build to parse 40-line files; Halide and TVM cannot express static recursion (so no Strassen); Wasm/RISC-V solve a sandboxing problem this design does not have. Winner inside C: an embedded Python builder library emitting canonical schedule text.

The recommendation is B ⊕ C: a Python builder as the authoring layer (outside the trust boundary, eager validation — a non-affine index is a TypeError on the submitter's machine), emitting canonical schedule text styled after MLIR-affine notation as the reviewed and scored artifact, consumed by a Rust judge of ~2–3k new lines atop dally-eval. Hand-written schedule text stays a legal submission, so the trace-tuning culture is not orphaned. "In every route surveyed, the grammar itself is the rung enforcement: while is unrepresentable — verification by absence, the cheapest verifier there is."

Explicitly rejected: - Python-script submissions with instrumentation — "allowed too many ways to cheat; unrestricted hosts are not sandboxable" (RFC #48 §1; you reached the same conclusion in Telegram five months earlier). - Traces alone at scale — the 27 TB wall. - Compiler-chosen placement — "under the Dally model placement is the algorithm".

5.6 What Andy has actually built, and what he is asking you

Andy Zhang (@zh4ngx) has eight open PRs, all awaiting you. You have left zero comments on any of them.

PR What Needs you?
#48 RFC: the Scheduled Dally Language — the whole submission-format answer yes
#58 MNIST tier-1 (3×3) demo: 280 ops, cost 4,625, ~0.005 nJ/inference, float 53.0% = int8 53.0% yes
#61 Chebyshev-vs-Manhattan provenance audit yes — see §6.1
#49 three-engine benchmark: Rust CPU 69–77× Python; GPU LDS 0.38–0.90× Rust CPU on static traces, but 2× on search workloads reference
#51 / #52 / #55 / #56 matmul search runs (858 floor), access heatmaps reference
#60 BUILD_NOTES.md + reviewer triage reference

His four named forks, verbatim from PR #48's body: "MNIST data representation, accumulate semantics, reduce primitive, home repo."

His framing of your Sep 2 ask, verbatim from RFC §1: "The request (Yaroslav, 2026-09-02): a language that covers both matmul and MNIST, is immune to cheating (interpreted in a sandbox), and makes memory placement an explicit part of program design."

5.7 What your dictated notes say about this axis

Nothing directly — the words "submission", "instruction stream", "IR", "language", "SDL", and "PR" do not appear in the 14-day window. Two indirect signals, both pointing at generator, not trace:

"I can just generate this competition, and then people will create these data flow algorithms that do learning and prediction, and I'll want to understand how they're doing it, but probably not backprop." — 20260903-183033 (opens in Russian: "Right now I think that I have seen how to get rid of backprop")

A flat trace is precisely the artifact you cannot understand.

"what if we reinvent all of linear algebra, all of algorithms for a new abstraction? And then we have compilers" — 20260904-125521

And on who builds it:

"maybe with Andy I … figure out how to create a data flow computer that does this." — 20260903-073220


6. Cross-cutting decisions you will need anyway

6.1 The metric: Manhattan or Chebyshev — currently self-contradictory on main

This is the sharpest live inconsistency and it sits in published documents.

The identity that settles it, and it is exact:

Shell at radius c Cell count
legacy {a : ⌈√a⌉ = c} c² − (c−1)² = 2c − 1
upper-half-plane Manhattan {|x| + y = c, y > 0} 2c − 1 ✓
first-quadrant Chebyshev {max(x,y) = c} 2c + 1 ✗

So Manhattan is the exact discrete realisation of the existing ⌈√a⌉ law — no substitution is needed. The explicit embedding: k = ⌈√a⌉, t = a − (k−1)² − 1, x = −(k−1) + t, y = k − |x|. Physically, Chebyshev underprices a rectilinear diagonal route by up to 2× ((d,d) costs d under Chebyshev, 2d under Manhattan), which breaks the 1 fJ/byte-micron calibration.

Also uncorrected on main: the competition design report's time formula prints 1.875 µm·ns⁻¹ where c/160 = 1,874 µm/ns — a factor of 1,000.

And a second, subtler metric error that PR #61 also flags. RFC #48 §8.1 charges "each edge … ceil(sqrt(distance)) once." But ⌈√a⌉ is the function that converts a linear address rank into the radius of its packed 2-D shell — applying another square root to an already geometric distance conflates rank with route length. Under native placement the correct form is simply

E_move = Σ_reads bytes(r) · d₁(src, dst) · 1 fJ/byte-µm

This one is easy to miss because it only bites once addresses become coordinates — i.e. exactly when you adopt native 2-D. The correction was drafted in PR #61 and never posted.

6.2 The scoring function

6.3 Anti-cheat — the decision that shapes the competition

Your stated policy, 20260903-090056, which is worth writing into the rules verbatim:

"the reward hack just means they do something and somebody looks at this and they're like, there is no way I can use this technique in real life. … It's like, hey, that's a cool technique, [re]use it. And they hard code — it's like, well, you can't reuse it. And we want to isolate as many reward hacks and fix it on the competition. And I think for the integrity of the competition, allow any other reward hacks because ultimately it's subjective. So it's unfair to the participants to change the rules in the middle of the game."

Decoded: the test is subjective ("could I reuse this?"); enumerate and close hacks up front; then allow the rest and never change rules mid-game.

Reinforced at the Sep 7 meeting: "the goal is algorithms … they love to cheat."

Mitigation triage from the red team (design report §7):

Mitigation Fate Why
Held-out QMNIST pool fails public since 2019; attacker pretrains on the union
Fetch-priced immediates fails kills Θ(K) table bombs, not 80 KB of compressed hypothesis
Program-length caps fails any cap excluding 8 KB of sets excludes honest inference
Per-episode label permutation keep, insufficient alone recoverable by vote-table matching (25–30% tax at tier 1, <1% at tier 3) — but it forces Ω(100) train reads, killing zero-read programs
Per-episode pixel permutation bypassable ~8% tax to attacker, ~2× tax to honest convs

The three options that actually work, to be chosen explicitly: (a) accept it — rename the closed division episodic inference energy, run in-episode learning separately; (b) procedural episode generation (MNIST-1D-style parameters redrawn per episode, so no finite pool exists); (c) per-episode invertible GF(2)-affine mixing — makes recovery equivalent to learning but destroys spatial priors and every calibration number. Recommendation: (b) for tiers 1–2; decide tier 3 after deciding whether convolutional architectures are a protected species.

Note: RFC #48 §9.1 claims D1 is "Resolved via procedural episode generation (S7) for Tiers 1-2"; the design report leaves tier 3 open. The two documents disagree about whether D1 is closed.

Seeds must be keyed, not merely derived: seed_j = SHA-256(server secret ‖ submission bytes ‖ tier ‖ j). Without the secret term a submitter simulates every episode offline and ships the answer sequence as immediates — 100% accuracy at near-zero energy. This was found as a fatal hole in the first draft. K = 50 / 10 / 3 shrinks the mined maximum to ~+0.8 pp at tier 1.

Two attacks recorded as dead: the tier-1 "Bayes table" (59,725 of 60,000 3×3 images are distinct at 8-bit depth, so a top-1000-pattern table scores ~12%), and pool-lookup by linear scan (~6×10¹² trace lines at tier 3, fetch-priced into oblivion).

Your own train/test idea (20260904-092239): "I just have a pool of examples, and I resample both training and test sets." Note the tension with the sparse-parity mask tier, which deliberately has no test set.

Nine pitfalls are already published (docs/index.html §Pitfalls) and were learned the hard way on sparse-parity. A redesign must not lose them:

  1. "The public suite is minable — never adjudicate on it." Demonstrated at n=12: an exploit passed the η ≥ 0.25 scorer at lower cost than the honest entry while its true advantage was 0.223.
  2. "Sampled secrets must be fresh or hidden." A ~33,000-instruction circuit (13% of the cap) enumerating only the dev suite's 128 sampled secrets "scores a perfect η = 1.0 on the dev suite — Pareto-dominating the honest frontier — while measuring exactly 0.0 on a fresh key."
  3. "Output truncation must be priced, not banned" — via the decode/predict energy balance; set m_test ≈ D/P of the reference decoder and re-derive it per tier.
  4. "Candidate enumeration is not the only brute force." Both exponentials must be checked against the cap. m_train = 18 at n = 32 was chosen specifically so the full 2¹⁴ scan exceeds it.
  5. "The instruction cap is part of the contract. Energy alone cannot kill brute force… Raising the cap is a benchmark-version change, not a tuning knob."
  6. Accepted noise is ±4 pp across suite keys — "Energy is static and exact; only the accuracy coordinate is noisy."
  7. Publish pooled curves with spread, not single deterministic points.
  8. Complement pairing needs odd k (both live tiers use k = 3, k = 5).
  9. "No data-dependent addressing cuts both ways" — submissions cannot hash, but the generator is ordinary Python.

Points 1, 2 and 5 are the ones that transfer directly to MNIST, and 5 is the one most likely to be forgotten: the line cap is a rule, not a tuning knob.

6.4 Accuracy targets — measurement overturned the drafts

An adversarial reviewer ran real 600/600 episodes rather than extrapolating, and two recommendations flipped:

Tier Measured Draft target Corrected target
1 (3×3) centroid 48.2 ± 2.1%, linear@600 50.5%, 1-NN@600 54.0% 50% (would kill centroid) ~45%
2 (9×9) linear 84.8% on full data 88% (would exclude linear) ~80–82%
3 (28×28) 1-NN 96.9% — 97.3%, and "never place a target at 97.0" (works only because the task pins a 60k test set, σ = 0.07 pp)

Per-episode binomial noise: ±2.0 / ±0.5 / ±0.09 pp.

The resident-vs-restream lever, genuinely contested. Holding the 60k×784 training set resident costs (2/3)·(4.70×10⁷)^1.5 ≈ 2.2×10¹¹ units per sweep; re-streaming costs 4.70×10⁷·E_in. Resident wins iff E_in ≳ 4.6×10³ — which is why E_in ≈ 5,000 fJ puts the two régimes within ~10% of each other and makes this a real strategic choice rather than a foregone one. "1,000 fJ is wrong."

Stream floors never decide a ranking: 11.4 kB / 0.98 MB / 94 MB in-streams cost 5.7×10⁷ / 4.9×10⁹ / 4.7×10¹¹ fJ — at tier 3 that is ~10⁻³ of MLP training energy.

6.5 Time and area

Internal log 06sep26: "report three things — energy, time, area, accuracy" (four listed).

6.6 Parallelism — decided in outline, P1–P8 open

docs/grid-vm-multiprocessor.html compares four ways to put more than one control unit on the grid, and recommends an adoption order:

  1. SPMD tile mesh with a collective library — adopt first. Smallest delta; keeps the judge a formula; makes parallel makespan, cross-core locality and idle-silicon cost visible. Every message-passing program using only collectives scores identically.
  2. Systolic coprocessor inside a tile ("AI-CPU tile") — the only scenario that prices Roune's thesis; closed-form; first showcase is 8192³ GEMM.
  3. Message-passing multicore — needs the reference simulator first.
  4. SIMT lane array — its own P ≥ 2 division.

Eight open decisions P1–P8, each with a stated cost: tape rate for P ≥ 2 (8 B/ns vs 1), fetch continuity, whether static energy A·T is scored, the idle-clock term (0.125 fJ per bit-micron of clock tree per ns = 168 fJ/ns per tile), the area rule, write pricing, whether SIMT/dp4a enters v1 ("adopting dp4a into v1 drops the 16×16 record 4× overnight"), and makespan on the general path.

Two migration items it flags as settled-but-unwritten: the port term (a recv.T occupies sizeof(T) ns at 1 B/ns with no issue slot) and the area function ((2k−1)² over the lin(a) half-diamond). And one correction: "5 pJ per byte" should be per bit.

6.7 What each proposed change costs the existing leaderboard

Every option has a stated re-scoring price. Collected in one place, because this is the constraint that has silently driven several design choices (D2 "writes free", D3 "keep lin(a)", RFC §8.3) toward compatibility rather than physics:

Change Records it moves
Manhattan lin(a) embedding, 2-D exposed none — bit-exact
Native (x,y) placement requires a versioned leaderboard even under Manhattan
Pricing ports "once ports are priced, legacy free-input records are not score-comparable anyway"
Instruction-fetch pricing every record; none has been re-scored under it
A located loop buffer (P2) every P = 1 record by 2–24%
Charging the idle clock at P = 1 (P4) every P = 1 energy record by 7–80%
8 B/ns tape at P = 1 (P1) every P = 1 time record (streamed sum 2,102,751 → 1,447,809 ns)
Adopting dp4a into v1 (P7) drops the 16×16 record 4× overnight
Priced writes both scorers charge reads only today

The corpus PR #48 §8.3 promises to keep bit-exact is "sparse-parity, 4×4 matmul at 675, 16×16 matmul at 64,431, and the Tier-1 PR #58 demo" — note 16×16 has since moved to 63,819, so that clause is already stale.

6.8 Hosting, and the shape of the competition

6.9 Two things you have said before that a redesign must not lose

The metric has been silently wrong twice. Both times a leaderboard leader was not counting its own work: - Yad, msg 1102 (2026-03-22): "GF2's measured DMC of 8,607 was artificially low. The harness only tracked write(A) → read(A) → write(solution), skipping the O(n²) row operations. Honest tracking puts GF2 at 189,056." - You, msg 1231 (2026-03-27): "It may affect the existing leaderboard because the Gaussian elimination didn't count the cost of the elimination step."

A new ISA should be judged partly on whether it makes this class of error impossible — which is exactly the argument for execution-free scoring over instrumented Python.

Iteration speed is a hard design input, repeatedly stated. msg 689 (2026-03-04): "baseline SGD 22 seconds is too slow. It should be <2 seconds. Making it 22 seconds makes your iteration time 10x slower." msg 925: "Only consider experiments runnable on 1980s hardware in 1 hour (less than 1 second today)."

And the benchmark-overfitting warning, from you (msg 767, 2026-03-09): "keep in mind the AI Radiology debacle where a team of (human) agents optimize the heck out of a benchmark, but the result didn't really work … Unless your evaluation is the actual thing you care about (extremely rare), [you] need to be exercising human judgement to tell if the direction your agents are moving is likely to make impact."


6½. Decisions nobody has named yet

These are not open questions — they are choices that have already been made by accident, or that a competition needs and this one does not have. Grouped by how badly they bite.

A. Made implicitly by code, never chosen by anyone

The accidental decision Why it matters
Free input placement The submitter chooses input addresses and placement is free (both scorers). At tier 3 that is free placement of 4.7×10⁷ cells — an enormous unpriced degree of freedom. The Grid VM report concedes the consequence in passing: "once ports are priced, legacy free-input records are not score-comparable anyway."
No registers addr = 0 is the off-grid ALU; address 1 costs 1. Whether the machine has a register file, and whether register traffic is free, was settled by an address convention rather than a design choice.
Free fan-out / broadcast Writes are free, so one read feeding many consumers is free. Broadcast is unpriced, which silently sets its complexity. The dally-spatial-report notes that pricing writes "restores broadcast to Θ(n^1.5)".
Store always beats recompute Free writes. Deletes Dally's headline example (§2.6).
Two engines, one silent fallback dally_eval.py returns None when the Rust binary is missing and the harness silently falls back to Python; DALLY_EVAL_BIN is environment-overridable, with no pinned version and no engine identity recorded per score. The vectorised engine computes in int16 while the scalar engine wraps to signed 8-bit — they are asserted to agree, not proven across all ops. A scoring harness that can quietly change engines is exactly the failure mode that produced the GF(2) 8,607 → 189,056 correction.

B. What is actually being measured — the biggest unnamed fork

C. Governance — the competition as an institution

Almost entirely absent, and cheap to fix:

D. Ranking — four axes and no rule

E. Machine-model forks still unnamed

F. MNIST-specific gaps


7. Contradictions ledger

Things that disagree in writing, right now. Each needs a ruling.

# The disagreement Sources
C1 Chebyshev vs Manhattan. D3 says native 2-D Chebyshev; the Sep 7 note and RFC §8.3 say Manhattan / keep 1-D lin(a); PR #61 proves Manhattan is the exact realisation (2c−1 vs 2c+1). Uncorrected on main. design report §5/§11 vs multiprocessor §2 vs RFC #48 §8.3 vs PR #61
C2 1.875 µm·ns⁻¹ vs 1,874 µm/ns for c/160 — 1000× design report §5 vs multiprocessor §2
C3 1 picojoule vs 1 femtojoule — resolved: a misspeak, not a fork. The dictated design description does say "the idealized computer is 1 picojoule, 1 micron" (verified across engines), but the same 134-word clip also says a memory access "takes milliseconds" (true value 0.5 ps), and every written source of that same week — the MNIST-on-grid doc (Sep 4), meeting #30 (Sep 8), RFC #48 §8.2, the multiprocessor note — says 1 fJ per byte per micron. Treat the clip as dictation error. 20260903-180729 vs meeting #30 / RFC §8.2
C5 30 vs 100 vs 280 vs 2.84 fJ/bit·mm. Dally's own figure went 100 (2020) → 30 (2022) → 100 (2023) — a 3.4× swing that node cannot explain (wire energy improved <2× from 45→7 nm). CACTI 22 nm repeated wire = 280; low-swing = 2.84. Our two derivations use different branches: the constant table is the 30 branch, but 1 fJ/byte·µm = 125 fJ/bit·mm is calibrated to the 100 branch with a ×2 round trip at 0.61 µm. Biggest single unresolved number. Dally CACM 2020 / 2022 / AHA 2023; internal log 03sep26; tmp/cacti-sram8k-22nm.txt
C5b 8 KB SRAM endpoint: 10 vs 32 vs 50 vs 117 vs 156 fJ/bit. The adopted 10 fJ/bit is the most optimistic in the literature by 5–12×; probably a bitcell+local-wire figure being used where a macro figure belongs. §4.6
C6 c/120 vs c/160 vs c/1200. 400 ps/mm = c/120 exactly (global wire); c/160 = 533 ps/mm (adopted, "conservative"); c/1200 for a complete sub-array access; meeting #30 says 0.5 ps/µm = c/150. internal log 03sep26 vs meeting #30 vs RFC §8.2
C7 1 µm/byte vs 1.25 µm/step vs 1.77 µm/pitch vs 0.61 µm/byte vs 10 µm/bit. Five pitches in circulation. binary-matmul-regimes/derived.json and a100-grid-energy-report/grid-audit.json both derive 1.25 µm per step at 1 fJ per int8 value from 100 fJ/bit·mm (and 10 µm if the fJ is per bit); the multiprocessor note splits the difference at 0.98 fJ for a byte-wide round trip at the 0.61 µm area-derived pitch. And 1 byte/µm² is 800× the 10,000 bits/mm² figure in the same MNIST doc. meeting #30 vs internal log 03sep26 vs multiprocessor §3 vs MNIST-on-grid doc vs derived.json
C8 "5 pJ per byte" vs "5 pJ per bit" off-chip. 320 pJ/64 b = 5 pJ/bit = 40,000 fJ/byte (board-level); the 5,000 fJ/byte used for E_in is a chiplet/HBM-class crossing (0.6 pJ/bit). "The two must not be cited for each other." multiprocessor §3 (explicit correction)
C9 Technology node: 7 nm vs 14 nm vs 22 nm vs 28/45 nm. You wrote "use 7nm because Jouppi gives numbers / Dally also uses 7nm", but Dally 2020 says 14 nm, AHA 2023 uses 28/45 nm, and the CACTI runs are 22 nm. There is no 14 nm CACTI datapoint (8k-14nm.cfg actually sets 0.032). internal log 02sep26 vs the papers vs tmp/cacti-*
C10 Is D1 (embedded weights) closed? RFC §9.1 says resolved via procedural episodes for tiers 1–2; the design report leaves it as the open decision that defines the competition. RFC #48 §9.1 vs design report §11
C11 256 MB vs 256 MiB for the Dally anchor. CACM 2022 vs meeting #30
C12 "report three things — energy, time, area, accuracy" — four listed. internal log 06sep26
C13 PR #58 claims it emits "the exact op set the matmul competition scores" but uses set, which v0's matmul.py rejects, and re-implements static_cost(). PR #58 vs matmul/matmul.py
C14 Meeting time 18:00 vs 18:15. top-level doc vs meeting #29 doc

8. What is not written down anywhere

Genuine gaps, not disagreements.

  1. The grid's actual size. Per-tier memory caps exist (1/16/256 MB); a cell count, address-space bound, or "the machine is N×N" statement does not. Your dictated notes never mention it once.
  2. A migration story from the existing leaderboards to the new machine. The Sep 3 dataflow/MNIST rethink appears in your notes as a fresh start. The matmul and sparse-parity records, the .ir format, and the energy numbers are not discussed at all in the 14-day window — even while you were actively producing sparse-parity results. Your own PR-#18 precedent (new IR → new problems only) would resolve this cheaply if stated. The one useful outside analogy is Jouppi's Lesson ⑩: Google deliberately chose source-level ("backwards ML") compatibility over binary compatibility, precisely so the hardware could change underneath the programs. Applied here, that says: when the competition switches from traces to SDL, re-score the programs, don't preserve the traces — which is the opposite of what "bit-exact record continuity" (RFC §8.3, D2, D3) is currently optimising for. Worth deciding on purpose rather than by default.
  3. Cited artifacts that do not exist. Verified absent from main and every remote branch: - "Proposal: a streamed instruction processor and a compact schedule language, this repository, September 4, 2026 — the four-port geometry, the canonical instruction encoding, charged reads and writes." The only cited source for the four-port geometry and for charged writes. - scoring-at-scale/ — cited in the Grid VM report as the source of the "0.3–3 ms to score a 1.1×10¹²-instruction schedule" measurement. The headline judge-cost claim has no artifact. - matmul/a100_8192_dtype_energy_results.json — cited as the measured A100 anchor; git log --all shows it has never existed. The content survives as a100-grid-energy-report/measurement-audit.json (which names it as source_artifact with its sha256), and the 0.904 J INT8 figure is verifiable there. - "S7" — cited by RFC #48 §9.1 as the resolution of the embedded-weights fork. Defined nowhere in any repo document. - docs/pr48-response-draft.md — cited by PR #61, not in the repo. - No SDL parser or interpreter exists in any PR. RFC #48's security claim ("no I/O, no host calls, no unbounded allocation") is a specification, not code.
  4. What "256 memory bank" refers to in your design description, and which Dally table the fJ anchor was fitted to. (sutro fJ report is the candidate; it is only 545 bytes.)
  5. The provenance of Dally's micron figures — your own inline note "(where is the micron measurement from?)" is still open.
  6. Number format. Bytes everywhere; no statement on precision, signedness, or accumulate width, beyond the design report's u8/i16/i32 proposal.
  7. Whether time and area are scored or only reported, and how a winner is defined across four axes.
  8. Off-chip / DRAM pricing. Only the pair "5 pJ/bit per boundary crossing" and "6 ps/mm" exist. Tier 3 has no meaning without it (PR #58 says so explicitly).
  9. Determinism of the eval suite. chatgpt-report.txt flags that the current scorer uses secrets.SystemRandom() — "good for unpredictability, but not for deterministic leaderboard comparisons" — and recommends a public deterministic suite (committed seed) plus a private one (fixed server-side key), with "Do not regenerate the final suite on every scoring run."
  10. Tie-breaking. chatgpt-report.txt proposes an escalation ladder (smoke 1k → dev 10k → leaderboard 50k → strong 100k → final 250k → tie-break 1M → audit ~2M) with "if two submissions are within ~0.3 pp, rerun those only at 1M–2M tasks." Not adopted.
  11. The recent Telegram record is not on this machine. Everything Aug–Sep lives on the intel Mac. PR #61's diff is the only local window.
  12. The Sep 8 memory-wall thread has not reached the design. 20260908-122822: "I got really nerd sniped by the numbers"; 20260908-141736: "there's some useful takeaways in terms of why we have the memory wall." Worth folding in before fixing axis (3).
  13. The Sep 7 meeting produced "a hierarchy of three" (20260907-163821) — probably the three MNIST tiers, but the clip does not say so, and meeting #30's own notes are thin.
  14. Load-bearing decisions that live only in private, uncitable threads. These cannot go into a public spec, and several are the only record of how a choice was made:
    • gemini.google.com/app/f39fddafa2095317 (11apr26) — the only cheat-proofing thread in the entire corpus; its content is nowhere summarised.
    • chatgpt.com/s/cx_6a99e9a8379c81918ee21b94acdd9235 — the "Compare matmul costs across chips" megathread that is the stated source of the whole 30 fJ/bit·mm / 400 ps/mm / 10 fJ/bit / 5 pJ/bit constant block.
    • claude.ai/code/session_01USznRPfqyKT1tbu8UeDJ1P — "divergence architecture in andy/yaroslav", the only record of how Andy's design and yours diverged.
    • The ISA threads (chatgpt.com/g/…/c/69d19662-…, and two others).
    • Dally 2020 and Jouppi are linked as private Drive files (drive.google.com/open?id=…) — unusable in a public spec.
    • http://100.70.243.24:31095/… — a tailnet IP, unreachable to anyone else.
  15. The primary source is not archived. Dally CACM 2022 returns 403 to automated fetch; the working copy is a university mirror plus a local tmp/pdfs/ that is not on origin/main. Every constant in the benchmark traces to a document the repo does not contain.

9. Appendix — the complete fork index

Every choice that has to be made, with its status. DECIDED = settled in code or in a document you wrote; PROPOSED = someone has recommended an answer and it awaits you; OPEN = no answer exists.

ISA

Fork Status
I1 Which instruction set is normative across all problems? OPEN
I2 Cell width and data representation OPEN
I3 Are writes priced? OPEN — and a documentation bug today
I4 Is arithmetic priced? DECIDED (free) in code; the 14.1× is an unaddressed known error
I5 Is instruction fetch priced? PROPOSED
I6 Do ports (recv/send) exist as instructions? PROPOSED
I7 What is E_in / E_out? PROPOSED (number chosen, not ratified)
I8 Read/write time asymmetry ("reads have 2× the time") PROPOSED, unimplemented anywhere
I9 Accumulate semantics at the macro level OPEN — the only one of Andy's four original forks still posed as a question
I10 Is there a reduce primitive? PROPOSED (self-resolved by the proposer, never ratified)
I11 Is mode admitted? PROPOSED
I12 Where is the expressivity ceiling (which rung)? PROPOSED
I13 Division-by-zero and trap semantics OPEN (shipped code contradicts the design note)
I14 Keep matmul's degree-≤2 symbolic check? DECIDED in code (PR #31); not revisited
I15 Are tensor/SIMD instructions (dp4a) admitted into v1? PROPOSED — adopting drops the 16×16 record 4× overnight
I16 Instruction/line cap for matmul OPEN (matmul has none; sparse-parity has 2 M)

Grid size

Fork Status
G1 What is the grid — subarray, chip, or nested pair? OPEN
G2 Address-space bound OPEN (two shipped scorers already disagree)
G3 What is the next problem? DECIDED in substance (MNIST)
G4 The MNIST tier ladder DECIDED by repetition; the 60k-test departure from standard MNIST is unflagged
G5 Accuracy targets per tier PROPOSED (D5)
G6 Episodes per tier (K) PROPOSED (D5)
G7 Off-chip / memory tier PROPOSED (intent stated, nothing implemented)
G8 Fuel / instruction caps per tier PROPOSED
G9 Resident dataset vs re-streaming PROPOSED (D6)
G10 256 MB or 256 MiB for the Dally anchor OPEN (trivial, but load-bearing for the derivation)
G11 Matmul tiers (does 8192³ join the board?) PROPOSED
G12 Sparse parity — retire, resize, or keep? OPEN

Physical dimensions

Fork Status
P1 Cell pitch and quantum de-facto adopted, but "pitch size is still open" was never retracted, and the 10 µm/bit block is still un-struck in the same doc as the 1 µm/byte one
P2 Movement energy per distance, and one-way vs round-trip OPEN (three literature values from one author)
P3 Propagation velocity PROPOSED, with a published 1000× unit error to correct
P4 Technology node DECIDED (7 nm) in the log, contradicted by every source it cites
P5 Is the memory endpoint priced at all? OPEN — softest constant in the model, and it moves the 8192³ verdict from 5.14× to 1.0034×
P6 Wire signalling style (repeated vs low-swing) OPEN — raised nowhere in any design doc, yet it spans ~100×
P7 Off-chip crossing energy PROPOSED
P8 Manhattan / Chebyshev / Euclidean OPEN — the highest-leverage one-line decision available; Chebyshev is live on Pages in a document you committed, and the PR retracting it has had no reply
P9 Scalar lin(a) vs native (x,y) PROPOSED (agreed by both parties)
P10 ceil(sqrt(distance)) applied to a geometric edge OPEN (correction drafted, never posted)
P11 Grid geometry (half-diamond at the origin) DECIDED
P12 Is TIME scored, and with what semantics? PROPOSED
P13 What is AREA? PROPOSED
P14 Static/leakage energy and the idle clock PROPOSED
P15 1 pJ vs 1 fJ in your spoken vs written record resolved as a misspeak (§7 C3)
P16 The "milliseconds" artifact resolved as a misspeak

Submission format

Fork Status
S1 Trace, generator, or both? the log sentence ("the generator, not the trace") and the two-track note are not reconciled
S2 If a language, which shape? PROPOSED — two competing proposals: Andy's SDL and the Grid VM stack
S3 Where does the language live? OPEN, and untracked — it exists only in a PR body
S4 Which scorer is normative? PROPOSED (dally-eval)
S5 I/O order — fixed or moving? OPEN — you wrote the question on 04sep26; the page answering it now exists on main
S6 Online or transductive? DECIDED (batched inference), with an unexamined hole
S7 Does the organizer ever execute submitted code? PROPOSED
S8 What is reported, and what is ranked? OPEN — the largest unwritten piece of the redesign
S9 Leaderboard continuity OPEN
S10 The submission bundle (IR + report + generator?) OPEN in rule, settled in practice
S11 Instruction caps as part of the contract DECIDED as a principle; per-tier numbers PROPOSED
S12 Adjudication protocol PROPOSED

Cross-cutting

Fork Status
X1 Coarse-and-memorable vs accurate DECIDED as philosophy, unstated in any spec
X2 Cheat-proofing policy DECIDED in voice, never written down anywhere
X3 Episode supply / embedded weights OPEN — the highest-stakes fork in the set
X4 Train/test resampling PROPOSED
X5 Single control unit or many? PROPOSED (P1–P8, your own commit 5706824)
X6 What is the model even called? OPEN — "simplified Dally model", "Grid VM", "PECM", "spatial computer", "stream processor with a scratchpad" are all in use
X7 Who adjudicates, and what happens to the nine open PRs? OPEN — a collaboration risk, not just a technical one
X8 The two parallel design tracks (yours and Andy's) OPEN
X9 Documentation consistency before any launch OPEN (pure cleanup, but it is what a visitor sees)
X10 Hosting, community, distribution OPEN
X11 Judge cost and hardware budget PROPOSED / mostly measured
X12 Launch timing, and whether the practical track survives OPEN

Public pages - Benchmark report — https://cybertronai.github.io/sutro-problems/docs/ - Spatial-model analysis — https://cybertronai.github.io/sutro-problems/docs/spatial-model-analysis.html - Grid VM competition design — https://cybertronai.github.io/sutro-problems/docs/grid-vm-competition-design.html - Expressivity–scorability ladder — https://cybertronai.github.io/sutro-problems/docs/expressivity-scorability-ladder.html - Multiprocessors on the Grid VM — https://cybertronai.github.io/sutro-problems/docs/grid-vm-multiprocessor.html - Fixed or moving I/O order? — https://cybertronai.github.io/sutro-problems/docs/dataflow-io-order.html - matmul energy report — https://cybertronai.github.io/sutro-problems/matmul/energy-report/ - matmul animations — https://cybertronai.github.io/sutro-problems/matmul/animation/ - sparse-parity — https://cybertronai.github.io/sutro-problems/sparse-parity/

Repos - https://github.com/cybertronai/sutro-problems - https://github.com/cybertronai/simplified-dally-model · instruction-sets - https://github.com/cybertronai/dally-eval

Open PRs needing you: #48 · #58 · #61 · #49 · #51 · #52 · #55 · #56 · #60

Literature - Dally, On the model of computation: point, CACM 65(9) 2022 — https://cacm.acm.org/opinion/on-the-model-of-computation-point/ - Dally, Turakhia, Han, Domain-Specific Hardware Accelerators, CACM 63(7) 2020 - Dally, AHA 2023 retreat keynote — https://aha.stanford.edu/sites/g/files/sbiybj20066/files/media/file/aha-retreat-2023_dally_keynote_en_eff_ai_hw_0.pdf - Horowitz, Life Post Moore's Law slides (55 pp, Jul 2023; proposal from slide ~37) — https://aha.stanford.edu/sites/g/files/sbiybj20066/files/media/file/aha_071923_horowitz_newcad_0.pdf · MICRO 2023 keynote — https://www.youtube.com/watch?v=q8WK63joI_Y - Jouppi et al., Ten Lessons From Three Generations… TPUv4i, ISCA 2021 - Aggarwal, Alpern, Chandra, Snir, A model for hierarchical memory, STOC 1987 — the HMM with f(a)=a^α; our model is α = ½ - Hong & Kung, I/O complexity: the red-blue pebble game, STOC 1981 - Buck, Scheduling dynamic dataflow graphs with bounded memory (PhD, Berkeley 1993) — the undecidability that fixes the (e)/(f) cliff - Karp, Miller, Winograd, JACM 14(3) 1967 — affine schedules for the v2 time division - Greydanus & Kobak, MNIST-1D, ICML 2024 — procedural generation as the memorization defense - Yadav & Bottou, Cold case: the lost MNIST digits (QMNIST), NeurIPS 2019 - Onur Mutlu dataflow lectures — https://www.youtube.com/@OnurMutluLectures/search?query=dataflow · https://www.youtube.com/watch?v=waqM1JM9GU0

Local paths worth remembering - ~/drive/gdocs-archive/<docId>/<date>.txt — nightly full text of 35 Google Docs - ~/My Drive/hiq-transcribed/ — dictated notes, four engines per clip - ~/…/sutro-problems/tmp/cacti* — the CACTI 22 nm / 32 nm SRAM runs - ~/…/SutroYaro/telegram.db — group chat, Feb 9 – Mar 28 2026 only - ~/…/flush-util/machines.toml — says the live Telegram session is on the intel Mac