Grid-model redesign: decision dossier
Every source bearing on the four open decisions — the instruction set, the grid size, the physical dimensions, and what competitors submit — with what is decided, what is proposed, what contradicts what, and a 58-item fork index. Companion to The Grid VM: competition design, The expressivity–scorability ladder and Multiprocessors on the Grid VM. September 8, 2026.
Assembled 2026-09-08 (week 37). Everything bearing on the four decisions you named: the instruction set, the grid size, the physical dimensions, and what competitors actually submit. This is a reference, not a proposal — it says what is written down, where, and what is still open. Nothing here is a recommendation unless a source made one, in which case the source is named.
Do not push this file as-is. It links private ChatGPT/Gemini/Claude sessions and internal Google Docs, and the repo root is served publicly by GitHub Pages. It is deliberately left uncommitted.
Read-order if you only have ten minutes: §0 (state of each decision) → §9 (the complete fork index — 58 choices with status) → §7 (contradictions) → §8 (gaps). If you have thirty: add §6½ (the forks nobody has named yet) — that is where the answer to "what am I forgetting?" mostly lives.
0. State of each decision
| # | Decision | Status | Where the current answer is written |
|---|---|---|---|
| 1 | ISA | v0/v3 live and frozen per-problem; the redesign ISA (13 ops + recv/send + mode) is proposed, not merged |
simplified-dally-model/instruction-sets/; docs/grid-vm-competition-design.html §4 |
| 2 | Grid size | Three MNIST tiers decided (3×3/600, 9×9/6k, 28×28/60k). Cell count/address-space caps proposed (1 MB / 16 MB / 256 MB per tier). Pitch-to-capacity mapping open | internal log 03sep26 + 06sep26; docs/grid-vm-competition-design.html §8 |
| 3 | Physical dimensions | Decided as a mnemonic triple: 1 µm pitch, 1 fJ/byte·µm, c/160. Their derivation from Dally is partly unsourced and the per-mm constants contradict each other by 3–9× | Sutro meeting #30 agenda; RFC #48 §8.2; docs/grid-vm-multiprocessor.html §2 |
| 4 | Submission format | Two-track, proposed and coherent: trace for episodic/tier-1, schedule language for at-scale/tiers 2–3. Awaiting your decision — Andy has built it and you have not reviewed PR #48 | PR #48; docs/expressivity-scorability-ladder.html §2 |
The single most consequential thing in the dossier, and it is not on your list of four:
"the one decision that actually shapes the competition is not in the language at all, but in where the episodes come from." —
docs/grid-vm-competition-design.html, §11 closing sentence
That is decision D1 (embedded weights). If the training pool is public, a submitter pretrains offline and ships int8 weights as immediates — ~100 B for a tier-1 linear head, ~80 KB for a tier-3 MLP at ~98%. Embedding costs a ~10⁻⁵ tax and skips roughly 15/16 of tier-3 energy, so the "learning" axis becomes dead code and the benchmark measures inference circuits. See §6.3.
1. Where everything lives
1.1 The 106-commit gap — read this first
Your local working tree at ~/Library/CloudStorage/Dropbox/git0/sutro-problems is
106 commits behind origin/main. The three design reports that settle most of this
exist only on the remote and are not in your checkout:
cd ~/Library/CloudStorage/Dropbox/git0/sutro-problems && git log --oneline HEAD..origin/main | wc -l
| Doc | Landed | What it settles |
|---|---|---|
docs/expressivity-scorability-ladder.html (00c8cb7) |
Sep 4 | rungs (a)–(f); where the scoring cliff is; the fetch fork |
docs/grid-vm-competition-design.html (4f4d3e2) |
Sep 4 | the whole machine: language, scoring function, adjudication, judge cost, targets, decisions D1–D6 |
docs/grid-vm-multiprocessor.html (5706824) |
Sep 7 | four multiprocessor scenarios + open decisions P1–P8; also silently corrects the metric and the c/160 figure |
docs/dataflow-io-order.html |
Sep 3 | ports, determinacy, closed/open division, the resident-vs-restream break-even |
docs/rfc-scheduled-dally-language.md |
branch rfc-scheduled-dally only (PR #48) |
SDL: grammar, closed-form scoring, sandbox |
Live at: - https://cybertronai.github.io/sutro-problems/docs/grid-vm-competition-design.html - https://cybertronai.github.io/sutro-problems/docs/expressivity-scorability-ladder.html - https://cybertronai.github.io/sutro-problems/docs/grid-vm-multiprocessor.html - https://cybertronai.github.io/sutro-problems/docs/dataflow-io-order.html - https://cybertronai.github.io/sutro-problems/docs/spatial-model-analysis.html - https://cybertronai.github.io/sutro-problems/docs/
The only origin-only directory is a100-grid-energy-report/. dally-eval is an
external repo (https://github.com/cybertronai/dally-eval); this repo carries a
74-line shim dally_eval.py. There is no scoring-at-scale/, no mnist/, no grid/ on
any of the 22 remote branches, despite references to them.
1.2 Repos
| Repo | Role |
|---|---|
cybertronai/sutro-problems |
the benchmark suite, leaderboards, design reports |
cybertronai/simplified-dally-model |
the cost model + the versioned ISA specs (v0–v3) |
cybertronai/dally-eval |
the fast Rust scorer (wired in by PR #44) |
~/Library/CloudStorage/Dropbox/git0/spatial-computer/ |
earlier long-form: dally-spatial-report.html, pram-tutorial.html, slides, paper.pdf |
~/Library/CloudStorage/Dropbox/git0/SutroYaro/ |
Yad's repo; telegram.db (Feb 9 – Mar 28 only), mirrored group docs under docs/google-docs/, docs/catchups/ |
~/Library/CloudStorage/Dropbox/git0/sutro/ |
sutro.org, docs/findings/, docs/research/ |
1.3 Google Docs
Full text of 35 of these is mirrored nightly to ~/drive/gdocs-archive/<docId>/<date>.txt
(with meta.json), which is faster and more reliable than opening them.
| Doc | id |
|---|---|
| Sutro internal Log (week 36 onward) | 1IjNHiv… |
| Sutro Group: top level | 1B9867EN… |
| sutro meeting #30 — the authoritative "design choices" list | 11Xj14ug… |
| sutro #29 | 1RYSonuC… |
| MNIST on grid model? | 1uDZ1Nh… |
| Numbers to know | 1vUbTxn… |
| sutro fJ report — how fJ were chosen | 1ZMx22vh… |
| Bjarke Roune, Designing AI Chip Software and Hardware | 1EUVW1tB… |
| sutro group challenge #1: sparse parity | 16eeltCa… |
| Yaroslav's sutro planning sprint #1 | 1oSTIM0h… |
| 08sep26 — interlude show | 1RvwoLX… |
1.4 Chat
- Group discussion: https://t.me/sutro_group. Topics:
General,chat-yaroslav,chat-yad,challenge #1: sparse parity,In-person meetings,Introductions. - Local copy
~/…/SutroYaro/telegram.dbcovers 2026-02-09 → 2026-03-28 only (1,005 messages). The Aug/Sep discussions are not on this machine — the live Telethon session lives on theintelMac (flush-util/machines.toml,telegram_fetch = { writer = "intel" }), read over the SSH forward. HTML exports of the group at~/Downloads/Telegram Lite/ChatExport_2026-04-06/and…_2026-04-20/(the 07-31 export is a different, personal chat). - Andy's PR #61 quotes a database-wide search of 2,588 messages, 2026-02-09 → 2026-09-05. That diff is currently the only local window into the recent Telegram record.
1.5 Dictated working notes
~/My Drive/hiq-transcribed/recording-<YYYYMMDD>-<HHMMSS>-<engine>.txt, engines
soniox / scribe / parakeet / groq (prefer scribe and soniox; groq hallucinates).
392 clips in the Aug 26 – Sep 8 window. Signal is concentrated in ~20 clips; the densest
by far is 20260903-180729, which you titled the design description.
1.6 Meeting cadence and people
Mondays 18:00, South Park Commons, 380 Brannan (meeting #29's doc says 18:15). #29 was Aug 31, #30 was Sep 7. Mission line as of Sep 8: "huge amount of energy waste due to legacy learning algorithms. Address this in a maximally open way."
Contributors on the current leaderboards: @zh4ngx (Andy Zhang), @cosminscn, @jurajselep, @npow, @b0nce, @sigkillme0, @SecurityQQ, @sjbaebae, @SethTS, @adotzh, @ab-10.
2. Decision 1 — the instruction set architecture
2.1 What exists today
Four versioned specs, each strictly extending the last
(instruction-sets/):
| Version | Ops added | Total | Used by |
|---|---|---|---|
| v0 | add, sub, mul, copy |
4 | matmul (4×4, 16×16) |
| v1 | + and, or, not, xor |
8 | — |
| v2 | + set (integer immediate; no read, so free) |
9 | — |
| v3 | + div, cmp, select, abs |
13 | sparse-parity, symmetry |
Three-address code, LLVM-flavoured: %d = add %a, %b ⇒ add d,a,b. Two-operand short
form add d,s ≡ add d,d,s. cmp takes a predicate from {eq,ne,lt,le,gt,ge}.
v3's stated purpose is Gaussian elimination with partial pivoting as straight-line code.
2.2 What the scorer actually implements — and where it differs from the spec
This matters because the code, not the README, is the ruler.
matmul/matmul.py — v0 only. unknown op: … (v0 supports add/sub/mul/copy).
- No cell width at all. Values are unbounded Python ints; correctness is checked
symbolically over a polynomial ring, and intermediates of degree > 2 are rejected
(which admits the usual bilinear algorithms). There is no wrapping.
- No address-space limit. Any integer ≥ 1; ≤ 0 raises.
- Cost: _cost(addr) = math.isqrt(addr - 1) + 1 = ⌈√addr⌉, summed over every source
operand of every instruction, plus one _cost per declared output address at exit.
Writes, arithmetic and input placement are free.
- Its own docstring already says: "the cell at linear index addr sits at Manhattan
distance ⌈√addr⌉ from the core."
sparse-parity / symmetry — v3, and a genuinely different machine:
- 8-bit signed cells, wrapped after every op:
_to_signed_8bit(v) = (v & 0xFF) - 0x100 if ≥ 0x80.
- Address cap bit_length() ≤ 64; IR length cap 2,000,000 lines (declarations included).
- set is free (zero reads); select charges 3 reads.
So "the ISA" is already two incompatible machines. Any redesign has to say which one it is continuous with.
2.3 Your own prior constraints on changing it
These are the only decision-bearing review comments you have ever left in the repo, and they all point the same way:
"Changing the instruction set changes the ruler, so algorithms' performance in the old instruction set aren't directly comparable against algorithms with the new instruction set, I guess maybe I should put the eval in a different repo so agents don't touch it." — PR #1, 2026-04-30
"Each problem has it's own leaderboard. I think evaluation metric should be fixed for each problem. Over time we may discover that something is wrong with current IR, and have v4 or v5 IR of Dally model and start using it for future problems, but for now v3 seems sufficient" — PR #18, 2026-05-14
"Can you modify this to not modify
matmul.py?matmul.pyis where evaluation code sits." — PR #2, 2026-04-30"Could you only include the first two files and not
matmul/exp_madd.py. The reason is that it re-implements the scoring function, but that changes the metric" — PR #1
Implication for the redesign: the precedent you set is new ISA version → new problems, old problems keep their ruler. That is a clean migration story for MNIST and it costs nothing; it also means the matmul and sparse-parity leaderboards need not move.
Note: PR #58 (MNIST tier 1) both uses set — which v0's matmul.py does not accept — and
re-implements static_cost(), which is exactly what you rejected in PR #1.
2.4 The proposed redesign ISA
From docs/grid-vm-competition-design.html §4 and RFC #48 §3:
- ~13 trace opcodes +
recv/send, plus the higher rungs:for(affine),def/static recursion, a fixed bijective index-map library (bitrev,stride(gcd 1),xorc), andmode sel -> c in 0..K. - Placement is syntax:
tile X[...] at (x,y),stage tmp from src at (x,y),place,release. RFC #48 §2.2: "The compiler (such as it is) never chooses addresses." - Ports (your 03sep26 formulation, verbatim from the log):
"A submission is a straight-line program (a circuit) over a distance-priced address space, with one input port and one output port. Ports are read and written in a declared, data-independent order; work cells are charged ⌈√a⌉ per read; port operations are charged a fixed crossing cost."
recv dcharges a flatE_in(the crossing, not ⌈√d⌉ — physically the price of a fresh sample is the crossing, not the destination).send scharges ⌈√s⌉ +E_out. - Four hard grammar constraints, each with a concrete exploit behind it:
1. Port balance — every arm of every
modemust have identical(recv, send)counts, or the tape position becomes data-dependent and you fall off the decidability cliff (Buck). 2. Mode dispatch must be priced — as a read at distance ⌈√(total static bytes of all arms)⌉. Unpriced, a nested 512-leaf mode is a free associative memory and a hardcoded pattern→label table wins tier-1 energy 50–100× over honest centroid. Cap arms (~64) and nesting (≤3). 3. Index maps banned inside mode arms — max-over-arms does not commute with the permutation-invariance rewrite; a four-iteration counterexample undercharges 8 → 5. 4. Recursion must be offset-affine — memoizing on the size tuple alone is wrong; Strassen at 8192 has ~14 memo nodes but 7¹³ ≈ 10¹¹ call sites at distinct offsets. - Semantics that must be pinned before anyone writes a program ("every worked example
produced during the design run had at least one overflow bug until these were fixed"):
u8/i16/i32as adjacent cells, little-endian; wrap mod 256 (two's complement); widen-before-op;x/0 = 0; all cells zero at start; one fresh run per episode; an out-of-range mode selector traps and the episode scores zero; blockingrecvonly — no emptiness test, no clock.
That last prohibition is load-bearing, not stylistic: given an emptiness test, output stops being a function of input alone (Brock–Ackerman), two runs of the same submission score differently, and the leaderboard is noise.
- Instruction fetch must be priced. Unpriced program text is free ROM: a lookup table as immediates costs Θ(K) instead of the Θ(K^1.5) the same table pays in priced memory. Two constants, one per track — ~1 unit per executed line (episodic), 4 fJ per static line + 0.1–1 fJ per issued instruction (schedule). Hutter Prize precedent. The current ISA does not price fetch, and adopting it re-scores every existing record.
2.5 What your own dictated notes add
The notes contain no mention of an instruction set, opcode, IR, or assembly in the whole Aug 26 – Sep 8 window. What they do contain are two design criteria:
"the real issue is that you want to be able to do algorithms with small memory footprint, because that's what ultimately it's about" —
20260907-164225"what if we reinvent all of linear algebra, all of algorithms for a new abstraction? … So the issue with linear algebra is that it … discourages algorithms with small memory footprint. … this was the big surprise of vanilla attention. It's like, well, no abstraction, the cost is invisible." —
20260904-125521
and one governing philosophy for how coarse the model should be:
"the signal cannot go from one [end to the other] in a single clock cycle, so we no longer can pretend that— … Or the danger of modeling is too finely. You can have super accurate models, but if there's many, it splits the community. So one somehow needs a model of computation which is at least conceptually in the right direction, that we can reuse the lessons from, and then fields can specialize." —
20260830-175235
2.6 Open forks on the ISA
- Cell width and datatype. 8-bit everywhere?
u8/i16/i32as adjacent cells? matmul currently has no width. PR #58 flags the concrete failure: "8-bit cells alias int16 logits (verified: wrapped argmax differs); argmax happens off-model." - Accumulate semantics — functional vs in-place rebind. RFC #48 §9.2, open; the 4×4
records indicate in-place with explicit
releasewins. - A reduce / combination primitive — one of Andy's four named forks.
- Fetch pricing on or off, and at what constant.
- Whether
modeships in v1 or only its syntax is reserved. - Read/write time asymmetry. Internal log 06sep26: "reads have 2x the time, round-trip with address, return sends the byte / writes have regular cost." This is a time asymmetry only; energy stays one-way. Not implemented anywhere.
- Whether writes stay free. D2 says free in v1 for record continuity — a compatibility choice, not a physics one. Note the cost of keeping it: free writes make store always beat recompute, which deletes the single most-cited worked example in Dally's own article (activations stored off-chip at 640 pJ per 16-b activation vs recomputed at 160 fJ/MAC; recompute wins below n ≈ 4,000, and by 16× at n = 256). If the benchmark is meant to reward the recompute-over-store instinct, free writes work against it.
- Is arithmetic still free? Inherited unexamined from Dally's PECM ("Arithmetic operations remain unit cost"). But priced at Horowitz's 45 nm figures (8-bit multiply 0.2 pJ, add 0.03 pJ), the 16×16 record's arithmetic alone is 934.4 pJ — 14.1× the entire movement proxy the leaderboard measures. At the current problem sizes the free-arithmetic assumption is not a small idealisation; it is the dominant term being discarded. It blocks any claim that the model ranks algorithms the way silicon does. (Dally's own "an add is worth 10 µm of movement" is the cheap middle option.)
- Division-by-zero. A direct conflict between shipped code and the design note:
mask_sparse_parity._safe_divraisesZeroDivisionError; the Grid VM report specifiesx/0 = 0. A raise is not a semantics a circuit can have. - Sub-byte widths. The grid is byte cells, so Dally's entire "do it with smaller data" branch (int8 / log8 / sym8 codebooks / per-64 scale factors) and Jouppi's 21× Int8-vs-Int32 multiply gap are inexpressible. Either admit that as a declared scope boundary or add sub-byte widths.
2.7 What each fork actually costs in code
The engineering price of every option, in the files that would have to change:
| Wanted change | What must change |
|---|---|
2-D (x,y) addressing |
matmul.py:_cost, _check_addrs, _parse (operands are int(x)); mask_sparse_parity.py:_cost, parse_addrs, _compile_ir (addr_to_idx sorts scalar ints), _compile_vector; addr_renamer.py (assumes cost is monotone in a scalar address); energy-report/eightk_scaling.py:shell_prefix/shell_sum — its closed form F(x) = q(q+1)(4q−1)/6 + (x−q²)(q+1) is 1-D only |
Manhattan lin(a) kept, 2-D merely exposed |
nothing. PR #61 shows the embedding is an exact isometry, so scores are bit-identical. This is the free option. |
| Ports / streaming I/O | _parse (line 1 = input list, last line = output list) disappears; recv/send need entries in the op tables of both scorers, plus E_in/E_out; test_matmul.py:_assert_submission_invariants hard-asserts exactly 32 inputs / 16 outputs |
| Wider words | _to_signed_8bit / _wrap8 in mask_sparse_parity.py; the int16 batch dtype; the set literal bound (−128..255). Note the matmul side is already unbounded-int, so it needs a width added, not widened |
| Priced writes | both simulators charge reads only — 3 cost += sites in matmul.py, 1 loop in _compile_ir |
| Instruction-fetch pricing | no code exists anywhere |
| Time / area metric | nothing exists. eightk-audit.json states it outright: "No throughput or clock rate is used, so the model produces energy but does not produce time." |
| An op cap for matmul | mask_sparse_parity has OP_CAP; matmul.py has none |
| Non-bilinear matmul algorithms | _poly_mul — the degree-2 restriction blocks them by construction |
Two observations worth carrying into the decision:
- Keeping
lin(a)and exposing 2-D coordinates costs literally nothing and preserves every record bit-exactly. That is the cheapest possible resolution of C1. - Byte-packing is the largest empirical lever the ISA has ever produced. The 100% sparse
parity band went 12,461,610 (unpacked, Aug 30) → 938,331 (septet-packed, Aug 31) →
392,666 (Sep 2) — a 32× improvement from packing alone.
docs/index.html: "byte-packing historically bought 5–15× in this repo." Any cell-width decision is therefore also a decision about how much of the leaderboard is packing skill. Note the packed records use 6-bit row masks, not 7 or 8, because signed 8-bit plusnotmakes the top bits hostile — a concrete argument for unsigned cells.
3. Decision 2 — the grid size
3.1 What is sized today
| Problem | Instance | Caps |
|---|---|---|
| sparse-parity | n = 32 bits, k = 5 secret positions, m = 18 strings; 32 output cells | ≤ 2,000,000 IR lines; address bit_length() ≤ 64; 1,024-instance dev suite, fresh 2,048-instance adjudication suite |
| matmul | 4×4 and 16×16 only | no address cap at all |
| symmetry | 6-bit palindrome (8 of 64 patterns) | — |
3.2 The MNIST tier ladder — decided
Stated twice within a minute on 2026-09-03 and again in the log on 06sep26:
| Tier | Image | Train / test | Grid cap (proposed) | K episodes | Fuel cap |
|---|---|---|---|---|---|
| 1 | 3×3 | 600 / 600 | 1 MB | 50 | 10⁹ |
| 2 | 9×9 | 6,000 / 6,000 | 16 MB | 10 | 10¹¹ |
| 3 | 28×28 | 60,000 / 60,000 | 256 MB | 3 | 3×10¹² laptop / 10¹³ CI |
"These scales roughly two orders of magnitude" (20260903-153539). QMNIST supplies the
extras. Memory stays under ~1.5 GB worst case, plus one shared ~110 MB corpus.
Area is defined as the origin-anchored bounding square of touched cells in µm² — deliberately not touched-cell count, which is gameable by scatter — unioned over mode arms.
3.3 Capacity anchors
The arithmetic a grid-size decision has to match, reverse-engineered from Dally in the internal log, 03sep26 (your own parenthetical "(where is the micron measurement from?)" is still unanswered):
- 256 MB ≈ 100 mm², built from 32,768 sub-arrays of 8 KB (64 Kb).
- 64 Kb sub-array = 1,024 words × 64 bits (= 8 KiB = 65,536 bits ✓); as a Manhattan grid with 32 concentric lines; dimensions 62 × 31; "Equivalent Bill Dally is 110 µm × 55 µm" ⇒ 1.77 µm per pitch (110/62 = 55/31 = 1.774 ✓).
⚠️ Two off-by-ones in this derivation, and the pitch inherits them. - The series is wrong: the log says "1+3+…+31 = 1024", but 1+3+…+31 = 256 (16 terms). The identity you want is Σ(2c−1) for c = 1..32 = 32² = 1024, whose last term is 63: "1+3+…+63". - The bounding box is wrong: a 32-shell upper-half-plane diamond spans x ∈ [−31, 31] and y ∈ [1, 32], i.e. 63 × 32, not 62 × 31. (62 × 31 = 1,922 cells, not 1,024.) - So 1.77 µm is derived from the wrong box. 110/62 = 55/31 = 1.774 exactly — which strongly suggests 110 × 55 µm was itself back-solved from 62 × 31 at 1.774. On the correct box the two axes stop agreeing: 110/63 = 1.746 and 55/32 = 1.719.
The conclusion (32 shells hold 1,024 word-sites) is right; the arithmetic under it is not. Worth fixing before it is quoted, because this is the same 2c−1 shell identity that settles the Manhattan-vs-Chebyshev question in §6.1. - ⇒ 181 × 181 grid of 64 Kb sub-arrays, ≈ 10 mm wide, sub-array ≈ 55 µm wide.
The parallel derivation in the multiprocessor note's calibration thread gets 55.2 µm square-equivalent per sub-array and 1.73 µm per 64-bit word site = 0.61 µm per byte, with the 8 KiB sub-array as 32 half-diamond shells of mean one-way distance 37.7 µm, and Dally's 15 mm needing a routing factor of ~2.24 over the 6.7 mm mean of a flat 100 mm² layout.
Note the tension: 0.61 µm/byte (area-derived) vs the adopted 1 µm/byte mnemonic.
Workload anchors:
- 8192³ matmul ≈ 1 T ops; measured A100 int8 = 0.904 J board energy above idle
= 1.64 pJ/MAC at 5.5×10¹¹ MACs. (fp32 15.757 J, fp16 1.426 J, b1 0.648 J —
a100-grid-energy-report/results.csv.)
- Judge-scored blocked 8192³ schedule = 0.119 J (216 fJ/MAC), 7.6× below the A100;
naive = 23.401 J (matmul/energy-report/eightk-audit.json).
- Ciresan MNIST 784→2500→2000→1500→1000→500→10, 99% achievable, ~0.7 T ops forward;
A100 energy ~2.5 J. One INT8 inference epoch over 60,000 images measured at 12.82 J
on an A100-SXM4-40GB.
- Meeting #30 closes with the scale target: "1 quadrillion moves."
3.4 Open forks on grid size
- What is the grid — one subarray, one chip, or a nested pair? This is not stated anywhere, and four incompatible readings are in circulation: (a) 1,024 cells = one 64 Kb subarray (the 03sep26 area math); (b) that subarray as the endpoint, with 181×181 of them = 256 MB on ~10×10 mm (same passage, one level up); (c) a flat unbounded half-plane — what all three scorers actually implement, where "writing to an address makes it exist"; (d) per-tier caps of 1 / 16 / 256 MB (the Grid VM report). Nothing in code has any capacity notion at all. Until this is settled, no area metric and no "how big is the machine" statement is possible, and the multiprocessor tile-pitch derivation has no base.
- Address-space bound. The two shipped scorers already disagree: matmul allows unbounded
positive ints, sparse-parity caps at
[1, 2^64). PR #48's floor-sum evaluator assumes A_max ≈ 10⁸ — and the 8192³ schedule already touches 201,359,490 cells (radius 14,191 steps), so this is not hypothetical. - Grid geometry is decided — upper half-plane / half-diamond with the processor at the origin — but note its area consequence: a 256 MiB half-diamond occupies 268 mm² against Dally's ~100 mm², and its rectangular envelope is ~32.77 × 16.38 = 536.8 mm², half of which is empty unless two triangles are packed together.
- The µm-per-bit vs µm-per-byte split. The MNIST-on-grid doc contains both "10 micron grid size in bits / 1 fJ 1 bit 10 microns / 10,000 bits per mm²" and "Grid of bytes, with 1 micron spacing / 1 fJ per byte moved per (one way) / 1 micron pitch." 1 byte/µm² = 8×10⁶ bits/mm² is 800× the 10,000 bits/mm² figure.
- Off-chip / memory tiers. PR #58 is blunt: "Off-chip DRAM/HBM streaming is not priced; tier 3 (47 MB dataset) needs an explicit memory-tier edge cost before its numbers mean anything." Meeting #30 design choice 1: "Keeping dataset in main memory is weird, augment grid model with streaming."
- 256 MB vs 256 MiB for the Dally anchor.
- Resident vs re-streamed training set — see §6.4.
4. Decision 3 — the physical dimensions
4.1 The adopted set
The authoritative statement is the meeting #30 agenda, verbatim:
Design choices: 1. Keeping dataset in main memory is weird, augment grid model with streaming 2. Use physically plausible dimensions: 1 micron spacing between bytes 1 fJ to transfer 1 byte by 1 micron in 0.5 ps (obtaining by matching latency, energy, area against 256 MiB in Dally CACM)
docs/grid-vm-multiprocessor.html §2 states the same machine as settled:
| Quantity | Value | Reading |
|---|---|---|
| Cell | one byte at a lattice site, 1 µm pitch | multi-byte values are adjacent cells, little-endian |
| Movement | 1 fJ per byte per micron of Manhattan path | the only priced quantity; arithmetic free |
| Propagation | c/160 = 1,874 µm/ns | 0.53 ns/mm (Dally's global-wire figure is 0.4 ns/mm = c/120) |
| Issue | 1 ns per instruction, no overlap with flight (time v1) | D4 |
| Tapes | input at (−1,0), output at (1,0), instructions at (0,−1); memory above the origin | serial, data-independent order, blocking receive only |
| Port crossing | E_in = E_out ≈ 5,000 fJ per byte | chiplet/HBM class (0.6 pJ/bit) |
| Fetch | 4 fJ per static line + 1 fJ per issued instruction | a loop buffer amortising fetch |
| Writes | free in v1 (D2) | the write of one op is the read of the next |
The calibration check that comes with it (added to the MNIST-on-grid doc 2026-09-04): one byte per grid location, 1 µm pitch, 1 fJ per byte per 1 µm Manhattan edge one-way, c/160. Applied to Dally's own 256 MiB benchmark, that machine gives 268 mm², 175 pJ per 64-bit access, 12 ns — against Dally's 100 mm², 58 pJ, 12.3 ns. So the adopted constants are 3× his energy and 2.68× his area, with latency matching almost exactly. That is the real "fitted to Dally" claim, and stating it in the spec is much stronger than "matched to a 256 memory bank", because a reader can check it.
And the constant that quietly decides the leaderboard: the endpoint term. Under the
current pure-distance rule (1 fJ × S) the best 8192³ schedule beats naive by 5.14×.
Under 2.5 fJ × (S + 160R) — i.e. adding a ~400 fJ fixed cost per paid read — the same
comparison collapses to 1.0034×. A fixed per-access endpoint charge (which is what every
real SRAM has, and what Dally's own 50 + 0.022√S law encodes as its additive 50 fJ/bit)
almost entirely erases the advantage the benchmark exists to measure. Whatever is chosen,
it should be chosen deliberately and published with the task, not inherited.
Your stated design method, which is worth writing into the spec (20260903-180729):
"I took the numbers to match the energy … of Bill Dally's model … match the energy of a 256 memory bank. And I chose the other to be within the same order of magnitude for time and area, while being easy to remember."
That is: energy is fitted; time and area are not fitted, they are mnemonic. Say so in the spec rather than letting a reader assume all three are derived.
4.2 Dally's published numbers — the calibration targets
CACM 65(9), Sept 2022 (On the model of computation: point), verbatim: - 32-bit add: 20 fJ and 150 ps. Fetching its two operands from main memory: 1.3 nJ — 64,000×. - Moving two 32-bit words 1 mm: 1.9 pJ and 400 ps. - 64 bits 40 mm corner-to-corner on a 400 mm² chip: 77 pJ and 16 ns. - Off chip: 320 pJ per 64 b, 6 ns per metre. - 8 KB (64 Kb) sub-array, 64 b access: 0.64 pJ, 300 ps. - 256 MB (≈100 mm²) from those sub-arrays: 58 pJ and 12.3 ns, of which 57.4 pJ and 12 ns are communication — 15 mm each way, 30 mm total. - CPU overhead turns a 20 fJ add into an 80 pJ add instruction. - PECM: "Each processing element and each memory is assigned a location l ∈ L and a distance function D: L × L ⇒ R … Two-dimensional Manhattan distance models on-chip communication … Off-chip communication can be modeled by adding additional distance when either coordinate spans a chip boundary." — note this model is explicitly parallel; the settled Grid VM is its P = 1 case.
CACM 63(7), July 2020 (Domain-Specific Hardware Accelerators), 14 nm: - 8-bit add 10 fJ and 4 µm²; double-precision FP multiply 5 pJ and 3,600 µm². - 8 KB local memory 50 fJ/bit, SRAM 0.013 µm²/bit. - On-chip communication 100 fJ/bit·mm. 100 MB memory ≈ 0.7 pJ/bit. - LPDDR4 ~4 pJ/bit; SDDR4 ~20 pJ/bit; off-chip SerDes ~10 pJ/bit. - And, crucially, the closed-form √-address law itself: the cost of accessing a memory of size S bits is 50 + 0.022·√S fJ per bit. That constant decomposes exactly: 0.022 = 2 (round trip) × 100 fJ/bit·mm × √(0.013 µm²/bit). Check at S = 800 Mbit: 50 + 0.022·28,284 = 672 fJ/bit ≈ the paper's own "0.7 pJ/bit for a 100 MB memory". ✓
This is the literature statement of the benchmark's pricing rule, in Dally's own words, and it is not currently cited anywhere in the repo. ⌈√a⌉ is not an approximation invented here — it is his formula with the additive endpoint term dropped.
AHA 2023 retreat keynote, in Dally's own preferred unit: - "Cost of an add (1 fJ/bit) = Cost of going 10 µm"; "An add is worth 10 µm of movement." Our grid charges 1 fJ/byte/µm = 0.125 fJ/bit/µm = 1.25 fJ per bit per 10 µm — i.e. 1.25× Dally's own figure. That is the tightest agreement anywhere in the constant set, and it is stated in the unit he prefers. - Instruction overhead, the axis that separates CPUs from TPUs: fetch + decode + operand fetch = 30 pJ; HFMA 1.5 pJ (2,000% overhead), HDP4A 6.0 pJ (500%), HMMA 110 pJ (22%), IMMA 160 pJ (16%). An out-of-order ARM A-15 integer add instruction: 250 pJ, ~4,000× the energy of the add itself. - Memory hierarchy per 64-b word: local SRAM (KB) 5 pJ, on-chip SRAM (MB) 50 pJ, LPDDR (GB) 640 pJ.
4.2b Dally's own number has moved, twice
The single largest unresolved discrepancy in the physical constants is not between us and him — it is within his own published work:
| Year | On-chip wire energy | Source |
|---|---|---|
| 2020 | 100 fJ/bit·mm | CACM 63(7) |
| 2022 | ~30 fJ/bit·mm (1.9 pJ / 64 b / mm = 29.69; cross-checks at 40 mm: 77 pJ/64 b/40 mm = 30.08) | CACM 65(9) |
| 2023 | ~100 fJ/b·mm | AHA retreat keynote, twice |
A 3.4× swing, same author, and node cannot explain it — Jouppi reports wire energy per unit length improved less than 2× from 45 nm to 7 nm.
This matters concretely, because our two derivations use different branches. The internal-log constant table (30 fJ/bit·mm, 400 ps/mm, 10 fJ/bit endpoint, 5 pJ/bit off-chip, 6 ps/mm) is CACM 2022 renormalised per bit. But 1 fJ/byte/µm = 125 fJ/bit·mm, which is calibrated against the 100 branch with a ×2 round-trip factor at a 0.61 µm byte pitch — not against the 30 branch that the same document quotes as its justification. Anyone re-deriving it from CACM 2022 gets a number 3.4× smaller and concludes there is a bug. Pick a branch, write the arithmetic into the spec, and state the round-trip convention.
Two further quoting hazards inside CACM 2022 itself: it says 1.3 nJ for a main-memory operand fetch in one paragraph and 1.2 nJ three paragraphs later; and it reports the same locality example as both 354× and 230×. Its 320 pJ is the crossing; its 640 pJ is the full DRAM access — do not cite them for each other.
4.3 The renormalisation that produced the working constants
Every line of your 03sep26 constant table is CACM 2022 re-expressed per bit:
| Your line | Dally source | Check |
|---|---|---|
| On-chip movement 30 fJ/bit·mm | 1.9 pJ / 64 b / mm | = 29.7 ✓ |
| On-chip latency 400 ps/mm | same move | ✓ (= c/120 exactly) |
| 8 KB SRAM endpoint 10 fJ/bit, 300 ps | 0.64 pJ / 64 b | ✓ |
| Off-chip 5 pJ/bit per crossing | 320 pJ / 64 b | ✓ |
| Off-chip propagation 6 ps/mm | 6 ns/m | ✓ |
Independently, CACTI 7.0.3 at 22 nm (tmp/cacti-sram8k-22nm.txt, 8 KB scratch RAM,
itrs-hp, 300 K, delay-optimal) gives: access 0.129 ns, read 2.048 pJ / 64 b = 32 fJ/bit,
sub-array 20.56 µm × 58.96 µm, data-array area 0.00683 mm² at 67.8% efficiency.
Wire at 22 nm: 280 fJ/bit·mm repeated full-swing, 2.84 fJ/bit·mm low-swing.
Two consequences worth stating in the spec:
1. "fJ per bit per mm" is undefined until the signalling style is fixed — the two CACTI
figures differ by ~100×, and Dally's own two papers differ by 3.4× (30 vs 100 fJ/bit·mm).
2. There is no 14 nm CACTI datapoint in tmp/ — the file named 8k-14nm.cfg actually
sets -technology 0.032. CACTI 7 tops out at 22 nm.
4.4 Technology node — unpinned
Internal log 02sep26: "use 7nm because Jouppi gives numbers / Dally also uses 7nm." But Dally CACM 2020 states 14 nm; the AHA 2023 slides use 28 nm for area/add and 45 nm for the instruction table; the CACTI runs are 22 nm. Nothing in the dictated notes names a node at all — only "conceivably implementable on existing CMOS process."
4.5 The pitch question, and the best available defence of 1 µm
The grid needs exactly one number: how many µm² does one byte of addressable memory occupy? Then pitch = √(µm²/byte). Every candidate, side by side:
| Basis | µm²/byte | ⇒ byte pitch | Node |
|---|---|---|---|
| Dally's stated SRAM area (0.013 µm²/bit) | 0.104 | 0.322 µm | 14 nm |
| TSMC N7 HD SRAM bitcell | 0.216 | 0.465 µm | 7 nm |
| Dally's 256 MB in ~100 mm² (array level) | 0.373 | 0.610 µm | not stated |
| TSMC 16 nm bitcell | 0.56 | 0.748 µm | 16 nm |
| Intel 22 nm SRAM cell | 0.736 | 0.858 µm | 22 nm |
| TPUv4i CMEM macro (128 MB = 28% of a <400 mm² die) | 0.834 | 0.913 µm | 7 nm |
| the adopted grid | 1.0 | 1.000 µm | — |
1 µm per byte is almost exactly the macro-level byte density of a shipping 7 nm inference chip's main scratchpad (TPUv4i's CMEM, 0.913 µm/byte). That is the strongest single justification for the choice and it is not cited anywhere in the repo — it is worth putting in the spec, because it converts "easy to remember" into "matches real silicon."
Two cautions that go with it: - Dally's 0.013 µm²/bit at "14 nm" is 5.4× denser than TSMC's 16 nm bitcell. It cannot be a bitcell figure; flag it as an idealisation. - Dally's 256 MB in 100 mm² is ~2.2× denser than TPUv4i achieves at 7 nm (2.56 vs 1.14 MB/mm²). Sizing the grid from his number is optimistic by that factor.
Sub-array cross-check: at 0.013 µm²/bit an 8 KB sub-array is 852 µm² of bitcell = 29.2 µm on a side, against the 55.2 µm square-equivalent the 100 mm² / 32,768 decomposition allocates it — ~28% array efficiency, which is plausible for SRAM with decoders, sense amps and an H-tree (CACTI measures 67.8% for a delay-optimal scratch RAM, so ~28% is on the conservative side).
4.6 The two softest constants
The 8 KB SRAM endpoint is the softest number in the model:
| Value | Source | Node |
|---|---|---|
| 10 fJ/bit (adopted) | Dally CACM 2022, 0.64 pJ / 64 b | not stated |
| 50 fJ/bit | Dally CACM 2020 | 14 nm |
| 117 fJ/bit | Jouppi ISCA 2021, 7.5 pJ / 64 b | 7 nm |
| 156 fJ/bit | Horowitz ISSCC 2014, 10 pJ / 64 b | 45 nm |
| 32 fJ/bit | CACTI 7.0.3, 2.048 pJ / 64 b | 22 nm |
10 fJ/bit is the most optimistic figure in the literature, by 5–12×. The plausible reconciliation: 0.64 pJ is bitcell + local bitline/wordline only, while 7.5–10 pJ is a macro access including decode, sense amps, periphery and the local H-tree. Say which one the model prices — a grid whose endpoint cost is 12× below the measured macro number will systematically under-reward locality relative to real silicon.
E_in = 5,000 fJ/byte is an HBM/chiplet crossing, not Dally's off-chip number. The repo's own correction, verbatim: "320 pJ per 64 b is 5 pJ per bit, 40,000 fJ per byte, a board-level crossing; the 5,000 fJ per byte the competition uses for E_in is a chiplet- or HBM-class crossing (0.6 pJ/bit), and the two must not be cited for each other." Cite HBM2 (250–450 pJ/64 b = 3.9–7.0 pJ/bit) for it, not CACM's 320 pJ.
4.7 Numbers available for free, if time and area become scored
Nothing needs deriving — the literature already supplies them: - Time: 150 ps (32-b add), 400 ps/mm (on-chip), 300 ps (8 KB sub-array), 12.3 ns (256 MB), 16 ns (40 mm), 6 ns/m (off-chip). - Area: 4 µm² (8-b add), 3,600 µm² (DP FP multiply), 0.013 µm²/bit (SRAM). - Node scaling, 45 → 7 nm (the only source giving the same quantity at two nodes, so the only basis for a "same program, different node" column): logic 2.1–4.3× (geomean 2.6×), SRAM only 1.3–2.4×, DRAM 6.3×, wire <2×.
One flag: the current time model's "1 ns per instruction, no overlap with flight" is an invention, not a literature number, and is ~3–7× slower than the 700–1,050 MHz TPU clock it is implicitly modelling.
4.8 What the HMM adds, beyond justifying ⌈√a⌉
docs/spatial-model-analysis.html already establishes that charging f(x) = x^α with
α = ½ is exactly [AACS 87]'s Hierarchical Memory Model, whose authors themselves motivate
it as 2-D physical distance. Its results at α = ½: scanning N inputs is Θ(N^1.5);
sorting and FFT are also Θ(N^1.5) (the log factors vanish); binary search collapses from
log N to Θ(√N); matmul is Θ(n³) for α < ½, Θ(n³ log n) at α = ½, Θ(n^(2α+2)) above —
i.e. α = ½ is precisely the critical exponent at which cubic linear algebra turns
movement-bound. Any time-T space-S RAM computation costs at most T·√S.
The one result with a direct design consequence is Thm 5.2: Block-LRU dynamic placement
is O(1)-competitive with the offline optimum. That says a compiler that relocates data
between phases is provably within a constant of optimal, and static single placement is not —
which is a real argument about whether release/stage should be the only relocation
mechanism, and it is measurable: the repo's own compiler audit puts the current hand layouts
at 1.41–2.79× over the provably optimal static assignment, where the worst assignment
would be 3.70–8.58×.
⚠️ That "1.41–2.79×" is the report's own summary line, but its own table lists 2.79, 2.78, 3.96, 1.68, 1.41, 2.46 — so the true range is 1.41–3.96×. The worst/optimal range (3.70–8.58×) does match the table. Fix the summary before quoting it.
4.9 Open forks on physical dimensions
- Which per-mm constant is normative (30 vs 100 vs 280 vs 2.84 fJ/bit·mm).
- c/120 vs c/160 vs c/1200 (global wire vs adopted vs full sub-array access).
- Pitch: internal log 02sep26 says outright "pitch size is still open" — 1 µm (adopted) vs 1.77 µm (your subarray derivation) vs 0.61 µm/byte (area-derived) vs 10 µm/bit (the other MNIST-doc section).
- The provenance of Dally's micron figures, which you flagged yourself.
- Whether time and area are scored or only reported — see §6.5.
5. Decision 4 — what competitors submit
5.1 What is submitted today
sparse-parity README, verbatim:
"A submission is one straight-line program (an IR) in the v3 instruction set: 8-bit cells, no loops, no branches, no data-dependent addressing, at most 2,000,000 lines (every line counts, declarations included). / An IR is plain text. Line 1 declares the input cell addresses … the last line declares the 32 output addresses, and each line between is one op…"
Delivery: "Submit a PR adding your .ir and generator under submissions/."
matmul adds: "the scorer (matmul.py) must not be modified."
So the de-facto practice is already the hybrid: .ir + a .md report + the Python
generator, with the generator outside the trust boundary. Andy's own sparse-parity records
are exactly this shape (optimize_layout(generate_sis_mask(1, 2))).
Current records (origin/main, 2026-09-08):
| Problem | Best | Date | By |
|---|---|---|---|
| matmul 4×4 | 675 | 2026-09-04 | @jurajselep |
| matmul 16×16 | 63,819 | 2026-09-07 | @SecurityQQ |
| sparse-parity 20% band | 86,753 | 2026-09-02 | @b0nce |
| 40% / 60% / 80% | 141,218 / 149,665 / 182,744 | 2026-09-01 | @npow |
| 100% | 392,666 | 2026-09-02 | @jurajselep + @b0nce |
| symmetry 6-bit / 8-bit | 20 / 29 | 2026-05-11 | baselines, no competitor entries |
(matmul/agent_prompt.md and directions.json are badly stale — they still quote
73,602 / 69,697.)
5.2 The wall that forces the decision
"Submissions … are straight-line programs … That choice makes energy a syntactic property — the judge scores by reading the program, never by running it … But it has a hard consequence: program length is proportional to work. A 10¹² -operation computation is a 10¹² -line program." —
docs/expressivity-scorability-ladder.html§1
1.7×10¹² ops as raw trace text is ~27 TB. Tier-2 MNIST is ~1.2×10⁷ trace lines; tier-3 is ~10⁹ lines, a 10–19 GB file.
Your own answer, internal log 04sep26:
"the submitted artifact has to be the generator of the straight-line program, not the expanded trace"
5.3 The two-track resolution — the actual proposal on the table
docs/expressivity-scorability-ladder.html §2 and RFC #48 §1:
Generators are unrestricted wherever the expansion is scored, and restricted exactly where the schedule is scored.
- Episodic track (tier 1, sparse-parity, matmul): submit the expanded trace. Generate it with anything — Python, an LLM, anything Turing-complete — because the organizer never runs it. This is already how the leaderboards work.
- Schedule track (tiers 2–3, 8192³ GEMM): the schedule itself is the submission, and its language must be restricted — for scorability, not sandboxing.
Made normative in the design report §10: "tiers 2 and 3 are schedule-track only, declared in the rules from day one; tier 1 runs dual-track and serves as the trace-vs-schedule cross-validation corpus."
5.4 The ladder — the vocabulary for arguing about this
| Rung | Generator may contain | Scoring | Unlocks |
|---|---|---|---|
| a | nothing (straight-line trace) | O(L) summation | episodic tasks |
| b | affine loop nests, static bounds | closed form | GEMM, conv, transformer fwd/bwd |
| c | + static recursion (compile-time depth) | recurrence relations | Strassen, recursive FFT, fractal tiling |
| d | + bijective index maps from a fixed library | free (permutation invariance) | FFT butterflies, bitonic networks |
| e | + mode: data-dependent selector among k static schedules |
worst-case over arms stays static | MoE with capacity factor, early exit |
| f | + while / data-dependent trip counts |
undecidable — simulation only | convergence-based iteration |
"The cliff is between (e) and (f), and it is a theorem, not a taste: once trip counts depend on data, boundedness — and with it cost — is undecidable (Buck)."
Rung (c) is not optional: "an energy leaderboard for matrix multiplication that cannot express the sub-cubic algorithms would prejudge the very question it poses." Rung (d) is free because Σ⌈√σ(i)⌉ = Σ⌈√i⌉ for a bijection σ — permuting which iteration touches which address cannot change the address multiset.
5.5 The recommended stack, and what was rejected
Three routes were compared:
| Route | Verdict |
|---|---|
| A — GridASM, flat opcode text, superset of today's IR | keep as the SIR level — the address is the cost, visible to writer and reviewer, no compiler in the trust boundary |
B — structured DSL (place W : i8[10][9] at (4,0), affine for, recursive def with a kind system, mode, recv/send) |
adopt surface ideas; lowers to a one-page Schedule IR which is the scored artifact |
| C — reuse survey | decisive negative result: no existing system computes this scoring function. MLIR affine is the right vocabulary but a multi-GB LLVM build to parse 40-line files; Halide and TVM cannot express static recursion (so no Strassen); Wasm/RISC-V solve a sandboxing problem this design does not have. Winner inside C: an embedded Python builder library emitting canonical schedule text. |
The recommendation is B ⊕ C: a Python builder as the authoring layer (outside the
trust boundary, eager validation — a non-affine index is a TypeError on the submitter's
machine), emitting canonical schedule text styled after MLIR-affine notation as the
reviewed and scored artifact, consumed by a Rust judge of ~2–3k new lines atop dally-eval.
Hand-written schedule text stays a legal submission, so the trace-tuning culture is not
orphaned. "In every route surveyed, the grammar itself is the rung enforcement: while is
unrepresentable — verification by absence, the cheapest verifier there is."
Explicitly rejected: - Python-script submissions with instrumentation — "allowed too many ways to cheat; unrestricted hosts are not sandboxable" (RFC #48 §1; you reached the same conclusion in Telegram five months earlier). - Traces alone at scale — the 27 TB wall. - Compiler-chosen placement — "under the Dally model placement is the algorithm".
5.6 What Andy has actually built, and what he is asking you
Andy Zhang (@zh4ngx) has eight open PRs, all awaiting you. You have left zero comments on any of them.
| PR | What | Needs you? |
|---|---|---|
| #48 | RFC: the Scheduled Dally Language — the whole submission-format answer | yes |
| #58 | MNIST tier-1 (3×3) demo: 280 ops, cost 4,625, ~0.005 nJ/inference, float 53.0% = int8 53.0% | yes |
| #61 | Chebyshev-vs-Manhattan provenance audit | yes — see §6.1 |
| #49 | three-engine benchmark: Rust CPU 69–77× Python; GPU LDS 0.38–0.90× Rust CPU on static traces, but 2× on search workloads | reference |
| #51 / #52 / #55 / #56 | matmul search runs (858 floor), access heatmaps | reference |
| #60 | BUILD_NOTES.md + reviewer triage |
reference |
His four named forks, verbatim from PR #48's body: "MNIST data representation, accumulate semantics, reduce primitive, home repo."
His framing of your Sep 2 ask, verbatim from RFC §1: "The request (Yaroslav, 2026-09-02): a language that covers both matmul and MNIST, is immune to cheating (interpreted in a sandbox), and makes memory placement an explicit part of program design."
5.7 What your dictated notes say about this axis
Nothing directly — the words "submission", "instruction stream", "IR", "language", "SDL", and "PR" do not appear in the 14-day window. Two indirect signals, both pointing at generator, not trace:
"I can just generate this competition, and then people will create these data flow algorithms that do learning and prediction, and I'll want to understand how they're doing it, but probably not backprop." —
20260903-183033(opens in Russian: "Right now I think that I have seen how to get rid of backprop")
A flat trace is precisely the artifact you cannot understand.
"what if we reinvent all of linear algebra, all of algorithms for a new abstraction? And then we have compilers" —
20260904-125521
And on who builds it:
"maybe with Andy I … figure out how to create a data flow computer that does this." —
20260903-073220
6. Cross-cutting decisions you will need anyway
6.1 The metric: Manhattan or Chebyshev — currently self-contradictory on main
This is the sharpest live inconsistency and it sits in published documents.
docs/grid-vm-competition-design.html§5, still on main, uncorrected: "distance is Chebyshev, d(x,y) = max(x,y)", and decision D3: "native 2-D Chebyshev with canonicallinlegacy embedding." Rationale given: max-of-affine is piecewise-affine, so rung-(b) sums become lattice sums instead of √-staircases.- RFC #48 §8.3: keeps 1-D
lin(a)for v1 for bit-exact record continuity; Chebyshev deferred to a v2 spatial division. docs/grid-vm-multiprocessor.html(Sep 7, merged) silently adopts the correction: "1 fJ per byte per micron of Manhattan path", and "c/160 = 1,874 µm/ns".- PR #61's evidence: across 2,588 Telegram messages (2026-02-09 → 2026-09-05),
zero contain "Chebyshev". Your own messages consistently say Manhattan (msg 1277,
2026-03-30: "Bill Dally advocates a model based on Manhattan distance…"). First repo
occurrence is commit
4f4d3e2— the 14-agent design run — from which it propagated into PR #48's text.
The identity that settles it, and it is exact:
| Shell at radius c | Cell count |
|---|---|
legacy {a : ⌈√a⌉ = c} |
c² − (c−1)² = 2c − 1 |
upper-half-plane Manhattan {|x| + y = c, y > 0} |
2c − 1 ✓ |
first-quadrant Chebyshev {max(x,y) = c} |
2c + 1 ✗ |
So Manhattan is the exact discrete realisation of the existing ⌈√a⌉ law — no substitution
is needed. The explicit embedding: k = ⌈√a⌉, t = a − (k−1)² − 1, x = −(k−1) + t,
y = k − |x|. Physically, Chebyshev underprices a rectilinear diagonal route by up to 2×
((d,d) costs d under Chebyshev, 2d under Manhattan), which breaks the 1 fJ/byte-micron
calibration.
Also uncorrected on main: the competition design report's time formula prints
1.875 µm·ns⁻¹ where c/160 = 1,874 µm/ns — a factor of 1,000.
And a second, subtler metric error that PR #61 also flags. RFC #48 §8.1 charges "each
edge … ceil(sqrt(distance)) once." But ⌈√a⌉ is the function that converts a linear address
rank into the radius of its packed 2-D shell — applying another square root to an already
geometric distance conflates rank with route length. Under native placement the correct form
is simply
E_move = Σ_reads bytes(r) · d₁(src, dst) · 1 fJ/byte-µm
This one is easy to miss because it only bites once addresses become coordinates — i.e. exactly when you adopt native 2-D. The correction was drafted in PR #61 and never posted.
6.2 The scoring function
- Energy: `E = Σ_reads d(p) + 4 fJ · static lines + 0.1–1 fJ · issued instructions
- n_recv·E_in + Σ_sends (d(s) + E_out)`.
- Closed forms, honestly stated: single-variable progressions use the ⌈√⌉ antiderivative (RFC #48). Multi-coefficient affine addresses give Ehrhart quasi-polynomials with floor-sum corrections, not pure polynomials — still exact and fast via banded lattice counting, ~10⁴ integer evaluations at A_max = 10⁸ — "but the spec must name the floor-sum evaluator as normative rather than claim 'O(1) polynomial.'"
- Cross-validation is the anti-divergence rule: for every program small enough to
expand (~1 M ops), the closed-form score must equal
dally-eval's static score of the expanded trace, bit-exact. - Judge cost: 0.1–4 ms per submission at every tier. Adjudication: ≈0.5–4 s (tier 1), 10–20 s (tier 2), 2–18 min (tier 3, JIT; LeNet worst ≈1 h). A 20-finalist tier-3 re-score is a ~3 h overnight job on an 8-core machine. A cluster is never needed.
- Fuel is free: executed-op count is itself closed-form, so over-budget submissions are rejected without being run — evaluation DoS is structurally impossible.
- One flag set consciously: the tier-3 laptop fuel cap has ~3× headroom over a plain MLP but below 1× at a realistic 4–6 byte-ops/MAC — as drafted it can exclude honest LeNet training, unlike the 10× margins at tiers 1–2.
6.3 Anti-cheat — the decision that shapes the competition
Your stated policy, 20260903-090056, which is worth writing into the rules verbatim:
"the reward hack just means they do something and somebody looks at this and they're like, there is no way I can use this technique in real life. … It's like, hey, that's a cool technique, [re]use it. And they hard code — it's like, well, you can't reuse it. And we want to isolate as many reward hacks and fix it on the competition. And I think for the integrity of the competition, allow any other reward hacks because ultimately it's subjective. So it's unfair to the participants to change the rules in the middle of the game."
Decoded: the test is subjective ("could I reuse this?"); enumerate and close hacks up front; then allow the rest and never change rules mid-game.
Reinforced at the Sep 7 meeting: "the goal is algorithms … they love to cheat."
Mitigation triage from the red team (design report §7):
| Mitigation | Fate | Why |
|---|---|---|
| Held-out QMNIST pool | fails | public since 2019; attacker pretrains on the union |
| Fetch-priced immediates | fails | kills Θ(K) table bombs, not 80 KB of compressed hypothesis |
| Program-length caps | fails | any cap excluding 8 KB of sets excludes honest inference |
| Per-episode label permutation | keep, insufficient alone | recoverable by vote-table matching (25–30% tax at tier 1, <1% at tier 3) — but it forces Ω(100) train reads, killing zero-read programs |
| Per-episode pixel permutation | bypassable | ~8% tax to attacker, ~2× tax to honest convs |
The three options that actually work, to be chosen explicitly: (a) accept it — rename the closed division episodic inference energy, run in-episode learning separately; (b) procedural episode generation (MNIST-1D-style parameters redrawn per episode, so no finite pool exists); (c) per-episode invertible GF(2)-affine mixing — makes recovery equivalent to learning but destroys spatial priors and every calibration number. Recommendation: (b) for tiers 1–2; decide tier 3 after deciding whether convolutional architectures are a protected species.
Note: RFC #48 §9.1 claims D1 is "Resolved via procedural episode generation (S7) for Tiers 1-2"; the design report leaves tier 3 open. The two documents disagree about whether D1 is closed.
Seeds must be keyed, not merely derived:
seed_j = SHA-256(server secret ‖ submission bytes ‖ tier ‖ j). Without the secret term a
submitter simulates every episode offline and ships the answer sequence as immediates —
100% accuracy at near-zero energy. This was found as a fatal hole in the first draft.
K = 50 / 10 / 3 shrinks the mined maximum to ~+0.8 pp at tier 1.
Two attacks recorded as dead: the tier-1 "Bayes table" (59,725 of 60,000 3×3 images are distinct at 8-bit depth, so a top-1000-pattern table scores ~12%), and pool-lookup by linear scan (~6×10¹² trace lines at tier 3, fetch-priced into oblivion).
Your own train/test idea (20260904-092239): "I just have a pool of examples, and I
resample both training and test sets." Note the tension with the sparse-parity mask tier,
which deliberately has no test set.
Nine pitfalls are already published (docs/index.html §Pitfalls) and were learned the
hard way on sparse-parity. A redesign must not lose them:
- "The public suite is minable — never adjudicate on it." Demonstrated at n=12: an exploit passed the η ≥ 0.25 scorer at lower cost than the honest entry while its true advantage was 0.223.
- "Sampled secrets must be fresh or hidden." A ~33,000-instruction circuit (13% of the cap) enumerating only the dev suite's 128 sampled secrets "scores a perfect η = 1.0 on the dev suite — Pareto-dominating the honest frontier — while measuring exactly 0.0 on a fresh key."
- "Output truncation must be priced, not banned" — via the decode/predict energy balance; set m_test ≈ D/P of the reference decoder and re-derive it per tier.
- "Candidate enumeration is not the only brute force." Both exponentials must be checked against the cap. m_train = 18 at n = 32 was chosen specifically so the full 2¹⁴ scan exceeds it.
- "The instruction cap is part of the contract. Energy alone cannot kill brute force… Raising the cap is a benchmark-version change, not a tuning knob."
- Accepted noise is ±4 pp across suite keys — "Energy is static and exact; only the accuracy coordinate is noisy."
- Publish pooled curves with spread, not single deterministic points.
- Complement pairing needs odd k (both live tiers use k = 3, k = 5).
- "No data-dependent addressing cuts both ways" — submissions cannot hash, but the generator is ordinary Python.
Points 1, 2 and 5 are the ones that transfer directly to MNIST, and 5 is the one most likely to be forgotten: the line cap is a rule, not a tuning knob.
6.4 Accuracy targets — measurement overturned the drafts
An adversarial reviewer ran real 600/600 episodes rather than extrapolating, and two recommendations flipped:
| Tier | Measured | Draft target | Corrected target |
|---|---|---|---|
| 1 (3×3) | centroid 48.2 ± 2.1%, linear@600 50.5%, 1-NN@600 54.0% | 50% (would kill centroid) | ~45% |
| 2 (9×9) | linear 84.8% on full data | 88% (would exclude linear) | ~80–82% |
| 3 (28×28) | 1-NN 96.9% | — | 97.3%, and "never place a target at 97.0" (works only because the task pins a 60k test set, σ = 0.07 pp) |
Per-episode binomial noise: ±2.0 / ±0.5 / ±0.09 pp.
The resident-vs-restream lever, genuinely contested. Holding the 60k×784 training set resident costs (2/3)·(4.70×10⁷)^1.5 ≈ 2.2×10¹¹ units per sweep; re-streaming costs 4.70×10⁷·E_in. Resident wins iff E_in ≳ 4.6×10³ — which is why E_in ≈ 5,000 fJ puts the two régimes within ~10% of each other and makes this a real strategic choice rather than a foregone one. "1,000 fJ is wrong."
Stream floors never decide a ranking: 11.4 kB / 0.98 MB / 94 MB in-streams cost 5.7×10⁷ / 4.9×10⁹ / 4.7×10¹¹ fJ — at tier 3 that is ~10⁻³ of MLP training energy.
6.5 Time and area
Internal log 06sep26: "report three things — energy, time, area, accuracy" (four listed).
- Time v1 = sequential issue:
T = N_issued·1 ns + Σ d(p)/1.875 µm·ns⁻¹ + port terms(constant is wrong by 1000×, see §6.1). "The honest consequence: TIME correlates with ENERGY and tier-3 sequential time is measured in days." Tier-3 MLP 784-256-10 at 30 epochs ≈ 9.5 days; 1-NN anti-strategy ≈ 240 days. - Time v2 = dataflow critical path (Karp–Miller–Winograd affine schedules) is well-posed and closed-form, but bijection invariance fails there (a bit-reversal stretches wires — physically real), so library maps would need per-map dilation lemmas. Run as a separately scored division; do not blend it into v1.
- Area: origin-anchored bounding square of touched cells in µm², unioned over mode arms.
- Nothing is currently scored except energy. The c/160 constant is unused by any scorer.
docs/dataflow-io-order.htmlnames the gap: "Nothing rewards parallelism. Order-fixed, time-free, energy-only scoring never penalises an arbitrarily serial submission — an odd property for a benchmark about a spatial machine." - Unanswered and important: with four axes (energy, time, area, accuracy), how is a winner defined? The design report says accuracy gates entry and energy/time/area get separate leaderboards. Nothing says how they combine, or what happens on a tie.
6.6 Parallelism — decided in outline, P1–P8 open
docs/grid-vm-multiprocessor.html compares four ways to put more than one control unit on
the grid, and recommends an adoption order:
- SPMD tile mesh with a collective library — adopt first. Smallest delta; keeps the judge a formula; makes parallel makespan, cross-core locality and idle-silicon cost visible. Every message-passing program using only collectives scores identically.
- Systolic coprocessor inside a tile ("AI-CPU tile") — the only scenario that prices Roune's thesis; closed-form; first showcase is 8192³ GEMM.
- Message-passing multicore — needs the reference simulator first.
- SIMT lane array — its own P ≥ 2 division.
Eight open decisions P1–P8, each with a stated cost: tape rate for P ≥ 2 (8 B/ns vs 1),
fetch continuity, whether static energy A·T is scored, the idle-clock term (0.125 fJ per
bit-micron of clock tree per ns = 168 fJ/ns per tile), the area rule, write pricing,
whether SIMT/dp4a enters v1 ("adopting dp4a into v1 drops the 16×16 record 4×
overnight"), and makespan on the general path.
Two migration items it flags as settled-but-unwritten: the port term (a recv.T occupies
sizeof(T) ns at 1 B/ns with no issue slot) and the area function ((2k−1)² over the
lin(a) half-diamond). And one correction: "5 pJ per byte" should be per bit.
6.7 What each proposed change costs the existing leaderboard
Every option has a stated re-scoring price. Collected in one place, because this is the
constraint that has silently driven several design choices (D2 "writes free", D3 "keep
lin(a)", RFC §8.3) toward compatibility rather than physics:
| Change | Records it moves |
|---|---|
Manhattan lin(a) embedding, 2-D exposed |
none — bit-exact |
| Native (x,y) placement | requires a versioned leaderboard even under Manhattan |
| Pricing ports | "once ports are priced, legacy free-input records are not score-comparable anyway" |
| Instruction-fetch pricing | every record; none has been re-scored under it |
| A located loop buffer (P2) | every P = 1 record by 2–24% |
| Charging the idle clock at P = 1 (P4) | every P = 1 energy record by 7–80% |
| 8 B/ns tape at P = 1 (P1) | every P = 1 time record (streamed sum 2,102,751 → 1,447,809 ns) |
Adopting dp4a into v1 (P7) |
drops the 16×16 record 4× overnight |
| Priced writes | both scorers charge reads only today |
The corpus PR #48 §8.3 promises to keep bit-exact is "sparse-parity, 4×4 matmul at 675, 16×16 matmul at 64,431, and the Tier-1 PR #58 demo" — note 16×16 has since moved to 63,819, so that clause is already stale.
6.8 Hosting, and the shape of the competition
- Today: a GitHub repo with PR-based submissions. Nothing decided beyond that.
- Your own suggestion, Telegram msg 1238 (2026-03-27): "near.market is a good fit for this kind of competition." You had already posted the red-blue pebble game as a near.market job (msg 314).
- Also floated: Prime Intellect RL hub / Anthropic RL envs (msg 1092: "Ben Mann seemed interested"), HuggingFace Spaces leaderboard (Yad).
- Andy's fourth fork is literally "home repo" — where SDL lives.
6.9 Two things you have said before that a redesign must not lose
The metric has been silently wrong twice. Both times a leaderboard leader was not counting its own work: - Yad, msg 1102 (2026-03-22): "GF2's measured DMC of 8,607 was artificially low. The harness only tracked write(A) → read(A) → write(solution), skipping the O(n²) row operations. Honest tracking puts GF2 at 189,056." - You, msg 1231 (2026-03-27): "It may affect the existing leaderboard because the Gaussian elimination didn't count the cost of the elimination step."
A new ISA should be judged partly on whether it makes this class of error impossible — which is exactly the argument for execution-free scoring over instrumented Python.
Iteration speed is a hard design input, repeatedly stated. msg 689 (2026-03-04): "baseline SGD 22 seconds is too slow. It should be <2 seconds. Making it 22 seconds makes your iteration time 10x slower." msg 925: "Only consider experiments runnable on 1980s hardware in 1 hour (less than 1 second today)."
And the benchmark-overfitting warning, from you (msg 767, 2026-03-09): "keep in mind the AI Radiology debacle where a team of (human) agents optimize the heck out of a benchmark, but the result didn't really work … Unless your evaluation is the actual thing you care about (extremely rare), [you] need to be exercising human judgement to tell if the direction your agents are moving is likely to make impact."
6½. Decisions nobody has named yet
These are not open questions — they are choices that have already been made by accident, or that a competition needs and this one does not have. Grouped by how badly they bite.
A. Made implicitly by code, never chosen by anyone
| The accidental decision | Why it matters | |
|---|---|---|
| Free input placement | The submitter chooses input addresses and placement is free (both scorers). | At tier 3 that is free placement of 4.7×10⁷ cells — an enormous unpriced degree of freedom. The Grid VM report concedes the consequence in passing: "once ports are priced, legacy free-input records are not score-comparable anyway." |
| No registers | addr = 0 is the off-grid ALU; address 1 costs 1. |
Whether the machine has a register file, and whether register traffic is free, was settled by an address convention rather than a design choice. |
| Free fan-out / broadcast | Writes are free, so one read feeding many consumers is free. | Broadcast is unpriced, which silently sets its complexity. The dally-spatial-report notes that pricing writes "restores broadcast to Θ(n^1.5)". |
| Store always beats recompute | Free writes. | Deletes Dally's headline example (§2.6). |
| Two engines, one silent fallback | dally_eval.py returns None when the Rust binary is missing and the harness silently falls back to Python; DALLY_EVAL_BIN is environment-overridable, with no pinned version and no engine identity recorded per score. The vectorised engine computes in int16 while the scalar engine wraps to signed 8-bit — they are asserted to agree, not proven across all ops. |
A scoring harness that can quietly change engines is exactly the failure mode that produced the GF(2) 8,607 → 189,056 correction. |
B. What is actually being measured — the biggest unnamed fork
- Is this a learning benchmark or an inference benchmark? The Grid VM report says that if embedded weights go unmitigated, "the 'learning' axis is dead code and the benchmark measures inference circuits", and its fallback is to rename the closed division episodic inference energy. Your voice notes say the target is "an actual MNIST training loop." PR #58 trains in numpy, off-model. "Training-in-the-IR vs training-off-model" is never posed as a decision anywhere.
- Is accuracy computed inside the model and priced? Today it is not: the harness computes it, and PR #58's argmax runs off-model in the pre-wrap integer domain because 8-bit cells alias int16 logits.
- What else may legally happen off-model? PR #58 set the precedent by accident. Without a rule, "off-model" is an unbounded escape hatch — and it is the same shape as the reward hacks the pitfalls list already catalogues.
- Does the practical track survive? PTX-on-real-GPUs, NVML vs RAPL, "total system energy under time and accuracy bounds" was decided in Apr/May and has been dormant since. The redesign never says whether it still exists.
- Does MNIST-on-grid serve the stated week-52 final exam? The exam is "demonstrate energy-efficient training of Karpathy's nanoGPT … without sacrificing wall-clock time and accuracy." No document connects the MNIST ladder to it.
C. Governance — the competition as an institution
Almost entirely absent, and cheap to fix:
- There is no
LICENSEand noCONTRIBUTINGonorigin/main. Entrants submit code by PR into a repo with no licence, no CLA, no copyright grant, and no statement of what rights anyone has. (SutroYaro was deliberately made Unlicense; sutro-problems never was.) - No prize, deadline, eligibility rule, anti-sybil measure, submission rate limit, or appeals process.
- The conflict of interest is unaddressed: you merge every PR and hold leaderboard records.
- Human-vs-AI provenance disclosure existed once (Sutro #12 agenda, "human-vs-AI provenance rule") and vanished. PR #60 volunteers a build footprint; nothing requires it.
- No deprecation/migration policy, and precedent runs both ways — sparse-parity already carries a "Legacy page" of retired tiers, while PR #48 argues for bit-exact preservation.
- Discoverability regression: commits
9e3175e/5e72dbbremoved the "Design notes" section from the root README, so all four live design notes are published on Pages and linked from nothing. The README also lists only two of the repo's four problem directories. - Your reward-hack policy is not written anywhere a participant can read it. It exists only as dictated audio (§6.3). It is the most complete policy statement in the corpus.
D. Ranking — four axes and no rule
- How is a winner defined? Options nobody has written down: lexicographic; energy at fixed accuracy (today's sparse-parity band scheme); Pareto frontier with ties unbroken; AT² (the theoretical backdrop — Thompson, and the Ryan Williams thread in your April log — never adopted); energy·time; energy under time and area caps.
- How do you rank on a noisy axis? Energy is exact; accuracy carries ±2.0 / ±0.5 / ±0.09 pp.
The only escalation ladder anywhere is in
chatgpt-report.txt, written for sparse-parity. - Area has no unit and no rule in any scorer. Bounding-square is proposed only in the Grid VM report; touched-cell count is explicitly gameable by scatter.
- Nothing versions the metric.
v0…v3versions the ISA. Your PR #18 comment is the closest thing to a metric-versioning rule and it is buried in a May PR thread. - The compatibility claim and the proposed pricing changes are mutually exclusive, and both live in the same document set (§6.7).
E. Machine-model forks still unnamed
- DRAM as a tier. Off-chip is priced in exactly one place (5 pJ/bit crossing, 6 ps/mm) and nothing in the current model crosses a chip boundary at all. Capacity, bandwidth ceilings (400 GB/s off-chip, 1 TB/s interposer), refresh, and page/row structure are all absent.
- Leakage and idle silicon. Static energy is reported not scored (P3, and "the constant has a ±10× band"); clock energy is scored only for P ≥ 2 (P4). Both make an idle grid free — so at P ≥ 2 the energy board's optimum is P_max unless the clock term lands.
- Operating point. Voltage, frequency and temperature are the entire basis of the constants (E = ½CV² at V ≈ 0.5 V) and are never declared.
F. MNIST-specific gaps
- No lower bound and no reference implementation for any tier — so there is no way to detect saturation, which was the stated motivation for the redesign in the first place.
- No rule on dataset representation. PR #58 chose signed 6-bit features from block averaging; that was the submitter's choice, not a spec.
- Whether convolutions are "a protected species" — the Grid VM report defers tier-3 D1 on it.
- Is tier 3 even runnable? Modelled sequential time is 2.3–240 days, and the laptop fuel cap is "below 1× at a realistic 4–6 byte-ops/MAC — as drafted it can exclude honest LeNet training."
7. Contradictions ledger
Things that disagree in writing, right now. Each needs a ruling.
| # | The disagreement | Sources |
|---|---|---|
| C1 | Chebyshev vs Manhattan. D3 says native 2-D Chebyshev; the Sep 7 note and RFC §8.3 say Manhattan / keep 1-D lin(a); PR #61 proves Manhattan is the exact realisation (2c−1 vs 2c+1). Uncorrected on main. |
design report §5/§11 vs multiprocessor §2 vs RFC #48 §8.3 vs PR #61 |
| C2 | 1.875 µm·ns⁻¹ vs 1,874 µm/ns for c/160 — 1000× |
design report §5 vs multiprocessor §2 |
| C3 | 1 picojoule vs 1 femtojoule — resolved: a misspeak, not a fork. The dictated design description does say "the idealized computer is 1 picojoule, 1 micron" (verified across engines), but the same 134-word clip also says a memory access "takes milliseconds" (true value 0.5 ps), and every written source of that same week — the MNIST-on-grid doc (Sep 4), meeting #30 (Sep 8), RFC #48 §8.2, the multiprocessor note — says 1 fJ per byte per micron. Treat the clip as dictation error. | 20260903-180729 vs meeting #30 / RFC §8.2 |
| C5 | 30 vs 100 vs 280 vs 2.84 fJ/bit·mm. Dally's own figure went 100 (2020) → 30 (2022) → 100 (2023) — a 3.4× swing that node cannot explain (wire energy improved <2× from 45→7 nm). CACTI 22 nm repeated wire = 280; low-swing = 2.84. Our two derivations use different branches: the constant table is the 30 branch, but 1 fJ/byte·µm = 125 fJ/bit·mm is calibrated to the 100 branch with a ×2 round trip at 0.61 µm. Biggest single unresolved number. | Dally CACM 2020 / 2022 / AHA 2023; internal log 03sep26; tmp/cacti-sram8k-22nm.txt |
| C5b | 8 KB SRAM endpoint: 10 vs 32 vs 50 vs 117 vs 156 fJ/bit. The adopted 10 fJ/bit is the most optimistic in the literature by 5–12×; probably a bitcell+local-wire figure being used where a macro figure belongs. | §4.6 |
| C6 | c/120 vs c/160 vs c/1200. 400 ps/mm = c/120 exactly (global wire); c/160 = 533 ps/mm (adopted, "conservative"); c/1200 for a complete sub-array access; meeting #30 says 0.5 ps/µm = c/150. | internal log 03sep26 vs meeting #30 vs RFC §8.2 |
| C7 | 1 µm/byte vs 1.25 µm/step vs 1.77 µm/pitch vs 0.61 µm/byte vs 10 µm/bit. Five pitches in circulation. binary-matmul-regimes/derived.json and a100-grid-energy-report/grid-audit.json both derive 1.25 µm per step at 1 fJ per int8 value from 100 fJ/bit·mm (and 10 µm if the fJ is per bit); the multiprocessor note splits the difference at 0.98 fJ for a byte-wide round trip at the 0.61 µm area-derived pitch. And 1 byte/µm² is 800× the 10,000 bits/mm² figure in the same MNIST doc. |
meeting #30 vs internal log 03sep26 vs multiprocessor §3 vs MNIST-on-grid doc vs derived.json |
| C8 | "5 pJ per byte" vs "5 pJ per bit" off-chip. 320 pJ/64 b = 5 pJ/bit = 40,000 fJ/byte (board-level); the 5,000 fJ/byte used for E_in is a chiplet/HBM-class crossing (0.6 pJ/bit). "The two must not be cited for each other." | multiprocessor §3 (explicit correction) |
| C9 | Technology node: 7 nm vs 14 nm vs 22 nm vs 28/45 nm. You wrote "use 7nm because Jouppi gives numbers / Dally also uses 7nm", but Dally 2020 says 14 nm, AHA 2023 uses 28/45 nm, and the CACTI runs are 22 nm. There is no 14 nm CACTI datapoint (8k-14nm.cfg actually sets 0.032). |
internal log 02sep26 vs the papers vs tmp/cacti-* |
| C10 | Is D1 (embedded weights) closed? RFC §9.1 says resolved via procedural episodes for tiers 1–2; the design report leaves it as the open decision that defines the competition. | RFC #48 §9.1 vs design report §11 |
| C11 | 256 MB vs 256 MiB for the Dally anchor. | CACM 2022 vs meeting #30 |
| C12 | "report three things — energy, time, area, accuracy" — four listed. | internal log 06sep26 |
| C13 | PR #58 claims it emits "the exact op set the matmul competition scores" but uses set, which v0's matmul.py rejects, and re-implements static_cost(). |
PR #58 vs matmul/matmul.py |
| C14 | Meeting time 18:00 vs 18:15. | top-level doc vs meeting #29 doc |
8. What is not written down anywhere
Genuine gaps, not disagreements.
- The grid's actual size. Per-tier memory caps exist (1/16/256 MB); a cell count, address-space bound, or "the machine is N×N" statement does not. Your dictated notes never mention it once.
- A migration story from the existing leaderboards to the new machine. The Sep 3
dataflow/MNIST rethink appears in your notes as a fresh start. The matmul and
sparse-parity records, the
.irformat, and the energy numbers are not discussed at all in the 14-day window — even while you were actively producing sparse-parity results. Your own PR-#18 precedent (new IR → new problems only) would resolve this cheaply if stated. The one useful outside analogy is Jouppi's Lesson ⑩: Google deliberately chose source-level ("backwards ML") compatibility over binary compatibility, precisely so the hardware could change underneath the programs. Applied here, that says: when the competition switches from traces to SDL, re-score the programs, don't preserve the traces — which is the opposite of what "bit-exact record continuity" (RFC §8.3, D2, D3) is currently optimising for. Worth deciding on purpose rather than by default. - Cited artifacts that do not exist. Verified absent from
mainand every remote branch: - "Proposal: a streamed instruction processor and a compact schedule language, this repository, September 4, 2026 — the four-port geometry, the canonical instruction encoding, charged reads and writes." The only cited source for the four-port geometry and for charged writes. -scoring-at-scale/— cited in the Grid VM report as the source of the "0.3–3 ms to score a 1.1×10¹²-instruction schedule" measurement. The headline judge-cost claim has no artifact. -matmul/a100_8192_dtype_energy_results.json— cited as the measured A100 anchor;git log --allshows it has never existed. The content survives asa100-grid-energy-report/measurement-audit.json(which names it assource_artifactwith its sha256), and the 0.904 J INT8 figure is verifiable there. - "S7" — cited by RFC #48 §9.1 as the resolution of the embedded-weights fork. Defined nowhere in any repo document. -docs/pr48-response-draft.md— cited by PR #61, not in the repo. - No SDL parser or interpreter exists in any PR. RFC #48's security claim ("no I/O, no host calls, no unbounded allocation") is a specification, not code. - What "256 memory bank" refers to in your design description, and which Dally table
the fJ anchor was fitted to. (
sutro fJ reportis the candidate; it is only 545 bytes.) - The provenance of Dally's micron figures — your own inline note "(where is the micron measurement from?)" is still open.
- Number format. Bytes everywhere; no statement on precision, signedness, or
accumulate width, beyond the design report's
u8/i16/i32proposal. - Whether time and area are scored or only reported, and how a winner is defined across four axes.
- Off-chip / DRAM pricing. Only the pair "5 pJ/bit per boundary crossing" and "6 ps/mm" exist. Tier 3 has no meaning without it (PR #58 says so explicitly).
- Determinism of the eval suite.
chatgpt-report.txtflags that the current scorer usessecrets.SystemRandom()— "good for unpredictability, but not for deterministic leaderboard comparisons" — and recommends a public deterministic suite (committed seed) plus a private one (fixed server-side key), with "Do not regenerate the final suite on every scoring run." - Tie-breaking.
chatgpt-report.txtproposes an escalation ladder (smoke 1k → dev 10k → leaderboard 50k → strong 100k → final 250k → tie-break 1M → audit ~2M) with "if two submissions are within ~0.3 pp, rerun those only at 1M–2M tasks." Not adopted. - The recent Telegram record is not on this machine. Everything Aug–Sep lives on the
intelMac. PR #61's diff is the only local window. - The Sep 8 memory-wall thread has not reached the design.
20260908-122822: "I got really nerd sniped by the numbers";20260908-141736: "there's some useful takeaways in terms of why we have the memory wall." Worth folding in before fixing axis (3). - The Sep 7 meeting produced "a hierarchy of three" (
20260907-163821) — probably the three MNIST tiers, but the clip does not say so, and meeting #30's own notes are thin. - Load-bearing decisions that live only in private, uncitable threads. These cannot go
into a public spec, and several are the only record of how a choice was made:
gemini.google.com/app/f39fddafa2095317(11apr26) — the only cheat-proofing thread in the entire corpus; its content is nowhere summarised.chatgpt.com/s/cx_6a99e9a8379c81918ee21b94acdd9235— the "Compare matmul costs across chips" megathread that is the stated source of the whole 30 fJ/bit·mm / 400 ps/mm / 10 fJ/bit / 5 pJ/bit constant block.claude.ai/code/session_01USznRPfqyKT1tbu8UeDJ1P— "divergence architecture in andy/yaroslav", the only record of how Andy's design and yours diverged.- The ISA threads (
chatgpt.com/g/…/c/69d19662-…, and two others). - Dally 2020 and Jouppi are linked as private Drive files (
drive.google.com/open?id=…) — unusable in a public spec. http://100.70.243.24:31095/…— a tailnet IP, unreachable to anyone else.
- The primary source is not archived. Dally CACM 2022 returns 403 to automated fetch;
the working copy is a university mirror plus a local
tmp/pdfs/that is not onorigin/main. Every constant in the benchmark traces to a document the repo does not contain.
9. Appendix — the complete fork index
Every choice that has to be made, with its status. DECIDED = settled in code or in a document you wrote; PROPOSED = someone has recommended an answer and it awaits you; OPEN = no answer exists.
ISA
| Fork | Status | |
|---|---|---|
| I1 | Which instruction set is normative across all problems? | OPEN |
| I2 | Cell width and data representation | OPEN |
| I3 | Are writes priced? | OPEN — and a documentation bug today |
| I4 | Is arithmetic priced? | DECIDED (free) in code; the 14.1× is an unaddressed known error |
| I5 | Is instruction fetch priced? | PROPOSED |
| I6 | Do ports (recv/send) exist as instructions? |
PROPOSED |
| I7 | What is E_in / E_out? |
PROPOSED (number chosen, not ratified) |
| I8 | Read/write time asymmetry ("reads have 2× the time") | PROPOSED, unimplemented anywhere |
| I9 | Accumulate semantics at the macro level | OPEN — the only one of Andy's four original forks still posed as a question |
| I10 | Is there a reduce primitive? |
PROPOSED (self-resolved by the proposer, never ratified) |
| I11 | Is mode admitted? |
PROPOSED |
| I12 | Where is the expressivity ceiling (which rung)? | PROPOSED |
| I13 | Division-by-zero and trap semantics | OPEN (shipped code contradicts the design note) |
| I14 | Keep matmul's degree-≤2 symbolic check? | DECIDED in code (PR #31); not revisited |
| I15 | Are tensor/SIMD instructions (dp4a) admitted into v1? |
PROPOSED — adopting drops the 16×16 record 4× overnight |
| I16 | Instruction/line cap for matmul | OPEN (matmul has none; sparse-parity has 2 M) |
Grid size
| Fork | Status | |
|---|---|---|
| G1 | What is the grid — subarray, chip, or nested pair? | OPEN |
| G2 | Address-space bound | OPEN (two shipped scorers already disagree) |
| G3 | What is the next problem? | DECIDED in substance (MNIST) |
| G4 | The MNIST tier ladder | DECIDED by repetition; the 60k-test departure from standard MNIST is unflagged |
| G5 | Accuracy targets per tier | PROPOSED (D5) |
| G6 | Episodes per tier (K) | PROPOSED (D5) |
| G7 | Off-chip / memory tier | PROPOSED (intent stated, nothing implemented) |
| G8 | Fuel / instruction caps per tier | PROPOSED |
| G9 | Resident dataset vs re-streaming | PROPOSED (D6) |
| G10 | 256 MB or 256 MiB for the Dally anchor | OPEN (trivial, but load-bearing for the derivation) |
| G11 | Matmul tiers (does 8192³ join the board?) | PROPOSED |
| G12 | Sparse parity — retire, resize, or keep? | OPEN |
Physical dimensions
| Fork | Status | |
|---|---|---|
| P1 | Cell pitch and quantum | de-facto adopted, but "pitch size is still open" was never retracted, and the 10 µm/bit block is still un-struck in the same doc as the 1 µm/byte one |
| P2 | Movement energy per distance, and one-way vs round-trip | OPEN (three literature values from one author) |
| P3 | Propagation velocity | PROPOSED, with a published 1000× unit error to correct |
| P4 | Technology node | DECIDED (7 nm) in the log, contradicted by every source it cites |
| P5 | Is the memory endpoint priced at all? | OPEN — softest constant in the model, and it moves the 8192³ verdict from 5.14× to 1.0034× |
| P6 | Wire signalling style (repeated vs low-swing) | OPEN — raised nowhere in any design doc, yet it spans ~100× |
| P7 | Off-chip crossing energy | PROPOSED |
| P8 | Manhattan / Chebyshev / Euclidean | OPEN — the highest-leverage one-line decision available; Chebyshev is live on Pages in a document you committed, and the PR retracting it has had no reply |
| P9 | Scalar lin(a) vs native (x,y) |
PROPOSED (agreed by both parties) |
| P10 | ceil(sqrt(distance)) applied to a geometric edge |
OPEN (correction drafted, never posted) |
| P11 | Grid geometry (half-diamond at the origin) | DECIDED |
| P12 | Is TIME scored, and with what semantics? | PROPOSED |
| P13 | What is AREA? | PROPOSED |
| P14 | Static/leakage energy and the idle clock | PROPOSED |
| P15 | 1 pJ vs 1 fJ in your spoken vs written record | resolved as a misspeak (§7 C3) |
| P16 | The "milliseconds" artifact | resolved as a misspeak |
Submission format
| Fork | Status | |
|---|---|---|
| S1 | Trace, generator, or both? | the log sentence ("the generator, not the trace") and the two-track note are not reconciled |
| S2 | If a language, which shape? | PROPOSED — two competing proposals: Andy's SDL and the Grid VM stack |
| S3 | Where does the language live? | OPEN, and untracked — it exists only in a PR body |
| S4 | Which scorer is normative? | PROPOSED (dally-eval) |
| S5 | I/O order — fixed or moving? | OPEN — you wrote the question on 04sep26; the page answering it now exists on main |
| S6 | Online or transductive? | DECIDED (batched inference), with an unexamined hole |
| S7 | Does the organizer ever execute submitted code? | PROPOSED |
| S8 | What is reported, and what is ranked? | OPEN — the largest unwritten piece of the redesign |
| S9 | Leaderboard continuity | OPEN |
| S10 | The submission bundle (IR + report + generator?) | OPEN in rule, settled in practice |
| S11 | Instruction caps as part of the contract | DECIDED as a principle; per-tier numbers PROPOSED |
| S12 | Adjudication protocol | PROPOSED |
Cross-cutting
| Fork | Status | |
|---|---|---|
| X1 | Coarse-and-memorable vs accurate | DECIDED as philosophy, unstated in any spec |
| X2 | Cheat-proofing policy | DECIDED in voice, never written down anywhere |
| X3 | Episode supply / embedded weights | OPEN — the highest-stakes fork in the set |
| X4 | Train/test resampling | PROPOSED |
| X5 | Single control unit or many? | PROPOSED (P1–P8, your own commit 5706824) |
| X6 | What is the model even called? | OPEN — "simplified Dally model", "Grid VM", "PECM", "spatial computer", "stream processor with a scratchpad" are all in use |
| X7 | Who adjudicates, and what happens to the nine open PRs? | OPEN — a collaboration risk, not just a technical one |
| X8 | The two parallel design tracks (yours and Andy's) | OPEN |
| X9 | Documentation consistency before any launch | OPEN (pure cleanup, but it is what a visitor sees) |
| X10 | Hosting, community, distribution | OPEN |
| X11 | Judge cost and hardware budget | PROPOSED / mostly measured |
| X12 | Launch timing, and whether the practical track survives | OPEN |
10. Link hub
Public pages - Benchmark report — https://cybertronai.github.io/sutro-problems/docs/ - Spatial-model analysis — https://cybertronai.github.io/sutro-problems/docs/spatial-model-analysis.html - Grid VM competition design — https://cybertronai.github.io/sutro-problems/docs/grid-vm-competition-design.html - Expressivity–scorability ladder — https://cybertronai.github.io/sutro-problems/docs/expressivity-scorability-ladder.html - Multiprocessors on the Grid VM — https://cybertronai.github.io/sutro-problems/docs/grid-vm-multiprocessor.html - Fixed or moving I/O order? — https://cybertronai.github.io/sutro-problems/docs/dataflow-io-order.html - matmul energy report — https://cybertronai.github.io/sutro-problems/matmul/energy-report/ - matmul animations — https://cybertronai.github.io/sutro-problems/matmul/animation/ - sparse-parity — https://cybertronai.github.io/sutro-problems/sparse-parity/
Repos - https://github.com/cybertronai/sutro-problems - https://github.com/cybertronai/simplified-dally-model · instruction-sets - https://github.com/cybertronai/dally-eval
Open PRs needing you: #48 · #58 · #61 · #49 · #51 · #52 · #55 · #56 · #60
Literature - Dally, On the model of computation: point, CACM 65(9) 2022 — https://cacm.acm.org/opinion/on-the-model-of-computation-point/ - Dally, Turakhia, Han, Domain-Specific Hardware Accelerators, CACM 63(7) 2020 - Dally, AHA 2023 retreat keynote — https://aha.stanford.edu/sites/g/files/sbiybj20066/files/media/file/aha-retreat-2023_dally_keynote_en_eff_ai_hw_0.pdf - Horowitz, Life Post Moore's Law slides (55 pp, Jul 2023; proposal from slide ~37) — https://aha.stanford.edu/sites/g/files/sbiybj20066/files/media/file/aha_071923_horowitz_newcad_0.pdf · MICRO 2023 keynote — https://www.youtube.com/watch?v=q8WK63joI_Y - Jouppi et al., Ten Lessons From Three Generations… TPUv4i, ISCA 2021 - Aggarwal, Alpern, Chandra, Snir, A model for hierarchical memory, STOC 1987 — the HMM with f(a)=a^α; our model is α = ½ - Hong & Kung, I/O complexity: the red-blue pebble game, STOC 1981 - Buck, Scheduling dynamic dataflow graphs with bounded memory (PhD, Berkeley 1993) — the undecidability that fixes the (e)/(f) cliff - Karp, Miller, Winograd, JACM 14(3) 1967 — affine schedules for the v2 time division - Greydanus & Kobak, MNIST-1D, ICML 2024 — procedural generation as the memorization defense - Yadav & Bottou, Cold case: the lost MNIST digits (QMNIST), NeurIPS 2019 - Onur Mutlu dataflow lectures — https://www.youtube.com/@OnurMutluLectures/search?query=dataflow · https://www.youtube.com/watch?v=waqM1JM9GU0
Local paths worth remembering
- ~/drive/gdocs-archive/<docId>/<date>.txt — nightly full text of 35 Google Docs
- ~/My Drive/hiq-transcribed/ — dictated notes, four engines per clip
- ~/…/sutro-problems/tmp/cacti* — the CACTI 22 nm / 32 nm SRAM runs
- ~/…/SutroYaro/telegram.db — group chat, Feb 9 – Mar 28 2026 only
- ~/…/flush-util/machines.toml — says the live Telegram session is on the intel Mac