sutro-problems

** Under construction ** This will be turned into a A100 GPUmode competition. Meanwhile, please submit solutions so we can figure out common parameters/pitfalls:

MNIST (information for humans)

The system takes train/test images + train labels, and produces test labels.

Screenshot 2026-09-11 at 6 39 12 PM

Provide a way to solve this problem at prescribed accuracy without hitting the memory wall.

Motivation

Today’s learning is based on backprop which was popularized in the 80s when we were bottlenecked by arithmetic. Today, we are bottlenecked by memory movement. This favors algorithms with small memory footprint. Backprop has a large memory footprint.

Footprint issue is mitigated by batching and gradient-checkpointing hacks, yet these come with costs. Is there an alternative solution?

To understand the memory wall, consider that the energy of an 8-bit add is comparable to the energy needed to move its operands 10 micrometers. Fetching a byte from the opposite side of a 16 mm chip is worth 1,600 adds. Bill Dally’s AHA retreat slides

Screenshot 2026-09-11 at 7 37 26 PM

Datasets

Accuracy targets:

Submissions

MNIST-small — 67% accuracy target

Date Accuracy Energy on A100 (mJ) Time on A100 (ms) Energy in grid model (mJ) Time in grid model (ms) submission
2026-09-11 67.08% ± 1.54 pp 3,300 130 0.24 3,100 H32 MLP
2026-09-14 67.35% ± 1.76 pp 2,600 110 0.19 1,800 NR-K8 Adam MLP
2026-09-15 67.11% ± 1.63 pp 3,200 120 0.22 3,200 Panel-cached H32 MLP (1,000/1,000)
2026-09-15 67.95% ± 1.64 pp 0.59 0.017 0.00089 10 QDA
2026-09-15 68.39% ± 1.44 pp — — 0.86 1.0 × 10⁴ 9-64-10 SGD MLP

The revised panel result uses the same source examples and draw seeds as H32, but regenerated resized-image hashes differ; see its report’s reproducibility caveat. The accuracy difference is not evidence of superiority.

The QDA row reports the fresh eleven-draw evaluation and corrected A100 energy, with a second-host cross-check. Its report retains the original evidence, discloses the preselection limitation, and documents the pending publicly committed beacon evaluation.

MNIST-medium — 2% error target

Date Accuracy Energy on A100 (mJ) Time on A100 (ms) Energy in grid model (mJ) Time in grid model (ms) submission
2026-09-16 98.12% ± 0.15 pp (A100); 98.11% ± 0.14 pp (grid) 65,000 260 1,500 1.2 × 10⁷ 512-filter CG pair

MNIST-medium — 3% error target

| Date | Accuracy | Energy on A100 (mJ) | Time on A100 (ms) | Energy in grid model (mJ) | Time in grid model (ms) | submission | | — | —: | —: | —: | —: | —: | — |

MNIST-medium — 5% error target

Date Accuracy Energy on A100 (mJ) Time on A100 (ms) Energy in grid model (mJ) Time in grid model (ms) submission
2026-09-11 96.41% ± 0.13 pp 290,000 9,300 440 3.7 × 10⁶ 512-unit MLP (96% target)
2026-09-15 95.57% ± 0.17 pp 174 3.3 0.19 2.0 × 10³ PCA-QDA

PCA-QDA uses the corrected A100 measurement: 174 mJ above idle; see its energy audit and rerun instructions.

MNIST-medium — 8% error target

| Date | Accuracy | Energy on A100 (mJ) | Time on A100 (ms) | Energy in grid model (mJ) | Time in grid model (ms) | submission | | — | —: | —: | —: | —: | —: | — |

MNIST-medium — 12% error target

Date Accuracy Energy on A100 (mJ) Time on A100 (ms) Energy in grid model (mJ) Time in grid model (ms) submission
2026-09-14 89.28% ± 0.36 pp 1,000 39 — — Reversible MLP, smaller workspace (12% error target)

MNIST-original — 1% test error target

Date Accuracy Energy on A100 (mJ) Time on A100 (ms) Energy in grid model (mJ) Time in grid model (ms) submission
2026-09-21 99.06% (A100); 99.11% (grid) 30,000 290 23,000 1.8 × 10⁸ PCANet 9-block + mixture, k=8 K=100
Information for agents ## Task and datasets Learn from the supplied training images and labels, then produce one digit label (0–9) for each test image. The computation being compared includes both training and prediction. Test labels are for evaluation only. - **MNIST-small:** 1,000 training / 1,000 test images, 3 × 3 pixels; at least **67% mean accuracy**. - **MNIST-medium:** 10,000 training / 10,000 test images, 9 × 9 pixels; the five error bands below. - **MNIST-original (MNIST-large):** 60,000 training / 10,000 test images, 28 × 28 pixels; at most **1% test error** (at least **99% accuracy**). For small and medium, use disjoint random training and test subsets of the original 60,000 MNIST training examples. Report **mean accuracy ± sample standard deviation over 11 independently sampled datasets**, with fresh training on each dataset. Record all dataset seeds, preprocessing, checksums, learner seeds, and individual `correct / total` counts. Choose the learning procedure before inspecting test results; do not transfer learned state between datasets. Standard deviation is in **percentage points (pp)**. Original uses the complete official training and test splits. ### MNIST-medium error bands Error is the fraction of incorrect test predictions. Each band is an inclusive maximum error, equivalent to the following minimum accuracy: Across 11 datasets (110,000 test predictions), the thresholds are: - **2% mean error:** at least 98% mean accuracy; 107,800 correct. - **3% mean error:** at least 97% mean accuracy; 106,700 correct. - **5% mean error:** at least 95% mean accuracy; 104,500 correct. - **8% mean error:** at least 92% mean accuracy; 101,200 correct. - **12% mean error:** at least 88% mean accuracy; 96,800 correct. Label each medium result with the error band it targets and whether it meets that band. Apply thresholds to the **unrounded 11-dataset mean**, calculated as `sum(correct) / 110000`; a rounded display percentage does not establish a pass. Report all 11 draws, including their mean and sample standard deviation, rather than selecting favorable draws. ### MNIST-small accuracy target The target is **at least 67% mean accuracy** over 11 independent datasets. Use the exact `sum(correct) / 11000` fraction: at least **7,370 correct** across 11,000 predictions. A rounded display percentage does not establish a pass. The existing H32 MLP result, 7,379 / 11,000, meets this target; its frozen report retains the original 60% study target. ### MNIST-original error target The target is **at most 1% test error**, equivalent to **at least 99% accuracy** on the official 10,000-image test split. A submission must predict at least **9,900 labels correctly**, with at most **100 errors**. Apply this inclusive threshold to the exact `correct / 10000` fraction, not a rounded percentage. ## Efficiency measurements Costs are per complete training-and-prediction run. The two MLP grid entries use a globally serialized schedule; their reports describe this baseline and its limits. Address the memory wall by reducing memory footprint and data movement. The two implementation goals are: - **A100:** provide a kernel or implementation that uses little energy. Use an ISA or toolchain of your choice, such as [pyptx](https://github.com/patrick-toulme/pyptx) or Triton. Report runtime and **idle-adjusted energy measured with NVML**, including the measurement commands and idle-baseline method. - **Bill Dally's 2D grid:** use the [spatial-computer model](https://github.com/cybertronai/simplified-dally-model/tree/main/models/spatial-computer) and its specified instruction set. Declare processor and memory placement, data representation, tape layout, and execution schedule. Count data movement in **word-node hops**, including local accesses, interprocessor traffic, and tape I/O according to that model. Report **peak scratch-memory use**, active processors, hop counts, and the resulting model energy and elapsed time. Identify the model revision and scoring method used. Report **time to score**: the host runtime of computing the theoretical metrics. Account for the full training-and-prediction computation, identifying any setup or preprocessing excluded from a measurement. Keep theoretical scores distinct from measured A100 costs; report an em dash for unmeasured values. Display execution times in **ms**, energies in **mJ** (1 J = 1,000 mJ), and time to score in **s**, using two significant figures for cost measurements. Report memory in bytes or KiB and retain exact counts, accuracies, and measurements in the accompanying data files. If reporting area, state its spatial-model definition and derivation; do not reuse the old single-core area conversion. ## Submission Open a pull request against `main` only for a result that qualifies for a table above, with every reported metric measured or shown as an em dash. Keep experiments, prototypes, measurement-harness proposals, and non-record results on a separate branch rather than merging them into `main`. A record pull request adds the source or generator, reproduction commands, and a standalone report under `mnist/submissions//`. Include the tier, error target for medium or original, accuracy evidence, dataset and learner seeds, model revision, memory layout, scoring calculations, hardware/software versions, A100 measurements, and any W&B runs. Identify which efficiency metrics have been measured and which remain unavailable. Add a row to the matching submission table above. For medium, use a table whose error bound the submission meets; a result may appear in multiple qualifying bands, as on the sparse-parity page. Include the submission date, measured accuracy, A100 energy and runtime, spatial-grid energy and runtime, and a link to the submission report. Put contributors, memory use, hop counts, the full measurement scope, and reproduction evidence in that report. Keep historical results under their original specification. ## Existing tooling The [older agent instructions](/sutro-problems/mnist/instructions.html), default dataset generator, and accuracy evaluator still describe the previous specification: 600/600 and 6,000/6,000 sizes and single-core scoring. The evaluator also lacks the current small target and the original tier's 99% target. Update or configure reproduction code to match the datasets, error bands, and spatial model above before claiming current results. For medium, select generator profile `medium-error-targets-v1` for 10,000/10,000 examples and evaluator `--error-target` for one of the five error bands. The generator's `reference-20260910` profile has the new counts but uses a different train/test split protocol; matching counts alone is insufficient. </details>
Historical submissions (previous specification) The entries below retain their original measurements. Small used 600 training and 600 test images, except Static dataflow 3x3, whose row reports its int8-safe model on 1,000 of each (the 600-example prototype had a logit overflow in bytecode and is documented in its report, not the table); medium used 6,000 of each, except Ordered ConvNets, which used 10,000 of each and meets the 3% mean-error target. Reported Dally scores and areas use the former **single-core-with-tape** model. These results do not establish spatial-computer costs; consult each report for its accuracy scope. Their reports preserve the single-core scores, areas, scoring times, and original evaluation scope, including whether an entry used one dataset or 11 draws. Grid columns are unmeasured: single-core scores are not spatial-grid scores. ### MNIST-small (historical) | Date | Accuracy | Energy on A100 (mJ) | Time on A100 (ms) | Energy in grid model (mJ) | Time in grid model (ms) | submission | | --- | ---: | ---: | ---: | ---: | ---: | --- | | 2026-09-14 | 48.4% | 0.11 | 0.0062 | — | — | [Static dataflow 3x3 int8-safe (1,000/1,000)](/sutro-problems/mnist/submissions/static-dataflow-3x3-20260914/report.html) | | 2026-09-12 | 65.0% ± 2.1 pp | 1,800 | 57 | — | — | [Panel-cached MLP (600/600)](/sutro-problems/mnist/submissions/mlp-panels-v4-20260911/report.html) | | 2026-09-10 | 62% | 2,000 | 71 | — | — | [32-unit MLP](https://cybertronai.github.io/sutro-problems/docs/submissions/mlp60-affine-20260911/) | | 2026-09-10 | 51% | 0.52 | 0.0069 | — | — | [1NN](https://cybertronai.github.io/sutro-problems/docs/submissions/1nn-v4-20260911/) | ### MNIST-medium (historical) | Date | Accuracy | Energy on A100 (mJ) | Time on A100 (ms) | Energy in grid model (mJ) | Time in grid model (ms) | submission | | --- | ---: | ---: | ---: | ---: | ---: | --- | | 2026-09-11 | 97.8% ± 0.1 pp | 780,000 | 13,000 | — | — | [Ordered ConvNets (3% error target)](https://cybertronai.github.io/sutro-problems/docs/submissions/medium-convnet-v4-20260911/) | | 2026-09-10 | 98.1% ± 0.1 pp | 1.7 × 10⁶ | 5.9 × 10⁴ | — | — | [Three ConvNets](https://cybertronai.github.io/sutro-problems/docs/submissions/medium-convnet-11draw-20260911/) | | 2026-09-10 | 96% | 150,000 | 4,700 | — | — | [512-unit MLP](https://cybertronai.github.io/sutro-problems/docs/submissions/medium-affine-20260911/) | ### MNIST-large (historical) | Date | Accuracy | Energy on A100 (mJ) | Time on A100 (ms) | Energy in grid model (mJ) | Time in grid model (ms) | submission | | --- | ---: | ---: | ---: | ---: | ---: | --- | Historical single-core-with-tape sketch: ![MNIST competition sketch: scoring metrics, dataset tiers, and a Bill Dally single-core model with tape](/sutro-problems/mnist/doc/competition-overview.png)