** Under construction **
This will be turned into a A100 GPUmode competition. Meanwhile, please submit solutions so we can figure out common parameters/pitfalls:
- Can we measure A100 energy correctly? (Otherwise, might forced to stick with time)
- What should the accuracy targets be? Is 2/3/5/8/12% good? 2% might too hard for mnist-medium
- Can we detect cheating easily (ie boring solutions like hardcoding weights)?
The system takes train/test images + train labels, and produces test labels.

Provide a way to solve this problem at prescribed accuracy without hitting the memory
wall.
- kernel that runs on A100 using few Joules (measured using NVML)
- (optional) an algorithm that runs with small memory footprint in Bill Dally’s 2D grid (measured by counting hops in Bill Dally’s 2D grid)
Motivation
Today’s learning is based on backprop which was popularized in the 80s when we were bottlenecked by arithmetic. Today, we are bottlenecked by memory movement. This favors algorithms with small memory footprint. Backprop has a large memory footprint.
Footprint issue is mitigated by batching and gradient-checkpointing hacks, yet these come with costs. Is there an alternative solution?
To understand the memory wall, consider that the energy of an 8-bit add is comparable to the energy needed to move its operands 10 micrometers. Fetching a byte from the opposite side of a 16 mm chip is worth 1,600 adds. Bill Dally’s AHA retreat slides

Datasets
- MNIST-small: 1k train, 1k test, 3x3 images
- MNIST-medium: 10k train, 10k test, 9x9 images
- MNIST-large: original 60k train, 10k test, 28x28 images
Accuracy targets:
- MNIST-small requires 67% test accuracy
- MNIST-medium comes with 5 test-set error targets, 2% error, 3% error, 5% error, 8% error, 12% error
- MNIST-large uses LeNet5 original 1% error rate
Submissions
MNIST-small — 67% accuracy target
| Date |
Accuracy |
Energy on A100 (mJ) |
Time on A100 (ms) |
Energy in grid model (mJ) |
Time in grid model (ms) |
submission |
| 2026-09-11 |
67.08% ± 1.54 pp |
3,300 |
130 |
0.24 |
3,100 |
H32 MLP |
| 2026-09-14 |
67.35% ± 1.76 pp |
2,600 |
110 |
0.19 |
1,800 |
NR-K8 Adam MLP |
| 2026-09-15 |
67.11% ± 1.63 pp |
3,200 |
120 |
0.22 |
3,200 |
Panel-cached H32 MLP (1,000/1,000) |
| 2026-09-15 |
67.95% ± 1.64 pp |
0.59 |
0.017 |
0.00089 |
10 |
QDA |
| 2026-09-15 |
68.39% ± 1.44 pp |
— |
— |
0.86 |
1.0 × 10⁴ |
9-64-10 SGD MLP |
The revised panel result uses the same source examples and draw seeds as H32,
but regenerated resized-image hashes differ; see its report’s reproducibility caveat.
The accuracy difference is not evidence of superiority.
The QDA row reports the fresh eleven-draw evaluation and corrected A100
energy, with a second-host cross-check. Its report retains the original
evidence, discloses the preselection limitation, and documents the pending
publicly committed beacon evaluation.
MNIST-medium — 2% error target
| Date |
Accuracy |
Energy on A100 (mJ) |
Time on A100 (ms) |
Energy in grid model (mJ) |
Time in grid model (ms) |
submission |
| 2026-09-16 |
98.12% ± 0.15 pp (A100); 98.11% ± 0.14 pp (grid) |
65,000 |
260 |
1,500 |
1.2 × 10⁷ |
512-filter CG pair |
MNIST-medium — 3% error target
| Date | Accuracy | Energy on A100 (mJ) | Time on A100 (ms) | Energy in grid model (mJ) | Time in grid model (ms) | submission |
| — | —: | —: | —: | —: | —: | — |
MNIST-medium — 5% error target
| Date |
Accuracy |
Energy on A100 (mJ) |
Time on A100 (ms) |
Energy in grid model (mJ) |
Time in grid model (ms) |
submission |
| 2026-09-11 |
96.41% ± 0.13 pp |
290,000 |
9,300 |
440 |
3.7 × 10⁶ |
512-unit MLP (96% target) |
| 2026-09-15 |
95.57% ± 0.17 pp |
174 |
3.3 |
0.19 |
2.0 × 10³ |
PCA-QDA |
PCA-QDA uses the corrected A100 measurement: 174 mJ above idle; see its energy audit and rerun instructions.
MNIST-medium — 8% error target
| Date | Accuracy | Energy on A100 (mJ) | Time on A100 (ms) | Energy in grid model (mJ) | Time in grid model (ms) | submission |
| — | —: | —: | —: | —: | —: | — |
MNIST-medium — 12% error target
MNIST-original — 1% test error target
| Date |
Accuracy |
Energy on A100 (mJ) |
Time on A100 (ms) |
Energy in grid model (mJ) |
Time in grid model (ms) |
submission |
| 2026-09-21 |
99.06% (A100); 99.11% (grid) |
30,000 |
290 |
23,000 |
1.8 × 10⁸ |
PCANet 9-block + mixture, k=8 K=100 |
Information for agents
## Task and datasets
Learn from the supplied training images and labels, then produce one digit label
(0–9) for each test image. The computation being compared includes both training
and prediction. Test labels are for evaluation only.
- **MNIST-small:** 1,000 training / 1,000 test images, 3 × 3 pixels; at least
**67% mean accuracy**.
- **MNIST-medium:** 10,000 training / 10,000 test images, 9 × 9 pixels; the five
error bands below.
- **MNIST-original (MNIST-large):** 60,000 training / 10,000 test images,
28 × 28 pixels; at most **1% test error** (at least **99% accuracy**).
For small and medium, use disjoint random training and test subsets of the
original 60,000 MNIST training examples. Report **mean accuracy ± sample standard
deviation over 11 independently sampled datasets**, with fresh training on each
dataset. Record all dataset seeds, preprocessing, checksums, learner seeds, and
individual `correct / total` counts. Choose the learning procedure before
inspecting test results; do not transfer learned state between datasets.
Standard deviation is in **percentage points (pp)**. Original uses the complete
official training and test splits.
### MNIST-medium error bands
Error is the fraction of incorrect test predictions. Each band is an inclusive
maximum error, equivalent to the following minimum accuracy:
Across 11 datasets (110,000 test predictions), the thresholds are:
- **2% mean error:** at least 98% mean accuracy; 107,800 correct.
- **3% mean error:** at least 97% mean accuracy; 106,700 correct.
- **5% mean error:** at least 95% mean accuracy; 104,500 correct.
- **8% mean error:** at least 92% mean accuracy; 101,200 correct.
- **12% mean error:** at least 88% mean accuracy; 96,800 correct.
Label each medium result with the error band it targets and whether it meets
that band. Apply thresholds to the **unrounded 11-dataset mean**, calculated as
`sum(correct) / 110000`; a rounded display percentage does not establish a pass.
Report all 11 draws, including their mean and sample standard deviation, rather
than selecting favorable draws.
### MNIST-small accuracy target
The target is **at least 67% mean accuracy** over 11 independent datasets.
Use the exact `sum(correct) / 11000` fraction: at least **7,370 correct** across
11,000 predictions. A rounded display percentage does not establish a pass.
The existing H32 MLP result, 7,379 / 11,000, meets this target; its frozen report
retains the original 60% study target.
### MNIST-original error target
The target is **at most 1% test error**, equivalent to **at least 99% accuracy**
on the official 10,000-image test split. A submission must predict at least
**9,900 labels correctly**, with at most **100 errors**. Apply this inclusive
threshold to the exact `correct / 10000` fraction, not a rounded percentage.
## Efficiency measurements
Costs are per complete training-and-prediction run. The two MLP grid entries
use a globally serialized schedule; their reports describe this baseline and
its limits.
Address the memory wall by reducing memory footprint and data movement. The two
implementation goals are:
- **A100:** provide a kernel or implementation that uses little energy. Use an
ISA or toolchain of your choice, such as
[pyptx](https://github.com/patrick-toulme/pyptx) or Triton. Report runtime and
**idle-adjusted energy measured with NVML**, including the measurement commands
and idle-baseline method.
- **Bill Dally's 2D grid:** use the
[spatial-computer model](https://github.com/cybertronai/simplified-dally-model/tree/main/models/spatial-computer)
and its specified instruction set. Declare processor and memory placement,
data representation, tape layout, and execution schedule. Count data movement
in **word-node hops**, including local accesses, interprocessor traffic, and
tape I/O according to that model. Report **peak scratch-memory use**, active
processors, hop counts, and the resulting model energy and elapsed time.
Identify the model revision and scoring method used. Report **time to score**:
the host runtime of computing the theoretical metrics. Account for the full
training-and-prediction computation, identifying any setup or preprocessing
excluded from a measurement. Keep theoretical scores distinct from measured
A100 costs; report an em dash for unmeasured values.
Display execution times in **ms**, energies in **mJ** (1 J = 1,000 mJ), and time
to score in **s**, using two significant figures for cost measurements. Report
memory in bytes or KiB and retain exact counts, accuracies, and measurements in
the accompanying data files. If reporting area, state its spatial-model
definition and derivation; do not reuse the old single-core area conversion.
## Submission
Open a pull request against `main` only for a result that qualifies for a table
above, with every reported metric measured or shown as an em dash. Keep
experiments, prototypes, measurement-harness proposals, and non-record results
on a separate branch rather than merging them into `main`.
A record pull request adds the source or generator, reproduction commands, and
a standalone report under `mnist/submissions//`. Include the tier, error
target for medium or original, accuracy evidence, dataset and learner seeds,
model revision, memory layout, scoring calculations, hardware/software versions, A100
measurements, and any W&B runs. Identify which efficiency metrics have been
measured and which remain unavailable.
Add a row to the matching submission table above. For medium, use a table whose
error bound the submission meets; a result may appear in multiple qualifying
bands, as on the sparse-parity page. Include the submission date, measured
accuracy, A100 energy and runtime, spatial-grid energy and runtime, and a link
to the submission report. Put contributors, memory use, hop counts, the full
measurement scope, and reproduction evidence in that report. Keep historical results under their original
specification.
## Existing tooling
The [older agent instructions](/sutro-problems/mnist/instructions.html), default dataset generator, and
accuracy evaluator still describe the previous specification: 600/600 and
6,000/6,000 sizes and single-core scoring. The evaluator also lacks the current
small target and the original tier's 99% target. Update or
configure reproduction code to match the datasets, error bands, and spatial
model above before claiming current results.
For medium, select generator profile `medium-error-targets-v1` for 10,000/10,000
examples and evaluator `--error-target` for one of the five error bands.
The generator's `reference-20260910` profile has the new counts but uses a
different train/test split protocol; matching counts alone is insufficient.
</details>
Historical submissions (previous specification)
The entries below retain their original measurements. Small used 600 training
and 600 test images, except Static dataflow 3x3, whose row reports its
int8-safe model on 1,000 of each (the 600-example prototype had a logit
overflow in bytecode and is documented in its report, not the table); medium
used 6,000 of each, except Ordered ConvNets, which
used 10,000 of each and meets the 3% mean-error target. Reported Dally scores
and areas use the former **single-core-with-tape** model. These results do not
establish spatial-computer costs; consult each report for its accuracy scope.
Their reports preserve the single-core scores, areas, scoring times, and
original evaluation scope, including whether an entry used one dataset or 11
draws. Grid columns are unmeasured: single-core scores are not spatial-grid
scores.
### MNIST-small (historical)
| Date | Accuracy | Energy on A100 (mJ) | Time on A100 (ms) | Energy in grid model (mJ) | Time in grid model (ms) | submission |
| --- | ---: | ---: | ---: | ---: | ---: | --- |
| 2026-09-14 | 48.4% | 0.11 | 0.0062 | — | — | [Static dataflow 3x3 int8-safe (1,000/1,000)](/sutro-problems/mnist/submissions/static-dataflow-3x3-20260914/report.html) |
| 2026-09-12 | 65.0% ± 2.1 pp | 1,800 | 57 | — | — | [Panel-cached MLP (600/600)](/sutro-problems/mnist/submissions/mlp-panels-v4-20260911/report.html) |
| 2026-09-10 | 62% | 2,000 | 71 | — | — | [32-unit MLP](https://cybertronai.github.io/sutro-problems/docs/submissions/mlp60-affine-20260911/) |
| 2026-09-10 | 51% | 0.52 | 0.0069 | — | — | [1NN](https://cybertronai.github.io/sutro-problems/docs/submissions/1nn-v4-20260911/) |
### MNIST-medium (historical)
| Date | Accuracy | Energy on A100 (mJ) | Time on A100 (ms) | Energy in grid model (mJ) | Time in grid model (ms) | submission |
| --- | ---: | ---: | ---: | ---: | ---: | --- |
| 2026-09-11 | 97.8% ± 0.1 pp | 780,000 | 13,000 | — | — | [Ordered ConvNets (3% error target)](https://cybertronai.github.io/sutro-problems/docs/submissions/medium-convnet-v4-20260911/) |
| 2026-09-10 | 98.1% ± 0.1 pp | 1.7 × 10⁶ | 5.9 × 10⁴ | — | — | [Three ConvNets](https://cybertronai.github.io/sutro-problems/docs/submissions/medium-convnet-11draw-20260911/) |
| 2026-09-10 | 96% | 150,000 | 4,700 | — | — | [512-unit MLP](https://cybertronai.github.io/sutro-problems/docs/submissions/medium-affine-20260911/) |
### MNIST-large (historical)
| Date | Accuracy | Energy on A100 (mJ) | Time on A100 (ms) | Energy in grid model (mJ) | Time in grid model (ms) | submission |
| --- | ---: | ---: | ---: | ---: | ---: | --- |
Historical single-core-with-tape sketch:
