Submission date: September 15, 2026 (UTC).
This submission reports 7,523 / 11,000 = 68.39% mean accuracy with sample standard deviation 1.44 pp over eleven independently sampled datasets, meeting the 67% target (7,370 required) with 153 predictions of margin. Each draw contains 1,000 training and 1,000 disjoint test images reduced to 3x3 from the official 60,000-image training split. The learner is a 9-64-10 ReLU MLP trained with plain constant-learning-rate squared-error minibatch SGD — deliberately the same optimizer family as the first accepted entry, so the grid score comes from the reviewed, unchanged SGD path of the shared scorer rather than a new optimizer extension.
Draws use Generator(PCG64(seed)).permutation(60000); rows 0-999 are training
and rows 1000-1999 are test, disjoint within each draw. Predictions were
written and hashed before any evaluation-label slice was opened (two-phase
freeze/score in run.py); evidence/accuracy/draw_manifest.json records
seeds, indices and input hashes, and evidence/accuracy/prediction_manifest.json
records per-draw prediction, parameter and score hashes.
| Draw | Dataset seed | Correct / total | Accuracy |
|---|---|---|---|
| 0 | 20261201 | 690 / 1,000 | 69.0% |
| 1 | 20261202 | 703 / 1,000 | 70.3% |
| 2 | 20261203 | 687 / 1,000 | 68.7% |
| 3 | 20261204 | 669 / 1,000 | 66.9% |
| 4 | 20261205 | 652 / 1,000 | 65.2% |
| 5 | 20261206 | 685 / 1,000 | 68.5% |
| 6 | 20261207 | 692 / 1,000 | 69.2% |
| 7 | 20261208 | 687 / 1,000 | 68.7% |
| 8 | 20261209 | 683 / 1,000 | 68.3% |
| 9 | 20261210 | 700 / 1,000 | 70.0% |
| 10 | 20261211 | 675 / 1,000 | 67.5% |
The exact mean is 7523 / 11000 = 0.6839090909...; sample standard deviation
is 1.4376748905 pp (ddof=1). Qualification uses the exact count against
the 7,370 required.
The learning procedure was fixed on disjoint pilot seeds 20261120-20261130
before the official draws were evaluated. A 24-configuration sweep (width
32/48/64 x epochs 300/500 x lr 0.1/0.2/0.3/0.4, batch 25, seed 101) ranked
width 64 / 500 epochs / lr 0.1 first at 7678/11000 = 69.80% +/- 1.42 pp
(pilot.py, evidence/pilot_results.json). On the same pilot draws the
accepted H32 configuration (width 32, 300 epochs, lr 0.2) scores 68.17%,
versus its official 67.08%, so the pilot set runs about 1 pp easy; the
official margin over the bar (68.39% vs 67%) matches that expectation. No
evaluation results influenced the procedure.
Architecture 9-64-10 (1,290 FP32 parameters), batch 25, learning rate 0.1,
500 epochs, fixed supplied sample order, fresh seed-101 initialization per
draw, ordered FP32 arithmetic, no state transfer between draws. Update rule
with step = lr / batch = 0.004 as one FP32 scalar:
g = gradients from pre-update weights, ascending-index FP32 reductions
p = p - step * g
Squared-error loss 0.5*sum_class(error^2) averaged over the complete
minibatch; strict ReLU comparison; first-index argmax. This is exactly the
learner semantics of the accepted H32 entry at a different width, epoch count
and learning rate. reference.py (CPU) and gpu_benchmark.py (GPU) implement
the same recurrence; verify.py re-derives draw 0 bit-exactly and re-checks
every manifest hash.
| Energy in grid model | Time in grid model | Peak allocated scratch | Time to score |
|---|---|---|---|
| 0.865 mJ | 10,361 ms | 100,076 bytes | 0.047 s |
Exact totals: 864,909,081,610 fJ and 10,361,243,280 cycles for the serialized
schedule under spatial-computer revision
01a0bd5e0d2564825b0f53dd766f763c82dbc7c0. The program and score are in
../grid-mlp-scoring-20260912/small64-sgd-20260915/, generated by the shared
scorer’s default (SGD) path with
--features 9 --width 64 --epochs 500 --batch 25 --n-train 1000 --n-test 1000
--learning-rate 0.1 --seed 101. No scorer code was modified: the same
command with the H32 configuration reproduces the accepted small60 score
byte-for-byte (0.242990182992 mJ / 3,073.508456 ms), re-verified in this
workspace. The fixed program has these same grid costs on all 11 datasets
(the trace depends only on the declared shape, not on input values). The
machine, placement, tape layout and serialization contract are those of the
shared scorer README: arithmetic on P(125,0), 250 instruction-issuing bottom
processors with at most one outstanding access, charged bottom-edge tapes,
19,000 input words and 1,000 output words.
| Accuracy | Energy on A100 | Time on A100 |
|---|---|---|
| 68.39% +/- 1.44 pp | — (pending) | — (pending) |
The paired A100 NVML audit (gpu_benchmark.py, identical protocol to the
accepted H32 entry: CUDA-graph replays of the complete task, NVML board
energy with paired 3-second idle intervals after 3-second settling gaps,
three trials, bitwise GPU-vs-CPU validation of parameters, scores and
predictions including changed-labels and changed-queries mutations) could not
run at packaging time: the Modal workspace carrying the only available A100
credentials exceeded its spend limit. The benchmark, payload generation
(run.py gpu-payload) and verification path are complete and will be
executed unchanged once spend authority is restored; this section and the
leaderboard row will then be updated with measured values.
verify.py re-checks draw seeds, train/test indices, raw-source SHA-256
hashes, the per-draw train-image, train-label and test-image preprocessed
input hashes against the draw manifest, prediction hashes and the
label-derived accuracy total; re-derives draw 0 end-to-end from the ordered
FP32 reference comparing parameter, score and prediction bits; confirms the
pilot sweep’s ascending-K einsum matmuls are bit-equivalent to the ordered
loops; and, when results/gpu_results.json exists, re-verifies GPU
predictions/parameters against the frozen CPU evidence and recomputes
idle-adjusted energy from raw counters.
Downsampling to 3x3 uses run.resize_recorded: the repository’s box-area
weights accumulated in increasing index order with float64 product/sum
intermediates and a float32 cast after each step. The operation is BLAS-free,
so the archived input hashes and the draw-0 re-derivation reproduce
bit-identically across platforms (Linux x86 verified; the same construction
is already merged in the medium PCA-QDA and MNIST-small QDA entries), where
the previous float32-matmul area_resize dispatched to different BLAS
kernels with different rounding. Limitations: one GPU class planned
(A100-SXM4-40GB); A100 measurement pending spend authorization; grid columns
use the conservative globally serialized schedule, not a parallel-placement
claim.
Run from this directory (Python 3.11+ with NumPy; the packaging host used Python 3.14.7 / NumPy 2.5.1 via nix develop; after the deterministic-resize fix, all checks were re-run on Linux x86_64 with Python 3.12.14 / NumPy 2.5.2; downsampling and the draw-0 re-derivation are BLAS-free and platform-independent; the score reproduction below is environment-insensitive because the scorer is integer/hop-exact):
python run.py prepare # official draws, indices and input hashes
python run.py freeze # train + hash predictions (no evaluation labels)
python run.py score # verify every prediction hash first, then count
python run.py gpu-payload # write generated/gpu-payload.json for the GPU runner
python verify.py # re-check manifests, inputs, accuracy, draw-0 bits
modal run gpu_benchmark.py # paired A100 NVML audit (needs Modal A100 access)
Grid score regeneration from the repository root:
python mnist/submissions/grid-mlp-scoring-20260912/score.py \
--features 9 --width 64 --epochs 500 --batch 25 --n-train 1000 --n-test 1000 \
--learning-rate 0.1 --seed 101 --output /tmp/small64-grid-score
Evidence: accuracy, draw manifest, prediction manifest, evaluation freeze, pilot sweep, grid score, submission summary.