sutro-problems

9-64-10 SGD MLP: MNIST-small, 67% target

Submission date: September 15, 2026 (UTC).

This submission reports 7,523 / 11,000 = 68.39% mean accuracy with sample standard deviation 1.44 pp over eleven independently sampled datasets, meeting the 67% target (7,370 required) with 153 predictions of margin. Each draw contains 1,000 training and 1,000 disjoint test images reduced to 3x3 from the official 60,000-image training split. The learner is a 9-64-10 ReLU MLP trained with plain constant-learning-rate squared-error minibatch SGD — deliberately the same optimizer family as the first accepted entry, so the grid score comes from the reviewed, unchanged SGD path of the shared scorer rather than a new optimizer extension.

Accuracy evidence

Draws use Generator(PCG64(seed)).permutation(60000); rows 0-999 are training and rows 1000-1999 are test, disjoint within each draw. Predictions were written and hashed before any evaluation-label slice was opened (two-phase freeze/score in run.py); evidence/accuracy/draw_manifest.json records seeds, indices and input hashes, and evidence/accuracy/prediction_manifest.json records per-draw prediction, parameter and score hashes.

Draw Dataset seed Correct / total Accuracy
0 20261201 690 / 1,000 69.0%
1 20261202 703 / 1,000 70.3%
2 20261203 687 / 1,000 68.7%
3 20261204 669 / 1,000 66.9%
4 20261205 652 / 1,000 65.2%
5 20261206 685 / 1,000 68.5%
6 20261207 692 / 1,000 69.2%
7 20261208 687 / 1,000 68.7%
8 20261209 683 / 1,000 68.3%
9 20261210 700 / 1,000 70.0%
10 20261211 675 / 1,000 67.5%

The exact mean is 7523 / 11000 = 0.6839090909...; sample standard deviation is 1.4376748905 pp (ddof=1). Qualification uses the exact count against the 7,370 required.

Predeclared configuration

The learning procedure was fixed on disjoint pilot seeds 20261120-20261130 before the official draws were evaluated. A 24-configuration sweep (width 32/48/64 x epochs 300/500 x lr 0.1/0.2/0.3/0.4, batch 25, seed 101) ranked width 64 / 500 epochs / lr 0.1 first at 7678/11000 = 69.80% +/- 1.42 pp (pilot.py, evidence/pilot_results.json). On the same pilot draws the accepted H32 configuration (width 32, 300 epochs, lr 0.2) scores 68.17%, versus its official 67.08%, so the pilot set runs about 1 pp easy; the official margin over the bar (68.39% vs 67%) matches that expectation. No evaluation results influenced the procedure.

Frozen learner

Architecture 9-64-10 (1,290 FP32 parameters), batch 25, learning rate 0.1, 500 epochs, fixed supplied sample order, fresh seed-101 initialization per draw, ordered FP32 arithmetic, no state transfer between draws. Update rule with step = lr / batch = 0.004 as one FP32 scalar:

g = gradients from pre-update weights, ascending-index FP32 reductions
p = p - step * g

Squared-error loss 0.5*sum_class(error^2) averaged over the complete minibatch; strict ReLU comparison; first-index argmax. This is exactly the learner semantics of the accepted H32 entry at a different width, epoch count and learning rate. reference.py (CPU) and gpu_benchmark.py (GPU) implement the same recurrence; verify.py re-derives draw 0 bit-exactly and re-checks every manifest hash.

Grid-model cost

Energy in grid model Time in grid model Peak allocated scratch Time to score
0.865 mJ 10,361 ms 100,076 bytes 0.047 s

Exact totals: 864,909,081,610 fJ and 10,361,243,280 cycles for the serialized schedule under spatial-computer revision 01a0bd5e0d2564825b0f53dd766f763c82dbc7c0. The program and score are in ../grid-mlp-scoring-20260912/small64-sgd-20260915/, generated by the shared scorer’s default (SGD) path with --features 9 --width 64 --epochs 500 --batch 25 --n-train 1000 --n-test 1000 --learning-rate 0.1 --seed 101. No scorer code was modified: the same command with the H32 configuration reproduces the accepted small60 score byte-for-byte (0.242990182992 mJ / 3,073.508456 ms), re-verified in this workspace. The fixed program has these same grid costs on all 11 datasets (the trace depends only on the declared shape, not on input values). The machine, placement, tape layout and serialization contract are those of the shared scorer README: arithmetic on P(125,0), 250 instruction-issuing bottom processors with at most one outstanding access, charged bottom-edge tapes, 19,000 input words and 1,000 output words.

Measured A100 cost

Accuracy Energy on A100 Time on A100
68.39% +/- 1.44 pp — (pending) — (pending)

The paired A100 NVML audit (gpu_benchmark.py, identical protocol to the accepted H32 entry: CUDA-graph replays of the complete task, NVML board energy with paired 3-second idle intervals after 3-second settling gaps, three trials, bitwise GPU-vs-CPU validation of parameters, scores and predictions including changed-labels and changed-queries mutations) could not run at packaging time: the Modal workspace carrying the only available A100 credentials exceeded its spend limit. The benchmark, payload generation (run.py gpu-payload) and verification path are complete and will be executed unchanged once spend authority is restored; this section and the leaderboard row will then be updated with measured values.

Verification and limitations

verify.py re-checks draw seeds, train/test indices, raw-source SHA-256 hashes, the per-draw train-image, train-label and test-image preprocessed input hashes against the draw manifest, prediction hashes and the label-derived accuracy total; re-derives draw 0 end-to-end from the ordered FP32 reference comparing parameter, score and prediction bits; confirms the pilot sweep’s ascending-K einsum matmuls are bit-equivalent to the ordered loops; and, when results/gpu_results.json exists, re-verifies GPU predictions/parameters against the frozen CPU evidence and recomputes idle-adjusted energy from raw counters.

Downsampling to 3x3 uses run.resize_recorded: the repository’s box-area weights accumulated in increasing index order with float64 product/sum intermediates and a float32 cast after each step. The operation is BLAS-free, so the archived input hashes and the draw-0 re-derivation reproduce bit-identically across platforms (Linux x86 verified; the same construction is already merged in the medium PCA-QDA and MNIST-small QDA entries), where the previous float32-matmul area_resize dispatched to different BLAS kernels with different rounding. Limitations: one GPU class planned (A100-SXM4-40GB); A100 measurement pending spend authorization; grid columns use the conservative globally serialized schedule, not a parallel-placement claim.

Reproduce and audit

Run from this directory (Python 3.11+ with NumPy; the packaging host used Python 3.14.7 / NumPy 2.5.1 via nix develop; after the deterministic-resize fix, all checks were re-run on Linux x86_64 with Python 3.12.14 / NumPy 2.5.2; downsampling and the draw-0 re-derivation are BLAS-free and platform-independent; the score reproduction below is environment-insensitive because the scorer is integer/hop-exact):

python run.py prepare        # official draws, indices and input hashes
python run.py freeze         # train + hash predictions (no evaluation labels)
python run.py score          # verify every prediction hash first, then count
python run.py gpu-payload    # write generated/gpu-payload.json for the GPU runner
python verify.py             # re-check manifests, inputs, accuracy, draw-0 bits
modal run gpu_benchmark.py   # paired A100 NVML audit (needs Modal A100 access)

Grid score regeneration from the repository root:

python mnist/submissions/grid-mlp-scoring-20260912/score.py \
  --features 9 --width 64 --epochs 500 --batch 25 --n-train 1000 --n-test 1000 \
  --learning-rate 0.1 --seed 101 --output /tmp/small64-grid-score

Evidence: accuracy, draw manifest, prediction manifest, evaluation freeze, pilot sweep, grid score, submission summary.