A 60→256→256→10 MLP trained from scratch passes difficulty 1 in 61.388 ms per call, the median of three sandboxed Modal A100-80GB runs. The three runs used the unchanged official scorer from repository commit 5714d1a. Each run contains 11 MNIST calls and four foreign hold-out calls, with a fresh process and foreign warmup for every timed call.
| Evidence attempt | Score, ms/call | MNIST accuracy | Hold-out | Hold-out accuracy | Result |
|---|---|---|---|---|---|
| 2 | 61.388 | 95.07% (104,582/110,000) | Fashion-MNIST | 85.98% | Pass |
| 3 | 61.269 | 95.13% (104,638/110,000) | Fashion-MNIST | 85.67% | Pass |
| 4 | 61.590 | 95.30% (104,828/110,000) | KMNIST | 91.56% | Pass |
Scores are the scorer’s final ranked times: the slower of the MNIST and hold-out means, including its wall-clock floor where applicable. Passing also checks the worst MNIST draw, hold-out accuracy, timing dispersion, and per-call limits. Only difficulty 1 was tested.
On 2026-09-26, after main gained the energy column, the same file passed difficulty 1 again in three sandboxed Modal A100-80GB runs of mnist.py 1.2.0. Their medians are the ones in the records table: 61.701 ms and 2,131 mJ per call above idle.
| Run | GPU | Score, ms/call | Energy, mJ/call | MNIST accuracy | Hold-out |
|---|---|---|---|---|---|
| 1 | A100-SXM4-80GB | 61.280 | 2,131 | 95.00% (104,502/110,000) | KMNIST 91.62% |
| 2 | A100-SXM4-80GB | 61.701 | 2,185 | 95.01% (104,513/110,000) | Fashion-MNIST 85.75% |
| 3 | A100-SXM4-80GB | 61.942 | 2,113 | 95.26% (104,783/110,000) | Fashion-MNIST 85.85% |
Each run landed on a different board and passed every energy check. One container failed inside Modal before it received its input (modal.exception.InternalError: failed to get new inputs, in runs.log); Modal ran that input again. Records: run 1, run 2, run 3.
This submission ports the repository’s mlpg-k1-w256-s200-b512 cutoff experiment (credited below) to the three-argument mnist-a100 API. Its algorithm is unchanged: 200 minibatches of 512, AdamW, a cosine learning-rate schedule with warmup, input noise, dropout, label smoothing, and EMA inference. One training step is captured in a CUDA graph during the untimed warmup.
Every invocation copies the current training data and labels and resets model weights, Adam moments and counters, EMA weights, the schedule counter, and CUDA random state. The graph is reused within the process; learned state is not reused. The submission includes no training examples or pretrained weights. Its source is 5,729 bytes and passes the source checker with no review flags.
From mnist-a100/, with Modal configured:
python run_modal.py submissions/graph-mlp-20260926/fast_mlp.py:fast_mlp --difficulty 1 --runs 3
The included verify_modal.py runs that same scorer with sandboxing required, records source/scorer hashes and complete output, and reserves budget before launching each attempt. It uses separate single-use containers, function retries disabled (retries=0), a 600-second function timeout and a 120-second startup timeout. The checked-in ledger preserves this task’s reservations. A rerun writes a separate ledger and evidence directory:
python submissions/graph-mlp-20260926/verify_modal.py --runs 3 --output /tmp/mnist-a100-verification
The verification controller sets cpu=(1, 2) and memory=(4096, 8192), with the same CUDA/PyTorch image as run_modal.py. It runs the unchanged mnist.py in a clean subprocess with MNIST_SANDBOX=required.
Four launch attempts reserve $3.70: $0.80 each plus $0.50 for CPU image building, below the requested $5 limit. Attempt 1 failed during controller startup and was stopped before any scoring calls; its full reservation remains in the ledger. The controller import was repaired without changing the submission or scorer. Attempts 2–4 are the three scored runs. These are conservative reservations, not a provider billing receipt. Modal’s billing API reports $0.22687950 in compute across these four app IDs (about $0.23), including the failed startup. All four apps are stopped with zero tasks; the itemized billing snapshot is included below.
All three passing containers reported NVIDIA A100-SXM4-80GB, Python 3.13.0, PyTorch 2.12.0+cu130, and sandbox on. Their GPU UUIDs are distinct and recorded in the evidence. The median is 11.40× faster than the listed 700.1 ms example baseline; aggregate MNIST accuracy across 330,000 predictions is 95.166%.
The local CPU scorer suite initially passed 49 tests; its stdout-isolation test hit the timing-dispersion gate on macOS. That test passed on rerun. The scorer was not modified. Local validation records both outcomes.
SHA-256:
44cb0c03d4ac869e99186c584d234c8743c12565e0a3b49cf3172bd5c8a823252833ba470776314eabd00f8315561f31da16f839740f5776b58fb8b1322480221bc6d8d96a776565bc70f3ccdb91be7de57187d4495b2b0437fac98f5e120064