Measured board energy · movement model · literature
A100 and grid energy for 8192³ GEMM and Ciresan MNIST
A controlled comparison of exact 8192 × 8192 × 8192 matrix multiplication across four arithmetic formats, followed by one INT8 forward-only pass over all 60,000 MNIST training images.
Report date: September 2, 2026GPU: NVIDIA A100-SXM4-40GB, 400 W capEnergy: idle-adjusted NVML board counter
Exact workload2⁴⁰ ops549.756 billion MACs
8K A100 range0.648–15.757 JB1 through strict FP32
MNIST, full batch7.774 ms · 2.569 Jall 60k examples at once
MNIST, batch 16780.364 ms · 12.106 J3,750 logical batches
Comparison contract
What is—and is not—being compared
Measured A100 values are physical GPU-board observations. Grid values are movement-energy estimates using exactly 10−15 J per source-operand grid step, with all results converted to joules. Literature values are estimates derived from the cited experiments. They are three different evidence classes, not interchangeable measurements.
A100 boundary. Operands, outputs, weights, activations, and workspaces are already resident on the GPU. Allocation, warm-up, host transfer, input conversion, quantization, and B1 packing are outside the measured window. The result includes GPU-board and HBM activity, then subtracts the mean of paired loaded-idle measurements.
Grid boundary. Each source read travels ceil(sqrt(address)) abstract grid steps, each priced at 10−15 J. Arithmetic, destination writes, bit width, physical memory hierarchy, clocks, control, leakage, and elapsed time cost zero. Its scalar cells are unbounded, so changing FP32 to FP16, INT8, or B1 does not change the modeled movement energy.
N = 8192 = 2¹³
MAC = N³ = 2³⁹ = 549,755,813,888
ops = 2N³ = 2⁴⁰ = 1,099,511,627,776
E_grid [J] = 10⁻¹⁵ × Σ(reads × grid steps)
Table 1
Exact 8192³ matrix multiplication
All four rows execute the same 549,755,813,888 pair contributions. The A100 paths differ in arithmetic and output semantics; the grid proxy deliberately does not.
Measured A100Modeled gridLiterature-derived
Times and A100 energies are five-trial aggregate results. “Baseline range” varies only the two paired idle estimates; trial-to-trial standard deviation is also shown. Literature entries are not direct replicas of this benchmark.
≈0.5–0.9 Jcross-benchmark planning band; no paired exact-shape energy result
Packed B1 Tensor Coreone-bit AND + population count → INT32; native SM80 WMMA BMMA
2.944 ms
0.648 Jtrial SD 0.0205 J baseline 0.646–0.651 J
0.118953 Jone bit per abstract scalar cell
≈0.089 J†aggressive extrapolation from a much larger nonsquare binary beamformer with different useful-op semantics
Precision helps, but not monotonically in time
Relative to strict FP32, measured energy falls 11.0× for FP16, 17.4× for INT8, and 24.3× for B1. The custom B1 kernel is slightly slower than the mature INT8 path, yet draws less incremental board power.
The model is intentionally widthless
The retuned modeled movement is 0.118953 J for every row. Scaling that energy by 32, 16, 8, or 1 bits would be an extra assumption, not part of the current widthless model.
No 8192³ instruction file
The grid result is evaluated in closed form for a retuned sa_cache schedule with Ti=256, Tj=128. No enormous unrolled 8K IR is materialized; the downloadable script and audit JSON are the complete projection artifacts.
Table 2
One INT8 forward pass over the 60k MNIST training set
The recovered local network has logical dimensions 784–2500–2000–1500–1000–500–10: 11,965,000 MACs per image and 717.9 billion logical MACs per full-dataset inference epoch. This is inference only—no labels, loss, backward pass, or optimizer.
Weights are deterministic dense synthetic INT8 values, and real MNIST bytes are centered into signed INT8. After each of six layers, an INT32 bias is added, ReLU is applied, and the result is saturated to [0,127] and written as INT8. This measures shape and execution cost, not classifier accuracy.
Batch strategy
GEMM + epilogue launches
Physical MACs
A100 time
A100 idle-adjusted energy
Grid movement energy (J)
Comparable literature
Logical batch 163,750 calls; each M=16 input is pre-padded to physical M=32 for this PyTorch A100 INT8 path
22,500 + 22,500
1.445 T2.01× logical, including row and width padding
780.364 ms
12.106 Jtrial SD 0.515 J baseline 11.430–12.782 J
0.134310 Jone persistent epoch address space
—no comparable published A100 INT8 epoch-energy result found
One full batchall 60,000 images in a single forward call
6 + 6
722.504 G0.641% width padding
7.774 ms
2.569 Jtrial SD 0.0411 J baseline 2.549–2.590 J
0.133503 Jone jointly packed layout
—no comparable published A100 INT8 epoch-energy result found
The A100 bottleneck is tiny launches
Batch 16 is 100.4× slower and 4.71× more energy than the full batch. The main causes are 22,500 small GEMMs plus 22,500 epilogues and the native-path M=32 padding—not a change in logical network work.
The grid result is almost tied
With weights and all 60,000 inputs held in one epoch-wide address space, batch 16 is only 0.60% above full batch. The model prices reads from every image’s distinct resident cells; it does not multiply a 16-row input calculation. It still omits GPU launches, elapsed time, idle power, and physical row padding.
MAC-only sensitivity
Simply multiplying 717.9B MACs by the 8K grid energy per MAC gives 0.155335 J for either batching. The main table uses the more informative shape-aware rectangular schedule instead.
Table 3
Batching in one persistent 2D grid
Every row performs the same 717.9 billion logical MACs over the same 60,000 images. All 47,040,000 input cells (60,000 × 784) and one copy of every weight are resident at distinct addresses before execution; 600,000 final-logit output cells are preallocated. The initial materialization energy is outside the boundary, while every subsequent input read is charged from that epoch-wide storage. Only computed hidden activations and scratch use layer-specific batch-row buffers that can be overwritten by the next invocation.
Movement intensity is the average abstract distance traveled by one charged read: total charged grid steps divided by all charged read operations, including the final output reads. Energy is automatically converted to joules as total charged grid steps × 10−15 J. Each row is the exact evaluation of the best schedule found from three deterministic coordinate-descent starts, not a global lower bound. Resident regions are frequency-packed once before each counterfactual run, so absolute addresses may differ between batch-size rows. The Dally small-RAM column adds 4 × 10−13 J per charged read to the same schedule as an illustrative sensitivity; it is not part of the movement-only model.
Logical batch
Movement intensity (grid steps / charged read)
Persistent-grid movement energy (J)
Vs. full batch
Dally small-RAM sensitivity (J)
Full: 60,000
46.21 steps/read
0.133503 J
1.000×
1.289060 J
64
25.11 steps/read
0.073034 J
0.547×
1.236646 J
16
45.63 steps/read
0.134310 J
1.006×
1.311717 J
4
141.40 steps/read
0.435668 J
3.263×
1.668119 J
1
461.52 steps/read
1.673598 J
12.536×
3.124118 J
Batch-64 input storage
The first 937 calls address 47,014,912 distinct input cells (937 × 64 × 784); the final 32-example call addresses another 25,088. All 47,040,000 cells exist before execution and each is read 20 times by the first layer. They are frequency-packed in the epoch-wide grid, not copied through or replaced by one reusable 64-row input buffer.
Why the full schedule has higher movement intensity
The 46.21 steps/read value is a mean travel distance, not a read count. Both rows keep the same 47,040,000 input cells and one weight copy resident. The difference is transient storage: the modeled monolithic full-batch schedule allocates 450 million hidden-activation cells—60,000 × (2,500 + 2,000 + 1,500 + 1,000 + 500)—whereas batch 64 reuses only 480,000 such cells. Its total modeled footprint is therefore 509.6 million cells versus 60.1 million. Under frequency packing around one processor and d(a)=ceil(sqrt(a)), reusable batch-64 activations and scratch stay much nearer the processor. Full batch performs 0.69% fewer charged reads (2.888892T versus 2.909029T), but its movement intensity is 84.1% higher (46.21 versus 25.11), so its energy is 0.133503 versus 0.073034 J—82.8% higher. An unrestricted full-batch program could stream internally like batch 64, so this is a result of the modeled schedules, not an inherent lower bound for full-batch inference.
What “harder to make efficient” means here
The grid model has no GPU lanes, occupancy, launch latency, clock, or leakage. It can express batch inefficiency only when the program performs more reads or moves operands farther. A hardware-efficiency penalty belongs in the A100 model, not as an unexplained surcharge on modeled grid energy.
Heuristic cross-check
How this squares with Dally’s 2022–2023 heuristics
Dally’s central point survives the comparison: communication and location dominate arithmetic. His example says a 32-bit add costs 20 fJ, moving its two 32-bit inputs 1 mm costs 1.9 pJ, and fetching two words from main memory costs 1.3 nJ. The stated 40 mm in 16 ns corresponds to about 2.5×10⁶ m/s—roughly c/120.
His MLP recomputation example uses 160 fJ/MAC. Applied mechanically, that is 0.087961 J for 8192³ and 0.114864 J for this 60k-image network. That network estimate is near the persistent full-batch and batch-16 grid energies, but the agreement is only an order-of-magnitude cross-check: the 10−15 J/grid-step coefficient is abstract, and neither estimate is A100 board energy.
Dally’s later 2023 AHA talk rounds on-chip communication to 100 fJ/(bit·mm) and a small-RAM access to 50 fJ/bit. At that calibration, 10−15 J per whole INT8 operand-step is energy-equivalent to only 1.25 μm of wire; adding 4 × 10−13 J for every charged read produces the last column of Table 3. That is only a what-if: it overprices reads served by registers and underprices any 32-bit accumulator read.
The article does not identify a semiconductor node for all of these heuristic numbers. The A100 is a 7 nm device, but a 2022 publication date alone does not make Dally’s examples A100- or 7 nm-calibrated.
Grid vs. Dally, 8192³: 0.118953 J is 1.35× the 160 fJ/MAC heuristic.
A100 vs. grid, 8192³: measured board energy is 5.45× grid for B1, 7.60× for INT8, 11.99× for FP16, and 132.47× for strict FP32.
A100 vs. grid, MNIST: 19.25× for the full batch and 90.14× for logical batch 16.
Interpretation: the gaps are where real data paths, HBM, output traffic, control, launches, clocking, and leakage enter. That is evidence for Dally’s hierarchy, not a contradiction.
Methods and reproducibility
Measurement and projection details
The numbers below are sufficient to reproduce the grid calculation and to interpret the A100 energy boundary without treating nvidia-smi samples as a high-rate wattmeter.
A100 measurement
A100-SXM4-40GB; 400 W cap; driver 580.95.05; CUDA 12.4; PyTorch 2.5.1.
Unit conversion: 10−15 J for one source operand moving one abstract address-grid step.
Best among all 196 ordered divisor-tile pairs for this schedule family, not a proof over every algorithm.
MNIST grid projection
Six rectangular GEMMs are jointly packed by read frequency.
Tile search keeps the best of three deterministic coordinate-descent starts: independent per-layer optima, all-minimum tiles, and all-maximum tiles.
Primary batching comparison uses one persistent epoch address space: 47,040,000 distinct source-image cells and one copy of all weights are resident; 600,000 output cells are allocated before execution.
Computed hidden activations and scratch use layer-specific batch-row buffers; input examples do not pass through a reusable batch-sized staging buffer.
Logical dimensions only; Tensor Core padding, bias, ReLU, saturation, writes, and launches are omitted.
“Literature estimate” means a calculation anchored to published measurements. None of the cited papers directly reproduces every datatype, shape, output type, software path, and energy boundary used here.