Measured board energy · movement model · literature

A100 and grid energy for 8192³ GEMM and Ciresan MNIST

A controlled comparison of exact 8192 × 8192 × 8192 matrix multiplication across four arithmetic formats, followed by one INT8 forward-only pass over all 60,000 MNIST training images.

Report date: September 2, 2026 GPU: NVIDIA A100-SXM4-40GB, 400 W cap Energy: idle-adjusted NVML board counter
Exact workload 2⁴⁰ ops 549.756 billion MACs
8K A100 range 0.648–15.757 J B1 through strict FP32
MNIST, full batch 7.774 ms · 2.569 J all 60k examples at once
MNIST, batch 16 780.364 ms · 12.106 J 3,750 logical batches

Comparison contract

What is—and is not—being compared

Measured A100 values are physical GPU-board observations. Grid values are movement-energy estimates using exactly 10−15 J per source-operand grid step, with all results converted to joules. Literature values are estimates derived from the cited experiments. They are three different evidence classes, not interchangeable measurements.

A100 boundary. Operands, outputs, weights, activations, and workspaces are already resident on the GPU. Allocation, warm-up, host transfer, input conversion, quantization, and B1 packing are outside the measured window. The result includes GPU-board and HBM activity, then subtracts the mean of paired loaded-idle measurements.

Grid boundary. Each source read travels ceil(sqrt(address)) abstract grid steps, each priced at 10−15 J. Arithmetic, destination writes, bit width, physical memory hierarchy, clocks, control, leakage, and elapsed time cost zero. Its scalar cells are unbounded, so changing FP32 to FP16, INT8, or B1 does not change the modeled movement energy.

N = 8192 = 2¹³
MAC = N³ = 2³⁹ = 549,755,813,888
ops = 2N³ = 2⁴⁰ = 1,099,511,627,776

E_grid [J] = 10⁻¹⁵ × Σ(reads × grid steps)

Table 1

Exact 8192³ matrix multiplication

All four rows execute the same 549,755,813,888 pair contributions. The A100 paths differ in arithmetic and output semantics; the grid proxy deliberately does not.

Measured A100 Modeled grid Literature-derived
Times and A100 energies are five-trial aggregate results. “Baseline range” varies only the two paired idle estimates; trial-to-trial standard deviation is also shown. Literature entries are not direct replicas of this benchmark.
Arithmetic path and semantics A100 time A100 idle-adjusted energy Grid movement energy (J) Literature energy estimate
Strict FP32FP32 × FP32 → FP32; TF32 disabled; cuBLAS 57.945 ms 15.757 Jtrial SD 0.477 J
baseline 15.650–15.865 J
0.118953 Jsame movement energy for every row 13.04–13.26 Jcubic interpolation between measured 4096³ and 16384³ A100 cases; 16.41 J from a separate efficiency result
FP16 Tensor CoreFP16 × FP16 → FP16; cuBLAS 4.362 ms 1.426 Jtrial SD 0.0149 J
baseline 1.420–1.432 J
0.118953 J 1.04–1.28 Jcubic interpolation between measured adjacent-size A100 cases
INT8 Tensor Coresigned INT8 × INT8 → INT32; torch._int_mm 2.779 ms 0.904 Jtrial SD 0.0102 J
baseline 0.901–0.907 J
0.118953 J ≈0.5–0.9 Jcross-benchmark planning band; no paired exact-shape energy result
Packed B1 Tensor Coreone-bit AND + population count → INT32; native SM80 WMMA BMMA 2.944 ms 0.648 Jtrial SD 0.0205 J
baseline 0.646–0.651 J
0.118953 Jone bit per abstract scalar cell ≈0.089 J†aggressive extrapolation from a much larger nonsquare binary beamformer with different useful-op semantics

Precision helps, but not monotonically in time

Relative to strict FP32, measured energy falls 11.0× for FP16, 17.4× for INT8, and 24.3× for B1. The custom B1 kernel is slightly slower than the mature INT8 path, yet draws less incremental board power.

The model is intentionally widthless

The retuned modeled movement is 0.118953 J for every row. Scaling that energy by 32, 16, 8, or 1 bits would be an extra assumption, not part of the current widthless model.

No 8192³ instruction file

The grid result is evaluated in closed form for a retuned sa_cache schedule with Ti=256, Tj=128. No enormous unrolled 8K IR is materialized; the downloadable script and audit JSON are the complete projection artifacts.

Table 2

One INT8 forward pass over the 60k MNIST training set

The recovered local network has logical dimensions 784–2500–2000–1500–1000–500–10: 11,965,000 MACs per image and 717.9 billion logical MACs per full-dataset inference epoch. This is inference only—no labels, loss, backward pass, or optimizer.

Weights are deterministic dense synthetic INT8 values, and real MNIST bytes are centered into signed INT8. After each of six layers, an INT32 bias is added, ReLU is applied, and the result is saturated to [0,127] and written as INT8. This measures shape and execution cost, not classifier accuracy.
Batch strategy GEMM + epilogue launches Physical MACs A100 time A100 idle-adjusted energy Grid movement energy (J) Comparable literature
Logical batch 163,750 calls; each M=16 input is pre-padded to physical M=32 for this PyTorch A100 INT8 path 22,500 + 22,500 1.445 T2.01× logical, including row and width padding 780.364 ms 12.106 Jtrial SD 0.515 J
baseline 11.430–12.782 J
0.134310 Jone persistent epoch address space —no comparable published A100 INT8 epoch-energy result found
One full batchall 60,000 images in a single forward call 6 + 6 722.504 G0.641% width padding 7.774 ms 2.569 Jtrial SD 0.0411 J
baseline 2.549–2.590 J
0.133503 Jone jointly packed layout —no comparable published A100 INT8 epoch-energy result found

The A100 bottleneck is tiny launches

Batch 16 is 100.4× slower and 4.71× more energy than the full batch. The main causes are 22,500 small GEMMs plus 22,500 epilogues and the native-path M=32 padding—not a change in logical network work.

The grid result is almost tied

With weights and all 60,000 inputs held in one epoch-wide address space, batch 16 is only 0.60% above full batch. The model prices reads from every image’s distinct resident cells; it does not multiply a 16-row input calculation. It still omits GPU launches, elapsed time, idle power, and physical row padding.

MAC-only sensitivity

Simply multiplying 717.9B MACs by the 8K grid energy per MAC gives 0.155335 J for either batching. The main table uses the more informative shape-aware rectangular schedule instead.

Table 3

Batching in one persistent 2D grid

Every row performs the same 717.9 billion logical MACs over the same 60,000 images. All 47,040,000 input cells (60,000 × 784) and one copy of every weight are resident at distinct addresses before execution; 600,000 final-logit output cells are preallocated. The initial materialization energy is outside the boundary, while every subsequent input read is charged from that epoch-wide storage. Only computed hidden activations and scratch use layer-specific batch-row buffers that can be overwritten by the next invocation.

Movement intensity is the average abstract distance traveled by one charged read: total charged grid steps divided by all charged read operations, including the final output reads. Energy is automatically converted to joules as total charged grid steps × 10−15 J. Each row is the exact evaluation of the best schedule found from three deterministic coordinate-descent starts, not a global lower bound. Resident regions are frequency-packed once before each counterfactual run, so absolute addresses may differ between batch-size rows. The Dally small-RAM column adds 4 × 10−13 J per charged read to the same schedule as an illustrative sensitivity; it is not part of the movement-only model.
Logical batch Movement intensity (grid steps / charged read) Persistent-grid movement energy (J) Vs. full batch Dally small-RAM sensitivity (J)
Full: 60,000 46.21 steps/read 0.133503 J 1.000× 1.289060 J
64 25.11 steps/read 0.073034 J 0.547× 1.236646 J
16 45.63 steps/read 0.134310 J 1.006× 1.311717 J
4 141.40 steps/read 0.435668 J 3.263× 1.668119 J
1 461.52 steps/read 1.673598 J 12.536× 3.124118 J

Batch-64 input storage

The first 937 calls address 47,014,912 distinct input cells (937 × 64 × 784); the final 32-example call addresses another 25,088. All 47,040,000 cells exist before execution and each is read 20 times by the first layer. They are frequency-packed in the epoch-wide grid, not copied through or replaced by one reusable 64-row input buffer.

Why the full schedule has higher movement intensity

The 46.21 steps/read value is a mean travel distance, not a read count. Both rows keep the same 47,040,000 input cells and one weight copy resident. The difference is transient storage: the modeled monolithic full-batch schedule allocates 450 million hidden-activation cells—60,000 × (2,500 + 2,000 + 1,500 + 1,000 + 500)—whereas batch 64 reuses only 480,000 such cells. Its total modeled footprint is therefore 509.6 million cells versus 60.1 million. Under frequency packing around one processor and d(a)=ceil(sqrt(a)), reusable batch-64 activations and scratch stay much nearer the processor. Full batch performs 0.69% fewer charged reads (2.888892T versus 2.909029T), but its movement intensity is 84.1% higher (46.21 versus 25.11), so its energy is 0.133503 versus 0.073034 J—82.8% higher. An unrestricted full-batch program could stream internally like batch 64, so this is a result of the modeled schedules, not an inherent lower bound for full-batch inference.

What “harder to make efficient” means here

The grid model has no GPU lanes, occupancy, launch latency, clock, or leakage. It can express batch inefficiency only when the program performs more reads or moves operands farther. A hardware-efficiency penalty belongs in the A100 model, not as an unexplained surcharge on modeled grid energy.

Heuristic cross-check

How this squares with Dally’s 2022–2023 heuristics

Dally’s central point survives the comparison: communication and location dominate arithmetic. His example says a 32-bit add costs 20 fJ, moving its two 32-bit inputs 1 mm costs 1.9 pJ, and fetching two words from main memory costs 1.3 nJ. The stated 40 mm in 16 ns corresponds to about 2.5×10⁶ m/s—roughly c/120.

His MLP recomputation example uses 160 fJ/MAC. Applied mechanically, that is 0.087961 J for 8192³ and 0.114864 J for this 60k-image network. That network estimate is near the persistent full-batch and batch-16 grid energies, but the agreement is only an order-of-magnitude cross-check: the 10−15 J/grid-step coefficient is abstract, and neither estimate is A100 board energy.

Dally’s later 2023 AHA talk rounds on-chip communication to 100 fJ/(bit·mm) and a small-RAM access to 50 fJ/bit. At that calibration, 10−15 J per whole INT8 operand-step is energy-equivalent to only 1.25 μm of wire; adding 4 × 10−13 J for every charged read produces the last column of Table 3. That is only a what-if: it overprices reads served by registers and underprices any 32-bit accumulator read.

The article does not identify a semiconductor node for all of these heuristic numbers. The A100 is a 7 nm device, but a 2022 publication date alone does not make Dally’s examples A100- or 7 nm-calibrated.

  • Grid vs. Dally, 8192³: 0.118953 J is 1.35× the 160 fJ/MAC heuristic.
  • A100 vs. grid, 8192³: measured board energy is 5.45× grid for B1, 7.60× for INT8, 11.99× for FP16, and 132.47× for strict FP32.
  • A100 vs. grid, MNIST: 19.25× for the full batch and 90.14× for logical batch 16.
  • Interpretation: the gaps are where real data paths, HBM, output traffic, control, launches, clocking, and leakage enter. That is evidence for Dally’s hierarchy, not a contradiction.

Methods and reproducibility

Measurement and projection details

The numbers below are sufficient to reproduce the grid calculation and to interpret the A100 energy boundary without treating nvidia-smi samples as a high-rate wattmeter.

A100 measurement

  • A100-SXM4-40GB; 400 W cap; driver 580.95.05; CUDA 12.4; PyTorch 2.5.1.
  • Authoritative counter: nvmlDeviceGetTotalEnergyConsumption.
  • Five approximately five-second trials per stage, synchronized only at aggregate boundaries.
  • Idle adjustment uses the mean of loaded-idle windows measured immediately before and after each stage.
  • nvidia-smi at 100 ms records supporting power, utilization, clocks, and temperature telemetry.

8K grid projection

  • Retuned generalized sa_cache schedule, Ti=256, Tj=128.
  • Movement energy: 0.118953 J.
  • Unit conversion: 10−15 J for one source operand moving one abstract address-grid step.
  • Best among all 196 ordered divisor-tile pairs for this schedule family, not a proof over every algorithm.

MNIST grid projection

  • Six rectangular GEMMs are jointly packed by read frequency.
  • Tile search keeps the best of three deterministic coordinate-descent starts: independent per-layer optima, all-minimum tiles, and all-maximum tiles.
  • Primary batching comparison uses one persistent epoch address space: 47,040,000 distinct source-image cells and one copy of all weights are resident; 600,000 output cells are allocated before execution.
  • Computed hidden activations and scratch use layer-specific batch-row buffers; input examples do not pass through a reusable batch-sized staging buffer.
  • Logical dimensions only; Tensor Core padding, bias, ReLU, saturation, writes, and launches are omitted.

Primary references

Sources and literature classification

“Literature estimate” means a calculation anchored to published measurements. None of the cited papers directly reproduces every datatype, shape, output type, software path, and energy boundary used here.

  1. Shared research report: “Estimate A100 GEMM Energy.” Used as the supplied research map; primary sources below were checked separately.
  2. Ma et al., “GEMM the New Gem.” A100 time and energy at 4096³ and 16384³; the FP32 and FP16 values in Table 1 are cubic interpolations to 8192³.
  3. Ootomo and Yokota, “Recovering single precision accuracy from Tensor Cores.” The 67 GFLOP/J cuBLAS FP32 result gives the 16.41 J cross-shape estimate.
  4. Luo et al., “Benchmarking and Dissecting the NVIDIA Hopper GPU Architecture.” A100 INT8 instruction-level throughput and power anchor.
  5. vllmbench A100 cuBLASLt INT8 exact-shape CSV. Timing anchor used by the supplied research report; output-type caveat applies.
  6. Oostrum et al., “The Tensor-Core Beamformer.” Binary A100 TOP/J anchor on a much larger nonsquare workload; the 0.089 J row is deliberately marked aggressive and non-comparable.
  7. William Dally, “On the Model of Computation: Point,” CACM 65(9), 2022. Communication-distance and 160 fJ/MAC MLP heuristics. Accessible PDF.
  8. Bill Dally, “Energy Efficiency and AI Hardware,” Stanford AHA Retreat, 2023. Later 100 fJ/(bit·mm), 50 fJ/bit small-RAM, and 1 fJ/bit add heuristics.
  9. NVIDIA Ampere architecture overview. A100 process and arithmetic-mode context.
  10. NVIDIA NVML device-query reference. Cumulative energy-counter definition.
  11. Cireșan et al., “Deep Big Multilayer Perceptrons.” Historical architecture context; it does not report A100 INT8 inference energy.