GitHub ↗
Sutro Problems · research note
28 Aug 2026

16 × 16 matrix multiplication · 4,096 MACs

One femtojoule per scored grid step.

Applied literally, the current 66,300-point record becomes a 66.3-picojoule synthetic movement term. It is a crisp measure of locality. It is not a measurement of the energy consumed by a complete matmul chip.

66.3 pJrecord score × exactly 1 fJ per operand-grid-step
5.14×lower movement term than the 340,704-point baseline
123.6TOPS/W-equivalent for this term alone across 8,192 conventional operations

Bottom line

The prototype is useful—and the label matters.

Use “movement-only proxy at 1 fJ per operand-grid-step.” The score captures the compiler win cleanly: more reads, placed dramatically nearer the ALU. Whole-device figures in the literature are mostly tens to thousands of times larger after optimistic per-operation normalization because they include some combination of arithmetic, SRAM endpoints, networks, control, clocks, leakage, conversion, and idle power.

80.54%reduction in the scored term

01 · Unit audit

The conversion is exact once the counted quantity is fixed.

The scorer charges each source-operand read and each final output read by ceil(sqrt(address)). The score is already the sum of those read charges. It should not be multiplied by a second “total distance” unless that second quantity has a documented, different definition.

Here, “grid-step” means the scorer’s one-way address-radius abstraction. It is not a trace of producer-to-consumer Manhattan routes, router hops, or physical millimetres of wire.

Eproxy = ε · S
ε = 1 fJ / (operand · grid-step)
S = Σpaid reads ceil(√address)
50 pJ

50,000 score units × 1 fJ/unit = 50,000 fJ = 5.0 × 10−11 J.

550 pJ

550,000 traversed units × 1 fJ/unit = 550,000 fJ = 5.5 × 10−10 J.

Reconciliation required: 50,000 fJ and a total traversed distance of 550,000 cannot both follow the same 1 fJ/unit rule. If 50,000 is a normalized objective, publish the normalization equation and call it “score units,” not raw distance.

Conversion calculator

For a 16×16 GEMM, vendors conventionally count 2mnk = 8,192 operations. The repository executes 4,096 multiplies and 3,840 explicit adds.

66.3 pJ6.63 × 10⁻¹¹ J8.093 fJ/op · 123.6 TOPS/W-equivalent

02 · Compiler report

The optimizer bought locality with explicit copies.

Naive baseline340.704 pJ16,128 paid reads · average distance 21.125
Current record66.300 pJ17,781 paid reads · average distance 3.729

The record performs 10.25% more paid reads, entirely through 1,653 copies, while cutting average distance by 82.35%. The weighted total falls by 80.54%. This is precisely the kind of scheduling trade that a distance-sensitive compiler objective should expose.

Histogram and cumulative distribution of baseline access distances
Baseline: most reads are far away; 340,704 total score. Tap to enlarge.
Histogram and cumulative distribution of record access distances
Record: most reads land in the first few tiers; 66,300 total score. Tap to enlarge.

Record instruction audit

Record instruction counts and score contribution
Instruction classCountPaid readsScore / fJ
Multiply4,0968,19214,272
Add3,8407,68026,209
Copy1,6531,65321,565
Output exit—2564,254
Total9,58917,78166,300

What the score currently makes free

  • 9,589 destination writes and their storage endpoints
  • 4,096 multiplies and 3,840 additions
  • initial placement of all 512 input elements
  • addresses, requests, acknowledgements, and control traffic
  • clocking, leakage, bandwidth limits, stalls, and time

The 66.3 pJ value therefore means “priced source-side movement under this abstract geometry,” not “energy at the wall” and not even “energy of the complete accelerator core.”

Record history under the 1 fJ rule

Selected matmul record history converted at 1 fJ per score unit
ProgramScoreMovement proxyvs. baseline
4×4 baseline1,3161.316 pJ—
4×4 outer product8000.800 pJ1.65× lower
16×16 baseline340,704340.704 pJ1.00×
16×16 recursive237,456237.456 pJ1.43× lower
16×16 tiled133,783133.783 pJ2.55× lower
16×16 hierarchical80,21780.217 pJ4.25× lower
16×16 aliased69,69769.697 pJ4.89× lower
16×16 current record66,30066.300 pJ5.14× lower

Reproducibility: inspect the record report, submitted IR, machine-readable audit, or comparison CSV.

Precision is not a footnote

The verifier uses unbounded Python integers. Current inputs reach 256 and outputs reach 161,368, requiring 18 unsigned bits. An INT8 chip comparison is therefore a scale reference, not semantic equivalence. A hardware experiment should first choose uint8 modular, int8 × int8 → int32, or exact wide-integer semantics; that choice changes both arithmetic and transfer energy.

03 · Hardware literature

Published “energy per matmul” is rarely one comparable thing.

The table below converts each source’s reported efficiency to an 8192-operation energy equivalent. It answers: “If the published large-workload or peak pJ/op held for 8,192 useful ops, what energy scale would result?” It does not predict the energy of an isolated 16×16 launch.

Published hardware figures normalized to 8,192 useful operations
Hardware / pathArchitecture and published basisEvidence8192-op equivalent÷ 66.3 pJ
Sutro recordAbstract one-ALU source-read distance; exact-integer verifierchosen model0.0663 nJ1×
Intel 22FFL analog SRAM CIM8-bit activation × 8-bit weight charge-domain analog MVM macro; measured computation error <0.5%; 32.2 TOPS/W peak. Assumes TOPS counts multiply + add as two ops.measured macro0.254 nJ3.84×
Apple M1 ANEFixed-function FP16 matrix engine; measured 0.37 pJ/FLOP optimum and about 0.5 pJ/FLOP sustained on saturated work.measured engine3.03–4.10 nJ45.7–61.8×
Google TPUv1256×256 systolic array; 86 TOPS achieved on CNN0 paired with 40 W incremental busy power.workload / busy W3.81 nJ57.5×
Tesla FSD computerFour broadcast 96×96 NNA grids across two chips; 144 TOPS and 72 W for the computer.peak / system W4.10 nJ61.8×
NM-CarusNon-systolic near-memory 8-bit SIMD; 6.8 pJ/MAC in 65 nm post-layout evaluation.post-layout27.9 nJ420×
AMD MI300X FP32 vectorConventional vector path; 163.4 TFLOP/s official peak and 750 W maximum TBP.peak / TBP37.6 nJ567×
EyerissNon-systolic row-stationary spatial array; 83.1 GMAC/s/W on measured convolution workload.measured chip49.3 nJ743×
Graphcore C2 IPU FP32Distributed MIMD tiles with local SRAM and specialized AMP units; 18.9 TFLOP/s per-IPU GEMM peak divided by 150 W nominal per-IPU power.throughput / nominal W65.0 nJ981×
NVIDIA H100 FP32 CUDA coresConventional SIMT path, not Tensor Cores; 60 TFLOP/s peak and 700 W TDP.peak / TDP95.6 nJ1,442×
NVIDIA A100 FP32 CUDA coresNon-Tensor-Core SGEMM; maximum 67.0 GFLOP/s/W over tested large shapes using device telemetry.measured device122 nJ1,844×
Dual Xeon E5-2650v4 AVX2 FP32OpenBLAS n=5000 SGEMM; 5.92 GFLOP/J from CPU package counters.measured package1.384 µJ20,872×

The especially relevant non-systolic case: Tesla

not systolic

Tesla’s IEEE Micro paper is unusually explicit: each 96×96 MAC keeps its own local accumulator; 96-element activation and weight vectors are broadcast along rows and columns; there is no intra-array operand movement from MAC to MAC. Two 2 GHz NNAs deliver 72 TOPS per chip with INT8 inputs and 30-bit local accumulators.

Why it is still not apples-to-apples

The safe public power figure is for the entire two-chip computer: 144 raw aggregate TOPS at 72 W. Converting that ratio gives 4.096 nJ per 8,192 executed operations, but a lone 16×16 multiply cannot fill even one 96×96 array. Both SoCs normally duplicate the same inference for safety, so normalizing by unique useful work would double that energy. This is an efficiency-scale comparison, not a tiny-call measurement.

The Hot Chips slide also shows a 15 W NNA bar, but its scope is not explicit enough to use as a per-chip energy denominator.

Repeated 16×16 measurements show the utilization cliff

A 2012 study reports energy and elapsed time for a window of 1,000 repeated 32-bit 16×16 multiplies. Dividing its table values by 1,000 gives 17.62 µJ on a Cyclone II systolic FPGA, 20.11 µJ on a Clarkdale Core i5, and 34.79 µJ on a Pineview Atom per multiply. These are old platforms and software stacks, not modern efficiency targets. The CPU figures subtract idle power, while the FPGA measurement excludes regulator loss. The experiment is still valuable because it shows why scaling large-GEMM TOPS/W down to 8,192 operations can miss small-call reality by orders of magnitude.

17.62 µJCyclone II FPGA

60.02 µs average · systolic implementation · 265,762× the proxy

20.11 µJIntel Core i5

3.90 µs average · MATLAB/LAPACK · 303,318× the proxy

34.79 µJIntel Atom

41.4 µs average · MATLAB/LAPACK · 524,736× the proxy

Modern fixed-function engines have the same qualitative issue: the Apple ANE study reports roughly a 0.23 ms host/firmware dispatch floor and a separate engine-rail reading near 0.9 W for tiny matmuls. Their rough cross-product is 207 µJ, not a measured complete-event energy and not evidence that an exact 16×16 shape is supported. Saturated pJ/op and isolated-call joules answer different questions.

How to read the evidence badges

measured means workload power or energy was reported, but its scope can still be a macro, accelerator, device, or CPU package. peak / power divides a throughput figure by a TDP, TBP, nominal, or system-power figure; it is not measured energy. modeled denotes post-layout or a deliberately chosen coefficient. None of these scopes are interchangeable.

04 · 8K INT8 extrapolation

A scalable Sutro schedule lands at 119 mJ. Blackwell B200 estimates 244 mJ.

For a full square C[8192×8192] = A[8192×8192] · B[8192×8192], the workload is 549,755,813,888 MACs, or 1,099,511,627,776 conventional operations. The numbers below assume INT8 inputs and INT32 accumulation/output for the hardware comparison.

0.119 Jretuned scalable Sutro schedule at 1 fJ per scored grid-step; movement only
0.244 JNVIDIA B200: 4.5 dense POPS at 1,000 W; peak/TDP estimate
7.92 JNVIDIA Rubin: preliminary 250 dense TOPS at a stated 1,800 W/GPU assumption

“8K model” is ambiguous. This section computes a synthetic full 8192³ GEMM. One-token autoregressive decode usually multiplies a 1×8192 activation by an 8192×8192 weight matrix—8,192× less arithmetic—and is much more likely to be memory-bound, so its energy cannot be obtained by simply dividing these joules by 8,192.

How the 0.118953 J Sutro estimate was derived #

This is an exact score for one parameterized schedule family, followed by the chosen 1 fJ/unit conversion. It is not the 66,300-point 16×16 record multiplied by a scale factor.

The checked-in 66,300-point record is a final physical-address program specialized to 16×16; it has no generator that can be rerun at 8K. Cubically scaling 66,300 gives 8.90 mJ, but incorrectly freezes average access distance while the live memory footprint grows. Instead, the calculation generalizes the repository’s sa_cache loop nest to N = 8192 and searches every tile pair whose dimensions divide N.

1 · Generalize the schedule

For every output tile (bi,bj) and reduction index k, load Tj B values into sB. Then stream each of the Ti A values through one near-ALU sA cell, multiply across the B strip, and accumulate a Ti×Tj output tile in sC.

2 · Retune for the larger footprint

Because 8192 = 213, it has 14 divisors. Evaluating all 14×14 = 196 legal Ti,Tj pairs selects 256×128. The formula reproduces the checked-in 16×16 sa_cache score of 73,602 at its original 8×4 tiles.

The sweep is exhaustive within this literal two-parameter sa_cache family, not across all possible matmul algorithms. The next two shapes are 128×128 at 121,105,737,911,286 and 256×64 at 121,681,290,169,654 score units. The 8K result is evaluated in closed form; it does not materialize a trillion-scale IR.

Exact scorer equation

Let d(a)=ceil(sqrt(a)) be the cost of reading address a, and let D(R) be the sum of d(a) over every address in region R. Cells are packed by descending reads per cell—sA, TMP, sB, sC, A, B, C—which is optimal for this fixed schedule because d(a) never decreases with address.

S = N³d(1) + N²(N−1)d(2)
+ (N³/Tj)D(sB) + (N³/(TiTj))D(sC)
+ (N/Tj)D(A) + (N/Ti)D(B) + D(C)

D([l,h]) = F(h) − F(l−1)
F(x) = q(q+1)(4q−1)/6 + (x−q²)(q+1),  q = floor(√x)

Substituting N = 8192, Ti = 256, and Tj = 128 gives this exact ledger:

Retuned sa_cache address layout and score contribution
Region and address rangeCellsPaid reads / cellScore unitsAt 1 fJ/unit
sA · 11549,755,813,888549,755,813,8880.000550 J
TMP · 21549,688,705,0241,099,377,410,0480.001099 J
sB · 3–1301284,294,967,2964,514,010,628,0960.004514 J
sC · 131–32,89832,76816,777,21666,997,983,379,4560.066998 J
A bulk · 32,899–67,141,76267,108,8646423,475,391,008,0000.023475 J
B bulk · 67,141,763–134,250,62667,108,8643221,448,665,777,1520.021449 J
C exit · 134,250,627–201,359,49067,108,8641867,899,704,6940.000868 J
Total201,359,4902,205,465,706,496 total reads118,953,083,721,3340.1189530837 J

The last step is only the proposed unit conversion: 118,953,083,721,334 × 10⁻¹⁵ J = 0.118953083721334 J. No throughput or clock rate is used, so the model produces energy but does not produce time. Arithmetic, destination writes, finite SRAM/HBM capacity, data conversion, control, clocks, and leakage remain unpriced.

Sutro 8192³ score extrapolations at exactly 1 fJ per score unit
ScheduleScaling treatmentScore unitsMovement termvs. naive
66,300 record × n³Fixed-distance local-kernel lower bound; not a generated 8K program8,898,635,366,4000.008899 J2,630× lower
Retuned sa_cacheExact analytic score, 256×128 tiles; formula reproduces the checked-in 16×16 program118,953,083,721,3340.118953 J196.7× lower
Retuned square tiles64×64 tiles317,312,291,631,1140.317312 J73.7× lower
Fixed 4×4 tilesLiteral generalization of the small tiled schedule2,136,290,066,409,4742.13629 J11.0× lower
Naive baselineExact closed form under the scorer23,401,284,397,481,98423.4013 J1×

The naive placement grows as Θ(n4). A size-retuned locality schedule is approximately Θ(n3.5) because the scorer’s ceil(sqrt(address)) distance grows with the footprint. The selected schedule averages 53.94 score units per paid read; its 0.118953 J is 108.2 fJ per conventional operation, or 9.24 TOPS/W-equivalent for movement alone.

Blackwell and Rubin: what the official INT8 tables imply #

Blackwell · B200

0.244 J

NVIDIA’s formal B200 datasheet reports 9 sparse INT8 POPS and says dense is half: 4.5 dense POPS. At the maximum 1,000 W TDP, ideal runtime is 0.244 ms and peak/TDP energy is 0.244 J—2.05× the Sutro movement-only term.

Roofline-style estimate, not measured kernel energy or a strict bound. Structured-sparse 9 POPS is deliberately not used.

Rubin · preliminary

7.92 J

NVIDIA currently lists 250 dense INT8 TOPS per Rubin GPU. Combining that with the datasheet’s 1,800 W/GPU application-model assumption gives 4.398 ms and 7.916 J—66.6× the Sutro movement-only term.

The 1,800 W figure is an application-comparison assumption, not a formally labeled TDP. Rubin specifications are preliminary.

The Rubin result is intentionally counterintuitive. The same preliminary product table emphasizes 50 PFLOPS NVFP4 and 17.5 PFLOPS FP8, but those formats are not interchangeable with exact INT8×INT8→INT32 work. A 72-GPU rack cross-check—18 dense INT8 POPS at 136 kW provisioned power—gives 8.31 J per GEMM when amortized across 72 simultaneous GEMMs. It is consistent with, but does not measure, the 7.92 J per-GPU estimate.

SKU names matter: NVIDIA’s current Blackwell Ultra B300 table lists only 153.5 dense INT8 TOPS and up to 1,100 W, which implies 7.88 J. That is radically different from B200’s 0.244 J and reflects the published legacy-INT8 specification—not a general claim that B300 or Rubin is less efficient on its target FP4/FP8 workloads.

Four bytes and four bits are different Blackwell paths #

Taking “four byte” literally means FP32: four bytes per matrix element. NVIDIA publishes 75 TFLOP/s FP32 for one HGX B200 GPU. The same FP32 storage can enter the much faster TF32 Tensor Core path, but TF32 reduces multiply precision. NVFP4 is the other common reading: four bits, or half a byte, plus block-scale metadata.

Literal 4-byte FP32

14.66 J

75 TFLOP/s at the 1,000 W maximum TDP gives 14.660 ms and 14.660 J—123× the Sutro movement-only term.

If “4-bit” was intended

0.122 J

9 dense PFLOP/s NVFP4 at the same 1,000 W gives 0.1222 ms and 0.1222 J. This has different numerical semantics and excludes quantization, scaling, and conversion work.

One B200, one full 8192³ GEMM: precision sensitivity from published dense rates
Stored input / math pathPublished dense rateMinimum logical dataIdeal timePeak/TDP energy
NVFP4 · 4-bit block-scaled9 PFLOP/s Tensor Cores64 MiB inputs, plus scales and higher-precision output0.1222 ms0.1222 J
INT8 · 1 byte4.5 POPS Tensor Cores384 MiB with INT32 output0.2443 ms0.2443 J
TF32 math · FP32 storage1.1 PFLOP/s Tensor Cores768 MiB for FP32 A, B, and C0.9996 ms0.9996 J
FP32 · literal 4 bytes75 TFLOP/s FP32768 MiB for FP32 A, B, and C14.660 ms14.660 J

All four B200 rows divide the same 1.0995 trillion conventional operations by NVIDIA’s published dense rate and multiply by the 1,000 W maximum TDP. They are matched roofline-style estimates, not measured per-precision power. FP32/TF32 traffic assumes C=A·B without reading a prior C matrix. For B300’s 1,100 W maximum TDP, the corresponding strict-FP32 and TF32 estimates are 16.13 J and 1.10 J; its per-GPU dense-FP4 figure gives 0.086 J.

8K comparison with evidence boundaries #

0.118953 J

Provenance of the Sutro row below: the exact 8K sa_cache score is 118,953,083,721,334 units after an exhaustive 196-pair tile sweep selects 256×128. Multiplying that score by the chosen 1 fJ/unit coefficient yields 0.118953 J. This is a movement-only model result—not measured chip energy and not a throughput/TDP estimate.

Read the complete derivation and score ledger above · Direct table permalink: #eightk-comparison

One full 8192³ dense INT8 GEMM (1.0995 trillion conventional operations)
Hardware or modelPublished or modeled basisEvidenceIdeal / derived timeEnergy
Sutro retuned sa_cacheExact score at 1 fJ per operand-grid-stepmovement onlynot modeled0.11895 J
NVIDIA B200 Blackwell4.5 dense POPS; 1,000 W maximum TDPpeak / TDP0.244 ms0.244 J
AMD MI300X2,614.9 dense INT8 TOPS; 750 W maximum TBPpeak / TBP0.420 ms0.315 J
NVIDIA H100 SXM1,979 dense INT8 TOPS; 700 W maximum TDPpeak / TDP0.556 ms0.389 J
Google TPUv192 dense INT8 TOPS; 40 W typical incremental busy powerpeak / busy W11.95 ms0.478 J
Tesla FSD, one chip72 INT8 TOPS; 40 W upper TDP; includes 96×96 edge paddingpeak / TDP15.51 ms≈0.620 J
NVIDIA T4, pre-transformed layoutExact 8192³ cuBLASLt graph-read ≈95 TOPS at measured 70 W card powerexact-size proxy≈11.57 ms≈0.81 J
Sutro endpoint sensitivity2.5 fJ × (S + 160R) on the retuned schedulemovement + readsnot modeled1.180 J
NVIDIA T4, standard layoutExact 8192³ including required device-side layout transformation; ≈55 TOPS at 70 Wexact-size proxy≈19.99 ms≈1.40 J
AMD VCK190 AutoMMMeasured repeated 16K INT8 energy/op applied to 8K worksize-scaled proxy39.06 ms2.38 J
NVIDIA A100 cuBLASMeasured repeated 16K INT8: 67.2 TOPS at 248.08 W; same energy/op applied to 8Ksize-scaled proxy16.36 ms4.06 J
NVIDIA B300 Blackwell Ultra153.5 dense INT8 TOPS; 1,100 W maximum TDPpeak / TDP7.16 ms7.88 J
NVIDIA Rubin250 dense INT8 TOPS; 1,800 W/GPU application-model assumptionpreliminary4.40 ms7.92 J
Sutro naive baselineExact closed-form scorer extrapolationmovement onlynot modeled23.40 J

Only the T4 rows use an exact 8192³ published run paired with measured card power—and even those joules are derived from a graph-read throughput. No primary source located reports integrated exact-8K INT8 kernel joules for B200, B300, Rubin, H100, MI300X, TPUv1, or Tesla FSD. Their rows are roofline-style estimates, not measurements.

Two INT8 inputs occupy 128 MiB and an INT32 output another 256 MiB, for at least 384 MiB of logical device traffic. Sutro’s widthless, unbounded 2-D memory charges none of the HBM transfers, writes, arithmetic, clocks, leakage, or dispatch. Also, the repository verifier currently uses unbounded Python integers, so this hardware table is a scale comparison rather than semantic equivalence.

Reproducibility: download the 8K comparison CSV, inspect the machine-readable derivations, or run the standalone scaling script.

05 · Calibration and extension

Distance optimization survives. Absolute energy needs more terms.

The 1 fJ coefficient is best treated as a transparent hypothesis. It answers “what if link movement costs 1 fJ per operand-grid-step?” It does not calibrate SRAM activation, arithmetic, or the physical length and width of a wire.

Sensitivity to the earlier endpoint-offset idea

Sensitivity to movement and fixed read-endpoint coefficients
ModelBaselineRecordImprovementInterpretation
1 fJ × S340.704 pJ66.300 pJ5.14×pure scored movement
2.5 fJ × (S + 160R)7.30296 nJ7.27815 nJ1.0034×movement plus fixed cost for every paid read

The second model charges 400 fJ of fixed endpoint energy per paid read. Because the record intentionally performs more reads, that term nearly erases its distance advantage. This does not prove that either calibration is correct; it proves that read count and read distance must remain separately visible.

Etotal = NMACeMAC + Σlevels(Nreaderead + Nwriteewrite)
+ Σlinks(bits · hops · ebit-hop) + Pleakt

Arithmetic sanity check

Horowitz’s illustrative 45 nm table assigns about 0.2 pJ to an 8-bit multiply and 0.03 pJ to an 8-bit add. The repository’s 4,096 multiplies plus 3,840 adds would therefore be about 934.4 pJ for arithmetic alone—14.1× the record movement proxy. It is an era- and precision-specific scale marker, not a modern chip prediction.

Physical calibration path

Use CACTI or characterized SRAM macros for endpoint reads/writes; extracted interconnect or router models for bit-hops; a synthesis/library estimate for arithmetic; and measured or simulated time for leakage. Name the transfer width: fJ/(operand·step) is abstract, while fJ/(bit·mm) or fJ/(byte·hop) is physically interpretable.

What a stronger compiler report should export

Recommended evolution of the compiler’s cost report
LayerCounts to retainClosest established analogue
Sutro nowpaid reads, read role, address tier, weighted distance, instruction countsone unicast movement action class
Next compiler schemareads and writes by memory level, operand bits, source→destination hops, multicast fanout, cyclesTimeloop statistics
Energy bindingenergy per read/write/MAC/router/link action, with technology and voltage metadataAccelergy ERT
Mapping explorationreuse, buffer capacity, NoC traffic, latency, bandwidthMAESTRO
Silicon validationinstructions, requests, sectors, cache/DRAM bytes, duration, sampled energyNsight Compute / PCM / RAPL

XLA’s “cost” is semantic FLOPs and bytes; TVM’s cost model generally predicts schedule latency; Timeloop’s counts are actions; Accelergy turns actions into modeled joules. Keeping those namespaces separate prevents an attractive unit conversion from acquiring more physical meaning than it has.

06 · Sources and method

Primary sources, explicit derivations.

All hardware conversions use one MAC = two conventional operations: 8,192 operations at n = 16 and 1,099,511,627,776 operations at n = 8,192. Reported sparse peaks were excluded. Values marked measured retain the source’s measurement boundary; peak/TDP values are explicitly labeled derived. Sources were reviewed 28 August 2026.

  1. Sutro scorer implementation and 66,300 record audit.
  2. Reproducible 8K closed-form scorer and the report’s derivation audit.
  3. NVIDIA B200 Blackwell datasheet, B300 Blackwell Ultra datasheet, and the current HGX specification table.
  4. NVIDIA Vera Rubin preliminary specifications, datasheet, and rack-power analysis.
  5. Exact 8192³ INT8 T4 study (DOI).
  6. AutoMM measured INT8 efficiency on A100 and VCK190.
  7. Tesla FSD computer, IEEE Micro; Hot Chips 31 slides.
  8. In-Datacenter Performance Analysis of a TPU.
  9. NVIDIA H100 Tensor Core GPU datasheet.
  10. A100 matmul power measurement study.
  11. AMD Instinct MI300X datasheet.
  12. Dissecting the Graphcore IPU Architecture.
  13. Apple Neural Engine reverse-engineering and energy study.
  14. Eyeriss row-stationary accelerator.
  15. NM-Carus near-memory SIMD accelerator.
  16. Intel 22FFL analog SRAM compute-in-memory macro.
  17. Energy efficiency of matrix multiplication on CPUs and GPUs.
  18. Direct 16×16 CPU/FPGA energy study.
  19. Horowitz, computing’s energy problem.
  20. CACTI memory and interconnect model.
  21. Timeloop, Accelergy, and MAESTRO.

Limitations: no cited chip implements the exact Sutro abstract machine or verifier contract. Published energy varies with matrix shape, data values, sparsity, placement, voltage/frequency, batching, and measurement scope. The normalized table is a scale comparison, not a ranking of products.