# Matrix Multiplication

**Author:** Codex and Cosmin<br>
**Date:** 2026-05-28<br>
**Problem:** 16x16 matmul<br>
**Cost:** 66,300<br>
**IR:** [`best_66300.ir`](best_66300.ir)<br>
**Method:** Claude-assisted simulated annealing over a prior leaderboard physical-address IR

## Summary

This IR computes ordinary `16x16` matrix multiplication: 256 `A` inputs,
256 `B` inputs, and 256 `C` outputs. It keeps the standard arithmetic shape:
4,096 multiplies and 3,840 additions.

The score is dominated by reads. A read from address `x` costs
`ceil(sqrt(x))`; writes and arithmetic themselves are free. This candidate was
produced by iterating with Claude and simulated annealing on top of a prior
leaderboard physical-address submission. It is checked in as a final
physical-address IR rather than as a raw semantic trace, so the companion script
verifies the submitted IR directly.

Compared with the prior promoted `66,400` submission, this artifact uses seven
fewer copy operations and seven fewer paid reads. The exact score improves by
100 points, from 66,400 to 66,300.

## Cost Breakdown By Address Tier

| tier | addrs | reads | cost |
|------|-------|------:|-----:|
| 1 | 1 | 5,169 | 5,169 |
| 2 | 2..4 | 4,991 | 9,982 |
| 3 | 5..9 | 2,359 | 7,077 |
| 4 | 10..16 | 852 | 3,408 |
| 5 | 17..25 | 1,089 | 5,445 |
| 6 | 26..36 | 1,325 | 7,950 |
| 7 | 37..49 | 326 | 2,282 |
| 8 | 50..64 | 110 | 880 |
| 9 | 65..81 | 119 | 1,071 |
| 10 | 82..100 | 133 | 1,330 |
| 11 | 101..121 | 147 | 1,617 |
| 12 | 122..144 | 161 | 1,932 |
| 13 | 145..169 | 114 | 1,482 |
| 14 | 170..196 | 100 | 1,400 |
| 15 | 197..225 | 87 | 1,305 |
| 16 | 226..256 | 93 | 1,488 |
| 17 | 257..289 | 96 | 1,632 |
| 18 | 290..324 | 70 | 1,260 |
| 19 | 325..361 | 74 | 1,406 |
| 20 | 362..400 | 78 | 1,560 |
| 21 | 401..441 | 82 | 1,722 |
| 22 | 442..484 | 44 | 968 |
| 23 | 485..529 | 45 | 1,035 |
| 24 | 530..576 | 47 | 1,128 |
| 25 | 577..625 | 49 | 1,225 |
| 26 | 626..676 | 21 | 546 |
| **total** | | **17,781** | **66,300** |

## Instruction Distribution

| instruction | count | paid reads | read cost |
|-------------|------:|-----------:|----------:|
| `mul` | 4,096 | 8,192 | 14,272 |
| `add` | 3,840 | 7,680 | 26,209 |
| `copy` | 1,653 | 1,653 | 21,565 |
| output exit | 256 | 256 | 4,254 |
| **total** | **9,589 ops** | **17,781** | **66,300** |

Additional shape checks:

- 512 distinct inputs.
- 256 distinct outputs.
- 127 input/output address overlaps.
- 646 read-bearing addresses, contiguous from `1..646`.

## Verification

```bash
/Users/cosmin/miniconda/bin/python matmul/submissions/best_66300.py
/Users/cosmin/miniconda/bin/python matmul/experiments/random_true_matmul_check.py matmul/submissions/best_66300.ir --n 16 --trials 100 --seed 20260528 --min -31 --max 31
```

Observed locally:

```text
best_66300.ir  cost=66,300
matmul/submissions/best_66300.ir: cost=66,300 ok 100 random trials
```
