Matrix multiply executed one memory access per clock under the simplified Dally model: an access to address a costs ⌈√a⌉; writes and arithmetic are free.
The cell at linear index a sits at Manhattan distance ⌈√a⌉ from the core, so ring k holds 2k−1 cells. The naive schedules place the arrays contiguously; the tiled schedule reserves the cheapest cells for its scratchpad and pushes the bulk arrays out.
Compare each algorithm’s memory address and access cost at every clock.