build1 distinct publisher
Reordering two loop indices beat every cache-tiled version on the same machine
A C benchmark on Apple Silicon puts a plain loop reorder 6.3x over naive and the best tile size 11% behind that, which suggests the first move on a hot loop is a stride check rather than a blocking rewrite.
Publishers:dev.to
Reality
- Evidence55
- Adoption10
- Hype gap+10
- Incentives25