Memory cost and capacity are significant issues for AI accelerators. Unlike game rendering, model i...
By @ID_AA_Carmack
John Carmack lays out a detailed technical argument that model inference has deterministic memory access, so NAND flash (far cheaper than HBM) could feed accelerator scratchpads via a specialized pipelined page-transfer protocol tolerant of millisecond cold starts.