Partitioning MLPs Along the Intermediate Dimension Let the Apple Neural Engine Run Two of Three Guidance Branches
How Diffio 4.0 reached 7.8x (pro) and 10.7x (flash) realtime on an M4 Pro Mac mini by splitting guidance branches across the GPU, the CPU's SME units and the Neural Engine, and what memory, precision and scheduling taught us.