Get the latest tech news

Getting 50 GB/S Back from the Apple Neural Engine


s Back Out of the ANE Introduction An RTL performance erratum in the Apple M3 Neural Engine throttles DRAM weight streaming throughput down to 17–19 GB/s from the nominal 45–60 GB/s, whenever the total weight size is an integer multiple of 1 MiB, which currently affects 7 of ANEMLL’s 15 models. Avoiding the problematic path in the kernel DMA engine's speculative prefetch ring increased Llama 3.2 1B token throughput from 10.0 to 24.3 tokens/s (DRAM usage from 24.7 to 60.0 GB/s), and Qwen3-8B from 1.36 to 2.97 tokens/s (DRAM usage from 22.4 to 48.7 GB/s).

None

Get the Android app

Or read this on Hacker News

Read more on:

Photo of Apple

Apple

Photo of neural engine

neural engine

Related news:

News photo

Apple Says iPhone 18 Pro Max Sold in U.S. Differs in One Way

News photo

Apple nailed iPhone Duo split screen, even as iPad multitasking feels convoluted

News photo

ESR’s iPhone accessory lineup is ready for Apple’s foldable era