~/llamay $ llamay bench -m qwen2.5:0.5b -gpu -forks 32 -prompt 2048 -gen 32 -reps 1 llamay v0.1.58-1-gb14f550 · Qwen2.5 0.5B Instruct · qwen2 · 7 threads · kernels neon+i8mm · matmul metal:Apple M4 · gemm own kernels · kv f32 weights 373.71 MiB · repack on · context 32768 · prompt 2048 · generate 32 phase rate best note prefill 650.4 tok/s 3.149s one command buffer per 256-token tile decode 41.7 tok/s 768ms 506 dispatches per token, one submission cpu encoding 40ms · gpu execution 3.854s (1% encoding) phase rate best note fork x32 65µs 0 pages allocated (a copy would be 1024) snapshot 129ms 48.01 MiB for 2048 tokens, at f32 restore 74ms vs 3.149s to re-prefill gpu time: staging 149ms · kernels 9.77s · copy-back 191ms Restoring a 2048-token context is 42.6x faster than recomputing it. ~/llamay $ # 0 pages allocated where a copy would have been 1024 — that is the fork