Blog

Engineering notes on evaluation, local inference, and building useful AI systems.

What I learned building memory-aware inference on Apple Silicon

OCTOBER 4, 2026

Memory budgets, fixture correctness, and why bounded native dispatch does not establish a whole-model speedup.


Get in touch