What I learned building memory-aware inference on Apple Silicon

ENGINEERING NOTES · OCTOBER 4, 2026

Strata explores a practical inference question: how should a runtime choose an execution plan when model weights, KV state, and the machine's available memory compete for the same budget?

Budget the whole request

Model weights are only one part of memory demand. Cached keys and values grow with the sequence. The planner therefore needs a request-specific budget, rather than a static claim that a model fits. Strata's baseline combines admission, KV state, and byte-budgeted weight residency.

  1. Admit a safe budget
  2. Load & pin needed weights
  3. Compute with qualified backend
  4. Synchronize & release residency

Correctness comes before promotion

The repository reports cached-decode logits checked against full causal forward passes in deterministic two-layer fixtures. Attention and RoPE have NumPy reference comparisons. These checks address whether the implementation preserves the calculation before measuring how quickly it runs.

A kernel result is a bounded result

The local M1 verification includes direct ANE dispatch and CPU numerical parity for one fixed fp16 projection. That establishes evidence for a particular operation. It leaves whole-model execution, data movement, scheduling, and end-to-end performance to be established separately.

Keep a fallback and label the scope

Strata uses evidence-scoped planning with an MLX fallback. Its README separates the verified baseline from production serving, complete native model execution, and SSD-paging targets. That separation is useful: it allows progress on a small component without implying that the complete runtime is ready.

My takeaway is to make execution decisions depend on measured workload evidence, and to carry the limits of that evidence into the claim. Fixture parity, bounded hardware dispatch, and a full-model benchmark each establish something different.

Based on the repository's documented verification scope reviewed October 4, 2026. No full-model speedup is claimed and no benchmarks were rerun for this article.

Strata source and verification scope · Case study


Get in touch