Strata: memory-aware inference on Apple Silicon
INDEPENDENT ENGINEERING · OPEN SOURCE · IN DEVELOPMENT
Strata is an Apple-Silicon LLM engine and learning laboratory for adaptive, memory-tiered inference. The objective is to fit useful models within a safe memory budget while retaining practical throughput.
Architecture and ownership
- Request & live memory budget
- Execution planner
- Weight residency & KV state
- MLX baseline or qualified native path
This diagram summarizes the current baseline and routing direction. The repository separates implemented components from targets such as complete native model execution and production SSD paging.
Evidence before speed claims
| Check | Evidence boundary |
|---|---|
| Cached decode versus full causal forward | Deterministic two-layer Llama and DeepSeek fixtures |
| Attention and RoPE numerical comparisons | NumPy reference checks |
| Fixed fp16 projection on the local M1 ANE | Dispatch and CPU numerical parity for one bounded operation |
These are correctness and bounded execution results. They do not establish a complete native model speedup. This website has not rerun the repository tests.
Engineering judgment
The runtime tracks memory budgets, pinned weights, warm-layer reuse, and prefetch. Routing is evidence-scoped, with an MLX fallback. A backend has to be qualified for the workload it receives; supporting one operation is insufficient to promote an entire model.
Remaining work
The README lists persistent production serving, multi-request continuous batching, complete native Metal/ANE model execution, and production SSD paging as targets. Real-model throughput and memory measurements remain distinct from fixture correctness.
Read the code and verification scope · Read the engineering notes