Strata: memory-aware inference on Apple Silicon

INDEPENDENT ENGINEERING · OPEN SOURCE · IN DEVELOPMENT

Strata is an Apple-Silicon LLM engine and learning laboratory for adaptive, memory-tiered inference. The objective is to fit useful models within a safe memory budget while retaining practical throughput.

Architecture and ownership

  1. Request & live memory budget
  2. Execution planner
  3. Weight residency & KV state
  4. MLX baseline or qualified native path

This diagram summarizes the current baseline and routing direction. The repository separates implemented components from targets such as complete native model execution and production SSD paging.

Evidence before speed claims

Repository-reported verification, reviewed October 4, 2026
CheckEvidence boundary
Cached decode versus full causal forwardDeterministic two-layer Llama and DeepSeek fixtures
Attention and RoPE numerical comparisonsNumPy reference checks
Fixed fp16 projection on the local M1 ANEDispatch and CPU numerical parity for one bounded operation

These are correctness and bounded execution results. They do not establish a complete native model speedup. This website has not rerun the repository tests.

Engineering judgment

The runtime tracks memory budgets, pinned weights, warm-layer reuse, and prefetch. Routing is evidence-scoped, with an MLX fallback. A backend has to be qualified for the workload it receives; supporting one operation is insufficient to promote an entire model.

Remaining work

The README lists persistent production serving, multi-request continuous batching, complete native Metal/ANE model execution, and production SSD paging as targets. Real-model throughput and memory measurements remain distinct from fixture correctness.

Read the code and verification scope · Read the engineering notes


Get in touch