Checkpoint 013 - provider matrix

Featherless is no longer
a testing dependency.

The exact promoted Nagato runtime completed active-memory proofs through Nous free routes, Ollama, and LM Studio. Every accepted route changed the decision only when the canonical run-bound Mnemosyne receipt was present.

6

Accepted routes

Four Nous free models and two local models completed the exact baseline, exact memory-backed answer, canonical trace, and receipt-binding gates.

3

Excluded routes

Ling failed baseline discipline, local Qwen 3.6 exceeded the bounded deadline, and the LM Studio 4B model ignored both exact-answer constraints.

Selection matrix

Route
Baseline
Memory
Decision
tencent/hy3:free
pass
pass
default
stepfun/step-3.7-flash:free
pass
pass
fallback
poolside/laguna-xs-2.1:free
pass
pass
fallback
poolside/laguna-s-2.1:free
pass
pass
transient retry
Ollama Qwen3-VL 4B 64k
pass
pass
local fallback
LM Studio Holo 9B 64k
pass
pass
local fallback

GPU result

The capped Ollama 4B route ran entirely on GPU at 65,536 tokens and passed. The uncapped Qwen 3.6 route allocated a 262k context across GPU and CPU and remained active past 180 seconds.

Harness corrected

The shadow launcher now pins the actual Python import path to the selected detached release. Redacted failure logs survive cleanup, and local keyless endpoints are first-class matrix inputs.

Current recommendation

Keep tencent/hy3:free as the live default. Use StepFun and Laguna XS as remote fallbacks. Keep capped Ollama Qwen3-VL 4B and LM Studio Holo 9B available as local fallback lanes. Do not route production work to Featherless, Ling, uncapped local Qwen 3.6, or the tested LM Studio 4B model.

Next subwave

Close the remaining Mnemosyne/GTC release gaps, reconcile the runtime-separation PR with current main, and complete independent review before asking for human gameplay testing.

Generated 25 July 2026, Pacific/Auckland. Sanitized, unlisted, noindex, and not authenticated. No provider secrets or raw model output are included.