Experiments
Track A — gpt-oss 20B context ladder
Verdict: partial pass. On Owlforge’s RTX 3090 Ti, gpt-oss:20b passed strict five-needle retrieval at 17,445 and 68,916 prompt tokens. At 107,832 prompt tokens it recovered all five values internally and emitted all five exact values, but violated the preregistered strict-JSON contract. That cell is therefore scored 0/5, not rescued after the fact.
| Nominal archive | Actual prompt tokens | Strict score | Wall time | Peak VRAM | Peak power | Verdict |
|---|---|---|---|---|---|---|
| 16,000 words | 17,445 | 5/5 | 2.384 s | 15,918 MiB | 380.63 W | PASS |
| 64,000 words | 68,916 | 5/5 | 16.658 s | 15,918 MiB | 474.16 W | PASS |
| 100,000 words | 107,832 | 0/5 | 30.081 s | 15,918 MiB | 475.17 W | FAIL — output contract |
The 107,832-token response was not a retrieval miss: the reasoning trace identified every needle correctly, and the final channel contained each value. It returned newline-separated text instead of the required JSON object. The preregistered evaluator therefore failed it closed.
Two earlier runs are recorded as infrastructure failures rather than model-quality results:
- v1: Ollama 0.32.5’s JSON-Schema response mode exposed an empty final channel for this model.
- v2: the 256-token generation budget truncated an otherwise correct final object.
Public receipt: experiments/track-a/public/gpt-oss-20b-v3-receipt.json
Receipt SHA-256: e4ef0db25d7aee05f2eaecbbd5b5a6e43ac7cde01d67d658ea49f209292ab148
This control measures exact synthetic retrieval only. It does not establish long-context summarization, reasoning, tool use, or million-token usefulness.
Iterative useful-context wave — Qwen 1M and gpt-oss controls
The next wave did not merely extend the retrieval ladder. Each result selected the next causal discriminator or a different capability class: output contracts, amendment tracking, distributed synthesis, chronology, adversarial instruction hierarchy, multi-hop state tracking, and KV-cache efficiency.
| Track | Capability | Prompt tokens | Preregistered verdict | Evidence |
|---|---|---|---|---|
| B | Five-needle retrieval | 18,010 | PASS 5/5 | Strict JSON |
| B | Five-needle retrieval | 71,353 | PASS 5/5 | Strict JSON |
| B | Five-needle retrieval | 142,542 | FAIL 0/5 | All 5 values present; malformed JSON |
| C | Output-contract discriminator control | 71,383 | INVALIDATED | New contract failed at the known-green control; no boundary cell run |
| D | Corrections + interference | 71,671 | FAIL 0/5 strict | Four amended values semantically correct; one stale initial value |
| E | Distributed ledger synthesis | 36,817 / 35,482 | INFRASTRUCTURE-FAILED | Qwen exposed no answer; gpt-oss hit 1,024 and 4,096-token reasoning caps |
| F | Chronology + conflicting memos | 53,985 | FAIL | Qwen returned revoked FROST-2509, not current TOKEN-GOLD-4821 |
| G | 24 embedded prompt injections | 36,494 | FAIL strict | Returned TOKEN TRUSTED-OWL-7391; trusted value recovered, bare-token contract missed |
| H | Eight-hop variable tracking | 36,163 | FAIL | Returned filler token willow |
| H control | Same task at short context | 9,463 | FAIL | Same willow result; no long-context boundary claim |
| I | Real q4_0 KV-cache sidecar | 18,010–142,542 | FAIL | Incoherent output at all three cells, including the short sanity control; q4_0 rejected |
| J | Predicate-selective retrieval | 75,170 | FAIL | Returned 1 of 5 matching records |
| J control | Same predicate task at short context | 12,783 | FAIL | Four true matches plus near-miss distractors; no context boundary claim |
| K | Cross-document consistency | 36,756 | FAIL | Qwen selected M01, not defective module M17 |
| K control | Same relational task at short context | 10,105 | FAIL | Qwen selected M24; no long-context boundary claim |
| K gpt-oss control | Same short relational task | 9,642 | INFRASTRUCTURE-FAILED | Empty final channel after 2,048 reasoning tokens; no cross-model quality claim |
| L | Minimal two-document relational floor | 9,027 | FAIL | Qwen selected M01, not uniquely inconsistent M02, with two modules and no comments |
| N | Minimal corrected-ledger synthesis floor | 9,033 | FAIL | Qwen returned visible integer 235, not exact total 464 |
| O | Byte-identical minimal-ledger model control | 8,762 | PASS | gpt-oss returned exact total 464 at the same 256-token budget |
| P | Context-allocation resource frontier | 71,353 | FAIL | 5/5 semantic values, but non-JSON; peak VRAM fell 41.136% |
What changed
The headline retrieval number did not generalize to stateful use. Qwen’s strict five-needle retrieval remained clean through 71,353 prompt tokens, but amendment tracking retained one revoked/stale state, chronology selected a revoked predecessor, and eight-hop derivation failed even in the 9,463-token short control. The prompt-injection trial was semantically encouraging—24 hostile archive instructions were ignored and the trusted value was recovered—but it still failed the exact output contract.
The arithmetic-synthesis family produced no valid model-quality comparison. Qwen emitted no visible answer after three generated tokens, while gpt-oss consumed bounded reasoning budgets of 1,024 and 4,096 tokens without reaching a final answer. Those cells remain infrastructure failures rather than scored zeros.
The first KV-cache trial was invalidated because --kv-cache q4_0 changed receipt metadata but did not reconfigure the running q8_0 server. A corrected isolated Ollama sidecar then proved the real configuration at process launch with --cache-type-k q4_0 --cache-type-v q4_0 -c 1010000. It still produced incoherent output and recovered 0/5 needles at 142,542, 71,353, and 18,010 server-observed prompt tokens. Because the short sanity cell also failed, this q4_0 configuration is rejected as globally unusable for the frozen task; it does not establish a useful-context boundary or an extended frontier. The predicate-filtering trial was valid and failed: at 75,170 prompt tokens, Qwen returned only the first of five matching records. Its 12,783-token short control also failed, finding four true matches while including near-miss distractors, so no filtering context boundary is claimed.
The cross-document consistency task required reconciling module records, review comments, and one defect report. Qwen chose the wrong module at both 36,756 and 10,105 prompt tokens, falsifying a long-context-only explanation and locating the failure at the task/model relational floor. The gpt-oss short control emitted no final answer before its 2,048-token reasoning cap, so that cell remains infrastructure-failed: it neither rescues nor falsifies gpt-oss quality, and no cross-model comparison is claimed. The next discriminator must change the output-channel/budget axis or reduce the join to a minimal two-document floor rather than repeat a Qwen length ladder.
Track L reduced that join from 24 modules plus 24 comment distractors to exactly two authoritative modules and no comments while holding the 8,000-word filler, model digest, q8_0 KV runtime, seed, and evaluator fixed. Version 1 aborted before inference when model discovery failed and remains an infrastructure receipt with no quality claim. The separately preregistered v2 ran once and failed closed: at 9,027 prompt tokens, Qwen returned M01 instead of M02. The run took 33.377 s, peaked at 21,652 MiB VRAM and 216.11 W, and scored failure despite parseable JSON because the identifier was wrong. This falsifies join-load/comment interference as a sufficient explanation and supports absence of a reliable relational floor for this task on the pinned Qwen/Ollama configuration. The family now stops; another length or quality retry would not be informative.
Track N rotated back to the unresolved distributed-synthesis class and reduced Track E’s load from 40 canonical records, 10 amendments, and 20 shadows to three canonical records, one amendment, and one shadow. After preserving a pre-inference endpoint-discovery abort separately, the preregistered Owlforge-local cell ran exactly once. At 9,033 prompt tokens, Qwen emitted a normal final answer (done_reason=stop) but returned 235 instead of the exact corrected total 464. The run took 31.756 s, peaked at 21,652 MiB VRAM and 218.3 W. This falsifies task load as a sufficient explanation and rules out Track E’s empty-answer mechanism for the minimal cell; it supports absence of a reliable corrected-synthesis floor on this pinned model/runtime. No quality retry is permitted.
Track O changed only the pinned model artifact while preserving Track N’s byte-identical prompt, 3+1+1 manifest, expected total, seed, q8_0 KV runtime, 256-token output budget, and strict evaluator. gpt-oss returned the exact integer 464 (done_reason=stop) at 8,762 server-observed prompt tokens in 6.764 s, peaking at 15,918 MiB VRAM and 436.42 W. This falsifies both a model-general task failure and an inevitable gpt-oss answer-channel failure at the frozen budget. The contrast supports a model-specific Qwen corrected-synthesis failure at this minimal floor. The family stops here; the next test rotates capability class rather than retrying quality.
Track P rotated to the protocol’s quality/resource-frontier class at Track B’s proven 71,353-token prompt. It changed only the context allocation from 1,010,000 to 131,072 tokens. Peak VRAM fell from 21,650 to 12,744 MiB (41.136%) and wall time from 130.937 to 26.821 s (79.516%), but strict quality failed closed: all five values were exact while the response used line-oriented text instead of the required JSON object. This falsifies the preregistered claim that the smaller allocation preserves the full strict quality point. The result is a contract/resource tradeoff, not a Pareto replacement, and Track P stops after this one cell rather than becoming an allocation ladder.
Public machine receipt: experiments/public/context-school-iterative-wave-20260805.json
Track N receipt: experiments/track-n/public/qwen25-minimal-ledger-8000.json
Track O receipt: experiments/track-o/public/gpt-oss-20b-minimal-ledger-8000.json
Track P receipt: experiments/track-p/public/qwen25-context-allocation-131072.json