Findings — Structure-Aware Retrieval (Registered Run)
Status: registered findings. Every number in §§1–5 is recomputable from the committed artifacts; the decision in §1 applies the preregistered §12.8 rule before any interpretation, and the one clause §12.8 left unoperationalized (recall ≈) is called out as interpretive where it is used. §7 is explicitly not part of the registered findings.
| Run record | |
|---|---|
| Run ID | 2026-09-01T20-46-12-716Z (registered mode) |
| Artifacts | experiments/results/registered/2026-09-01T20-46-12-716Z/ (run.json + scores.json), commit 7d1f1b6, exactly as produced |
| Gold freeze | a87dde5 (24 questions, schema v2; r2+r3 audited, check 0 blind-admitted 24/24) |
| Token budget | B = 1400 (calibrated, r4.1); measured package tokens 565–1352, 0 over-budget |
| Reps / order | 3 reps, generation order randomized, seed 20260901 |
| Models | generation + judges: claude-sonnet-4-6, temp-0 judges; RAGAS via local sidecar |
| Cells | 298 total: 267 generated; 24 C0 retrieval-only (§12.6); 7 C-rel retrieval-only where its package was identical to C's |
| Integrity | RAGAS 0/267 null, 0 judge soft-fails, provenance coverage 100%, constraint cache primed pre-run |
Conditions: A flat dense (topK 7) · B dense + applicability filter
(topK 7) · C0 filtered seeds S=4, no edges (control) · C structure-aware
primary (S=4 + typed closures + supersession replacement) · C-rel C plus
related edges (stress arm; reported, never gating).
1. Registered decision (§12.8 applied first)
Continue (deepen) requires C > A on requirement recall for the majority of
questions in each of prerequisite and exception where n ≥ 3.
prerequisite(n = 4): C > A on q2a (1.00 vs 0.50); ties on q2b, q2c (1.00 = 1.00); C < A on q5c (0.00 vs 0.50). 1 of 4 — no majority. Fails.exception(n = 3): q3a, q3b, q3c all tied at 1.00. 0 of 3 — fails. (The registered headline row excluding q3c has n = 2 and does not gate.)- The remaining Continue sub-criteria happen to pass — the C−C0 gap is entirely structure-attributable (§5.2), C's irrelevant-evidence rate 0.647 is better than B's 0.741, and there are no N3 regressions — but the recall-majority criterion fails, so Continue does not fire.
Stop the graph, keep the filter requires B ≈ C ≈ C0 on recall while B > A on applicability.
- Recall: B 0.952, C 0.929, C0 0.905 — a 4.7-point spread. §12.8 does not operationalize ≈ numerically, so this clause cannot be applied mechanically; the equivalence reading is stated transparently instead: the spread is under half of the only tolerance §12.8 quantifies anywhere (the ~10-point irrelevance clause), and C sits between B and C0 rather than beyond either. We read that as ≈.
- Applicability: B 1.000 vs A 0.542 — unambiguous. Fires under the stated equivalence reading.
Stop (publish negative) requires C ≈ A everywhere, or C winning only on the null-control subset, or an abstention regression. None hold: C beats A outright on q2a and on applicability everywhere; the null-control subset is flat across conditions; abstention is 9/9 for every condition. Does not fire.
Registered decision: stop the graph, keep the filter.
Honest applicability metadata captured essentially all of the value on this corpus at this budget. The typed graph produced one clean, fully attributable evidence recovery (q2a) and one symmetric loss (q5b/q5c seed starvation), netting slightly below the filtered baseline. Per §12.8 this is published as is; any re-run with adjusted parameters is exploratory (§5 control 5).
2. Preregistered prediction scorecard
Hypotheses and nulls from §2 (committed 2026-07, before any code), tag predictions from §12.3 (committed 2026-09-01, before the run).
| Prediction | Verdict | Evidence |
|---|---|---|
| H1 — C > A on required-evidence recall and applicability at equal budget | Half-miss | Applicability confirmed: C 1.000 vs A 0.542. Recall refuted: C 0.929 < A 0.952. |
| H2 — better packages → measurably better answers | No signal at this scale | Completeness A 0.978 / B 0.978 / C 0.948; RAGAS faithfulness 0.949–0.966; judge answerScore ≈ 0.98 everywhere. Differences are within noise except where recall failed outright. |
| H3 — filter captures applicability gains; graph must earn completeness gains or it isn't paying rent | Confirmed — this is the outcome | B fixed 100% of A's applicability failures with zero recall cost; the graph's completeness gains netted ≈ 0 (won q2a, lost q5b/q5c). |
| H4 — edge provenance enables diagnoses flat retrieval cannot express | Supported qualitatively | §5.2 ("prerequisite delivered via edge") and §5.3 ("seed starvation") are statements only expressible with provenance + graph. Label proposal in §7.1 (not registered). |
| N1 — null-control subset flat; C winning here = rigging smell | Held | n = 6: recall 1.000, completeness 1.000, pass 1.000 for A, B, C alike. |
| N2 — related-edge traversal raises irrelevance on noise-prone questions | Held | noise-prone irrel: C-rel 0.786 vs C 0.711; population irrel: C-rel 0.737 vs C 0.647. Strictly worse wherever packages differed; never a recall gain. |
| N3 — every condition abstains on unanswerable questions | Held, perfectly | 9/9 abstentions per condition, all conditions. |
| N3-extended — no condition asserts an unanswerable aspect; expansion must never flip a gap into an assertion | Split | The expansion half held: assertion rates are identical with and without the graph. The flat half failed for every condition — see §5.5. |
| N4 — structure may raise completeness while lowering precision | Inverted | C had the best precision of the full generation conditions (irrel 0.647 vs 0.741 A/B) and slightly lower completeness. |
Per-tag predictions (§12.3):
| Tag prediction | Verdict |
|---|---|
| null-control: A ≈ B ≈ C | Hit (dead flat) |
prerequisite: C > B ≈ A on recall |
Miss — all at 0.750; the q2a win and q5c loss cancel exactly |
exception: C > B ≈ A on recall |
Miss (ceiling) — every condition at 1.000; the authored exceptions were semantically close enough that dense retrieval never needed the edge |
supersession ∧ version-current: B ≈ C > A on applicability |
Hit — B/C 1.000, A 0.125 on the supersession tag |
version-historical: supersession handling must not override an explicit historical constraint |
Hit — q4b at 1.000 recall/applic/pass in every condition (n = 1, observation) |
cross-collection: C ≥ A |
Hit (observation-grade, n = 3) — recall tied 0.833; pass 0.667 vs 0.556 |
unanswerable: all abstain |
Hit — 9/9 everywhere |
partial: all conditions state facts, acknowledge gaps, assert nothing |
Miss for every condition — §5.5 |
noise-prone: C ≈ B ≈ A recall; C-rel strictly worse irrelevance |
Hit — recall 1.000 all; C-rel 0.786 vs C 0.711 |
Legacy §5 category note (superseded by tags, reported for the record): Q6's "C > B > A" was wrong in its C > B clause — B's filter alone fully handled the obsolete-strong-match questions; C's supply-the-successor mechanism added nothing B didn't already achieve.
3. Population table
All 24 questions; generation figures over 3 reps (72 cells/condition; C-rel 58 — identical-to-C packages ran retrieval-only). C0 is retrieval-only by design.
| Condition | Recall | Irrel. rate | Applicability | Completeness | Outcome pass | Asserted-aspect cells |
|---|---|---|---|---|---|---|
| A | 0.952 | 0.741 | 0.542 | 0.978 | 0.903 | 3 |
| B | 0.952 | 0.741 | 1.000 | 0.978 | 0.889 | 3 |
| C0 | 0.905 | 0.594 | 1.000 | — | — | — |
| C | 0.929 | 0.647 | 1.000 | 0.948 | 0.917 | 3 |
| C-rel | 0.908 | 0.737 | 1.000 | 0.917 | 0.863 | 2 |
RAGAS faithfulness (mean): A 0.966 · B 0.949 · C 0.958 · C-rel 0.954. Abstentions: 9 per condition — exactly the 3 unanswerable questions × 3 reps.
4. Per-tag tables
Tags are non-exclusive; a question contributes to every tag it carries. Tags with n ≤ 2 support observations, not findings (§12.3). C0 generation columns are structurally null and left blank.
| Tag (n) | Cond | Recall | Irrel | Applic | Compl | Pass |
|---|---|---|---|---|---|---|
| single-unit (15) | A | 1.000 | 0.761 | 0.533 | 1.000 | 1.000 |
| B | 1.000 | 0.764 | 1.000 | 1.000 | 1.000 | |
| C0 | 1.000 | 0.600 | 1.000 | |||
| C | 1.000 | 0.678 | 1.000 | 1.000 | 1.000 | |
| C-rel | 1.000 | 0.734 | 1.000 | 0.986 | 0.972 | |
| prerequisite (4) | A | 0.750 | 0.656 | 0.500 | 0.867 | 0.500 |
| B | 0.750 | 0.656 | 1.000 | 0.867 | 0.500 | |
| C0 | 0.625 | 0.438 | 1.000 | |||
| C | 0.750 | 0.455 | 1.000 | 0.750 | 0.750 | |
| C-rel | 0.500 | 0.607 | 1.000 | 0.000 | 0.000 | |
| exception (3) | A | 1.000 | 0.595 | 0.333 | 1.000 | 1.000 |
| B | 1.000 | 0.579 | 1.000 | 1.000 | 1.000 | |
| C0 | 1.000 | 0.333 | 1.000 | |||
| C | 1.000 | 0.467 | 1.000 | 1.000 | 1.000 | |
| C-rel | 1.000 | 0.556 | 1.000 | 0.944 | 0.889 | |
| exception, headline excl. q3c (2) | A | 1.000 | 0.607 | 0.500 | 1.000 | 1.000 |
| B | 1.000 | 0.583 | 1.000 | 1.000 | 1.000 | |
| C0 | 1.000 | 0.375 | 1.000 | |||
| C | 1.000 | 0.500 | 1.000 | 1.000 | 1.000 | |
| C-rel | 1.000 | 0.583 | 1.000 | 0.917 | 0.833 | |
| supersession (8) | A | 1.000 | 0.702 | 0.125 | 1.000 | 0.958 |
| B | 1.000 | 0.709 | 1.000 | 1.000 | 0.917 | |
| C0 | 0.938 | 0.594 | 1.000 | |||
| C | 0.938 | 0.646 | 1.000 | 0.969 | 0.875 | |
| C-rel | 0.917 | 0.723 | 1.000 | 0.950 | 0.800 | |
| version-current (11) | A | 0.909 | 0.692 | 0.000 | 0.952 | 0.788 |
| B | 0.909 | 0.679 | 1.000 | 0.952 | 0.758 | |
| C0 | 0.818 | 0.545 | 1.000 | |||
| C | 0.864 | 0.613 | 1.000 | 0.886 | 0.818 | |
| C-rel | 0.820 | 0.687 | 1.000 | 0.821 | 0.714 | |
| version-historical (1) | all | 1.000 | 0.714–0.875 | 1.000 | 1.000 | 1.000 |
| cross-collection (3) | A | 0.833 | 0.542 | 0.000 | 0.889 | 0.556 |
| B | 0.833 | 0.524 | 1.000 | 0.889 | 0.444 | |
| C0 | 0.667 | 0.417 | 1.000 | |||
| C | 0.833 | 0.440 | 1.000 | 0.917 | 0.667 | |
| C-rel | 0.700 | 0.507 | 1.000 | 0.750 | 0.000 | |
| partial (2) | A | 0.750 | 0.670 | 0.000 | 0.900 | 0.333 |
| B | 0.750 | 0.643 | 1.000 | 0.900 | 0.167 | |
| C0 | 0.250 | 0.625 | 1.000 | |||
| C | 0.250 | 0.625 | 1.000 | 0.375 | 0.000 | |
| C-rel | 0.250 | 0.714 | 1.000 | 0.375 | 0.000 | |
| unanswerable (3) | all | — | 0.917–0.944 | 1.000 | 1.000 | 1.000 |
| noise-prone (3) | A | 1.000 | 0.821 | 1.000 | 1.000 | 1.000 |
| B | 1.000 | 0.813 | 1.000 | 1.000 | 1.000 | |
| C0 | 1.000 | 0.667 | 1.000 | |||
| C | 1.000 | 0.711 | 1.000 | 1.000 | 1.000 | |
| C-rel | 1.000 | 0.786 | 1.000 | 1.000 | 1.000 | |
| null-control (6) | A/B/C | 1.000 | 0.725–0.821 | 1.000 | 1.000 | 1.000 |
| C-rel | 1.000 | 0.787 | 1.000 | 1.000 | 1.000 |
Reporting caveat for C-rel generation columns: where C-rel's package was
identical to C's, only a retrieval cell exists, so its per-tag completeness
and pass figures can rest on fewer questions than n suggests (e.g. its
prerequisite generation figures come from q5c alone).
5. Qualitative findings
5.1 The filter fixed every applicability failure and cost nothing
A retrieved a forbidden obsolete unit on 11 of 24 questions (q2a, q3a,
q3c, q4a, q4c, q5a, q5b, q5c, q6a, q6b, q6c) — applicability 0.000 across the
entire version-current tag. B's applicability filter removed every one, and
B's per-question requirement recall is identical to A's on all 24
questions: honest applies_to metadata cost zero recall at topK 7.
An honest wrinkle: A's outcome pass barely suffered (0.903 vs B's 0.889) — with both versions in context, the generator usually picked the current one. The obsolete-context hazard is real but mostly latent at this corpus size; applicability accuracy measures the exposure, not the realized damage.
5.2 q2a — the graph's one clean win, fully attributable
"Our order needs to be rushed — what do we have to do to set that up?"
requires request-rush-production and its prerequisite
approve-artwork-proof (a different collection, semantically distant from
rush vocabulary). In A and B, approve-artwork-proof ranked below the budget
cut (dropped: rank cut, both conditions) → recall 0.50, pass 0/3. C followed
prerequisite_of via request-rush-production and packaged it → recall 1.00,
pass 3/3 — the only condition to answer q2a correctly. (C-rel's package
was identical to C's here.)
C0 (same seeds, no edges) scored 0.50, and q2a is the entire population
C−C0 recall gap (0.929 vs 0.905). Attribution doesn't get cleaner: the
requires edge did exactly what H1 hypothesized — once.
5.3 q5b/q5c — seed starvation is the price of expansion headroom
S = 4 was fixed in the note to leave expansion headroom under B = 1400. On q5c (broken mugs — what
do we do, and does the replacement incur the setup fee again?) the required
units are file-damage-claim and inspect-delivery, but the question's
fee-and-mug vocabulary spent all four seeds on setup-fee-policy-2026,
setup-fee-reorder-exception, minimum-order-quantities, and
drinkware-imprint-areas — and no typed edge leads from any of them to the
damage-claim units. C recall 0.000. A and B, with topK 7, caught
file-damage-claim at rank 7 → 0.50. (inspect-delivery made no condition's
package — a miss shared by everyone.) q5b is the same shape: seeds went to
color/imprint units, setup-fee-policy-2026 sat just below the S = 4 cut, and
C's answers were missing the fee facts (stated pattern TTTf in all three
reps).
The graph can only expand from seeds it holds. The partial tag (recall C
0.25 vs A/B 0.75) is where that bill landed. Per §5 control 5, S and B were
fixed in the note and stand; any re-run with different seeding is exploratory.
5.4 C-rel — related edges gave back everything the typed closures earned
Wherever C-rel's package differed from C's (17 of 24 questions), the
difference was related_to units (q5b gained vector-artwork-basics,
volume-discount-policy, place-first-order; q5c gained
volume-discount-policy, place-first-order, product-material-safety).
The result: population irrelevance 0.737 — back at A's 0.741, erasing the
precision advantage C's typed traversal had built (0.647). It never bought a
recall point anywhere (0.908 ≤ C's 0.929), and on the one prerequisite
question it generated fresh (q5c) it passed 0/3. N2 wasn't just confirmed; it
was measured: the cost of untyped expansion begins with the first hop.
5.5 The fee fabrication — a generation behavior no retrieval condition caused or cured
q5c's unanswerable aspect (does a replacement for carrier-damaged goods incur another setup fee?) is stated nowhere in the corpus. It was asserted in 11 of 12 generated cells — A 3/3, B 3/3, C 3/3, C-rel 2/3. The mechanism is visible in the transcripts (B rep 3): the model chains the reorder stored-artwork exemption into "if the reprint uses your original, unchanged stored artwork … no setup fee would apply" — a plausible policy inference the handbook never makes. Identical rate with and without the filter, with and without the graph: retrieval architecture neither caused nor cured it.
The generic metrics did not reliably catch it. RAGAS on the asserted cells ran
0.70–0.91 (a mild dip from the ≈ 0.95 population mean); the eval judge flagged
missing_information for B/C/C-rel — but scored A's three asserted cells
0.95–0.97 with evidenceSufficient: true and no primary issue. All 11 were
flagged only under the registered aspect-specific check (the asserted
boolean — itself judged by an LLM, but against an explicit per-aspect rule).
Contrast q5b's aspect (the dollar amount): never asserted, but B failed to
acknowledge the gap in 2 of 3 reps. Two distinct partial-answer failure
modes — bluffing over the gap versus silence about it — and only the first is
a fabrication event.
This is the strongest public write-up candidate in the run: an evaluation-design finding squarely inside the sandbox mission ("make the behavior observable"), and it required the schema-v2 aspect booleans to see.
6. Registered conclusion
The current structure-aware strategy does not beat the filtered baseline overall, even though explicit relationships demonstrably recover evidence that similarity and filtering alone can miss. On this corpus, at this budget, honest applicability metadata plus a version filter captured essentially all of the retrieval-quality gains (H3's exact prediction); the typed graph's one clean, provenance-attributable win (q2a) was offset by the seed starvation its own budget economics induced (q5b/q5c). The §12.8 decision is stop the graph, keep the filter, scoped to this corpus, budget, and expansion policy. Abstention discipline was perfect everywhere; the null-control was flat (no rigging smell); and the run surfaced a retrieval-independent fabrication pattern on partially answerable questions that generic evaluation metrics did not reliably flag — RAGAS dipped on those cells and the judge caught some conditions, but only the aspect-specific check identified every fabrication event. The negative result is the deliverable, as §2 committed.
7. Not registered: follow-up candidates
Nothing below was run, and no code, gold, scoring, or interpretation rule was changed after the run. These are ideas for future, separately labeled work.
- H4 diagnosis labels (proposal). Extend the
FailureLabelvocabulary with mechanically derivable, provenance-dependent labels:prerequisite_not_retrieved(arequiresedge exists from a packaged unit to an absent required unit),seed_starvation(required unit neither seeded nor reachable by any typed edge from any seed), andsuperseded_context(forbidden version present in the package). Each is computable from run.json + the manifest alone; the first two are inexpressible without the graph — H4's claim in concrete form. - Hybrid seeding (exploratory re-run candidate). Use B's filtered topK-7 ranking as the seed pool and let typed expansion compete with ranks 5–7 for the remaining budget — directly targets the q5c failure without giving up the q2a win. Any such run is exploratory by §5 control 5.
- Generation-side partial guardrail. The §5.5 fabrication is a prompting/judging problem, orthogonal to retrieval; an answer-time "state what the evidence does not establish" instruction could be A/B-tested against the aspect booleans.
- Prose-degraded arm. Already named a follow-up in §12.10: condition A here is Markdown's well-migrated case; the "prose survives, semantics don't" story needs its own arm.