RAGLensResearch

RAGLens Research · Technical record

Structure-aware retrieval — registered findings

The registered findings document, rendered in full and unedited. For the readable version, see the research summary; to see the conditions side by side, see the interactive comparison.

Findings — Structure-Aware Retrieval (Registered Run)

Status: registered findings. Every number in §§1–5 is recomputable from the committed artifacts; the decision in §1 applies the preregistered §12.8 rule before any interpretation, and the one clause §12.8 left unoperationalized (recall ≈) is called out as interpretive where it is used. §7 is explicitly not part of the registered findings.

Run record
Run ID 2026-09-01T20-46-12-716Z (registered mode)
Artifacts experiments/results/registered/2026-09-01T20-46-12-716Z/ (run.json + scores.json), commit 7d1f1b6, exactly as produced
Gold freeze a87dde5 (24 questions, schema v2; r2+r3 audited, check 0 blind-admitted 24/24)
Token budget B = 1400 (calibrated, r4.1); measured package tokens 565–1352, 0 over-budget
Reps / order 3 reps, generation order randomized, seed 20260901
Models generation + judges: claude-sonnet-4-6, temp-0 judges; RAGAS via local sidecar
Cells 298 total: 267 generated; 24 C0 retrieval-only (§12.6); 7 C-rel retrieval-only where its package was identical to C's
Integrity RAGAS 0/267 null, 0 judge soft-fails, provenance coverage 100%, constraint cache primed pre-run

Conditions: A flat dense (topK 7) · B dense + applicability filter (topK 7) · C0 filtered seeds S=4, no edges (control) · C structure-aware primary (S=4 + typed closures + supersession replacement) · C-rel C plus related edges (stress arm; reported, never gating).


1. Registered decision (§12.8 applied first)

Continue (deepen) requires C > A on requirement recall for the majority of questions in each of prerequisite and exception where n ≥ 3.

  • prerequisite (n = 4): C > A on q2a (1.00 vs 0.50); ties on q2b, q2c (1.00 = 1.00); C < A on q5c (0.00 vs 0.50). 1 of 4 — no majority. Fails.
  • exception (n = 3): q3a, q3b, q3c all tied at 1.00. 0 of 3 — fails. (The registered headline row excluding q3c has n = 2 and does not gate.)
  • The remaining Continue sub-criteria happen to pass — the C−C0 gap is entirely structure-attributable (§5.2), C's irrelevant-evidence rate 0.647 is better than B's 0.741, and there are no N3 regressions — but the recall-majority criterion fails, so Continue does not fire.

Stop the graph, keep the filter requires B ≈ C ≈ C0 on recall while B > A on applicability.

  • Recall: B 0.952, C 0.929, C0 0.905 — a 4.7-point spread. §12.8 does not operationalize ≈ numerically, so this clause cannot be applied mechanically; the equivalence reading is stated transparently instead: the spread is under half of the only tolerance §12.8 quantifies anywhere (the ~10-point irrelevance clause), and C sits between B and C0 rather than beyond either. We read that as ≈.
  • Applicability: B 1.000 vs A 0.542 — unambiguous. Fires under the stated equivalence reading.

Stop (publish negative) requires C ≈ A everywhere, or C winning only on the null-control subset, or an abstention regression. None hold: C beats A outright on q2a and on applicability everywhere; the null-control subset is flat across conditions; abstention is 9/9 for every condition. Does not fire.

Registered decision: stop the graph, keep the filter.

Honest applicability metadata captured essentially all of the value on this corpus at this budget. The typed graph produced one clean, fully attributable evidence recovery (q2a) and one symmetric loss (q5b/q5c seed starvation), netting slightly below the filtered baseline. Per §12.8 this is published as is; any re-run with adjusted parameters is exploratory (§5 control 5).


2. Preregistered prediction scorecard

Hypotheses and nulls from §2 (committed 2026-07, before any code), tag predictions from §12.3 (committed 2026-09-01, before the run).

Prediction Verdict Evidence
H1 — C > A on required-evidence recall and applicability at equal budget Half-miss Applicability confirmed: C 1.000 vs A 0.542. Recall refuted: C 0.929 < A 0.952.
H2 — better packages → measurably better answers No signal at this scale Completeness A 0.978 / B 0.978 / C 0.948; RAGAS faithfulness 0.949–0.966; judge answerScore ≈ 0.98 everywhere. Differences are within noise except where recall failed outright.
H3 — filter captures applicability gains; graph must earn completeness gains or it isn't paying rent Confirmed — this is the outcome B fixed 100% of A's applicability failures with zero recall cost; the graph's completeness gains netted ≈ 0 (won q2a, lost q5b/q5c).
H4 — edge provenance enables diagnoses flat retrieval cannot express Supported qualitatively §5.2 ("prerequisite delivered via edge") and §5.3 ("seed starvation") are statements only expressible with provenance + graph. Label proposal in §7.1 (not registered).
N1 — null-control subset flat; C winning here = rigging smell Held n = 6: recall 1.000, completeness 1.000, pass 1.000 for A, B, C alike.
N2 — related-edge traversal raises irrelevance on noise-prone questions Held noise-prone irrel: C-rel 0.786 vs C 0.711; population irrel: C-rel 0.737 vs C 0.647. Strictly worse wherever packages differed; never a recall gain.
N3 — every condition abstains on unanswerable questions Held, perfectly 9/9 abstentions per condition, all conditions.
N3-extended — no condition asserts an unanswerable aspect; expansion must never flip a gap into an assertion Split The expansion half held: assertion rates are identical with and without the graph. The flat half failed for every condition — see §5.5.
N4 — structure may raise completeness while lowering precision Inverted C had the best precision of the full generation conditions (irrel 0.647 vs 0.741 A/B) and slightly lower completeness.

Per-tag predictions (§12.3):

Tag prediction Verdict
null-control: A ≈ B ≈ C Hit (dead flat)
prerequisite: C > B ≈ A on recall Miss — all at 0.750; the q2a win and q5c loss cancel exactly
exception: C > B ≈ A on recall Miss (ceiling) — every condition at 1.000; the authored exceptions were semantically close enough that dense retrieval never needed the edge
supersessionversion-current: B ≈ C > A on applicability Hit — B/C 1.000, A 0.125 on the supersession tag
version-historical: supersession handling must not override an explicit historical constraint Hit — q4b at 1.000 recall/applic/pass in every condition (n = 1, observation)
cross-collection: C ≥ A Hit (observation-grade, n = 3) — recall tied 0.833; pass 0.667 vs 0.556
unanswerable: all abstain Hit — 9/9 everywhere
partial: all conditions state facts, acknowledge gaps, assert nothing Miss for every condition — §5.5
noise-prone: C ≈ B ≈ A recall; C-rel strictly worse irrelevance Hit — recall 1.000 all; C-rel 0.786 vs C 0.711

Legacy §5 category note (superseded by tags, reported for the record): Q6's "C > B > A" was wrong in its C > B clause — B's filter alone fully handled the obsolete-strong-match questions; C's supply-the-successor mechanism added nothing B didn't already achieve.


3. Population table

All 24 questions; generation figures over 3 reps (72 cells/condition; C-rel 58 — identical-to-C packages ran retrieval-only). C0 is retrieval-only by design.

Condition Recall Irrel. rate Applicability Completeness Outcome pass Asserted-aspect cells
A 0.952 0.741 0.542 0.978 0.903 3
B 0.952 0.741 1.000 0.978 0.889 3
C0 0.905 0.594 1.000
C 0.929 0.647 1.000 0.948 0.917 3
C-rel 0.908 0.737 1.000 0.917 0.863 2

RAGAS faithfulness (mean): A 0.966 · B 0.949 · C 0.958 · C-rel 0.954. Abstentions: 9 per condition — exactly the 3 unanswerable questions × 3 reps.


4. Per-tag tables

Tags are non-exclusive; a question contributes to every tag it carries. Tags with n ≤ 2 support observations, not findings (§12.3). C0 generation columns are structurally null and left blank.

Tag (n) Cond Recall Irrel Applic Compl Pass
single-unit (15) A 1.000 0.761 0.533 1.000 1.000
B 1.000 0.764 1.000 1.000 1.000
C0 1.000 0.600 1.000
C 1.000 0.678 1.000 1.000 1.000
C-rel 1.000 0.734 1.000 0.986 0.972
prerequisite (4) A 0.750 0.656 0.500 0.867 0.500
B 0.750 0.656 1.000 0.867 0.500
C0 0.625 0.438 1.000
C 0.750 0.455 1.000 0.750 0.750
C-rel 0.500 0.607 1.000 0.000 0.000
exception (3) A 1.000 0.595 0.333 1.000 1.000
B 1.000 0.579 1.000 1.000 1.000
C0 1.000 0.333 1.000
C 1.000 0.467 1.000 1.000 1.000
C-rel 1.000 0.556 1.000 0.944 0.889
exception, headline excl. q3c (2) A 1.000 0.607 0.500 1.000 1.000
B 1.000 0.583 1.000 1.000 1.000
C0 1.000 0.375 1.000
C 1.000 0.500 1.000 1.000 1.000
C-rel 1.000 0.583 1.000 0.917 0.833
supersession (8) A 1.000 0.702 0.125 1.000 0.958
B 1.000 0.709 1.000 1.000 0.917
C0 0.938 0.594 1.000
C 0.938 0.646 1.000 0.969 0.875
C-rel 0.917 0.723 1.000 0.950 0.800
version-current (11) A 0.909 0.692 0.000 0.952 0.788
B 0.909 0.679 1.000 0.952 0.758
C0 0.818 0.545 1.000
C 0.864 0.613 1.000 0.886 0.818
C-rel 0.820 0.687 1.000 0.821 0.714
version-historical (1) all 1.000 0.714–0.875 1.000 1.000 1.000
cross-collection (3) A 0.833 0.542 0.000 0.889 0.556
B 0.833 0.524 1.000 0.889 0.444
C0 0.667 0.417 1.000
C 0.833 0.440 1.000 0.917 0.667
C-rel 0.700 0.507 1.000 0.750 0.000
partial (2) A 0.750 0.670 0.000 0.900 0.333
B 0.750 0.643 1.000 0.900 0.167
C0 0.250 0.625 1.000
C 0.250 0.625 1.000 0.375 0.000
C-rel 0.250 0.714 1.000 0.375 0.000
unanswerable (3) all 0.917–0.944 1.000 1.000 1.000
noise-prone (3) A 1.000 0.821 1.000 1.000 1.000
B 1.000 0.813 1.000 1.000 1.000
C0 1.000 0.667 1.000
C 1.000 0.711 1.000 1.000 1.000
C-rel 1.000 0.786 1.000 1.000 1.000
null-control (6) A/B/C 1.000 0.725–0.821 1.000 1.000 1.000
C-rel 1.000 0.787 1.000 1.000 1.000

Reporting caveat for C-rel generation columns: where C-rel's package was identical to C's, only a retrieval cell exists, so its per-tag completeness and pass figures can rest on fewer questions than n suggests (e.g. its prerequisite generation figures come from q5c alone).


5. Qualitative findings

5.1 The filter fixed every applicability failure and cost nothing

A retrieved a forbidden obsolete unit on 11 of 24 questions (q2a, q3a, q3c, q4a, q4c, q5a, q5b, q5c, q6a, q6b, q6c) — applicability 0.000 across the entire version-current tag. B's applicability filter removed every one, and B's per-question requirement recall is identical to A's on all 24 questions: honest applies_to metadata cost zero recall at topK 7.

An honest wrinkle: A's outcome pass barely suffered (0.903 vs B's 0.889) — with both versions in context, the generator usually picked the current one. The obsolete-context hazard is real but mostly latent at this corpus size; applicability accuracy measures the exposure, not the realized damage.

5.2 q2a — the graph's one clean win, fully attributable

"Our order needs to be rushed — what do we have to do to set that up?" requires request-rush-production and its prerequisite approve-artwork-proof (a different collection, semantically distant from rush vocabulary). In A and B, approve-artwork-proof ranked below the budget cut (dropped: rank cut, both conditions) → recall 0.50, pass 0/3. C followed prerequisite_of via request-rush-production and packaged it → recall 1.00, pass 3/3 — the only condition to answer q2a correctly. (C-rel's package was identical to C's here.)

C0 (same seeds, no edges) scored 0.50, and q2a is the entire population C−C0 recall gap (0.929 vs 0.905). Attribution doesn't get cleaner: the requires edge did exactly what H1 hypothesized — once.

5.3 q5b/q5c — seed starvation is the price of expansion headroom

S = 4 was fixed in the note to leave expansion headroom under B = 1400. On q5c (broken mugs — what do we do, and does the replacement incur the setup fee again?) the required units are file-damage-claim and inspect-delivery, but the question's fee-and-mug vocabulary spent all four seeds on setup-fee-policy-2026, setup-fee-reorder-exception, minimum-order-quantities, and drinkware-imprint-areas — and no typed edge leads from any of them to the damage-claim units. C recall 0.000. A and B, with topK 7, caught file-damage-claim at rank 7 → 0.50. (inspect-delivery made no condition's package — a miss shared by everyone.) q5b is the same shape: seeds went to color/imprint units, setup-fee-policy-2026 sat just below the S = 4 cut, and C's answers were missing the fee facts (stated pattern TTTf in all three reps).

The graph can only expand from seeds it holds. The partial tag (recall C 0.25 vs A/B 0.75) is where that bill landed. Per §5 control 5, S and B were fixed in the note and stand; any re-run with different seeding is exploratory.

5.4 C-rel — related edges gave back everything the typed closures earned

Wherever C-rel's package differed from C's (17 of 24 questions), the difference was related_to units (q5b gained vector-artwork-basics, volume-discount-policy, place-first-order; q5c gained volume-discount-policy, place-first-order, product-material-safety). The result: population irrelevance 0.737 — back at A's 0.741, erasing the precision advantage C's typed traversal had built (0.647). It never bought a recall point anywhere (0.908 ≤ C's 0.929), and on the one prerequisite question it generated fresh (q5c) it passed 0/3. N2 wasn't just confirmed; it was measured: the cost of untyped expansion begins with the first hop.

5.5 The fee fabrication — a generation behavior no retrieval condition caused or cured

q5c's unanswerable aspect (does a replacement for carrier-damaged goods incur another setup fee?) is stated nowhere in the corpus. It was asserted in 11 of 12 generated cells — A 3/3, B 3/3, C 3/3, C-rel 2/3. The mechanism is visible in the transcripts (B rep 3): the model chains the reorder stored-artwork exemption into "if the reprint uses your original, unchanged stored artwork … no setup fee would apply" — a plausible policy inference the handbook never makes. Identical rate with and without the filter, with and without the graph: retrieval architecture neither caused nor cured it.

The generic metrics did not reliably catch it. RAGAS on the asserted cells ran 0.70–0.91 (a mild dip from the ≈ 0.95 population mean); the eval judge flagged missing_information for B/C/C-rel — but scored A's three asserted cells 0.95–0.97 with evidenceSufficient: true and no primary issue. All 11 were flagged only under the registered aspect-specific check (the asserted boolean — itself judged by an LLM, but against an explicit per-aspect rule). Contrast q5b's aspect (the dollar amount): never asserted, but B failed to acknowledge the gap in 2 of 3 reps. Two distinct partial-answer failure modes — bluffing over the gap versus silence about it — and only the first is a fabrication event.

This is the strongest public write-up candidate in the run: an evaluation-design finding squarely inside the sandbox mission ("make the behavior observable"), and it required the schema-v2 aspect booleans to see.


6. Registered conclusion

The current structure-aware strategy does not beat the filtered baseline overall, even though explicit relationships demonstrably recover evidence that similarity and filtering alone can miss. On this corpus, at this budget, honest applicability metadata plus a version filter captured essentially all of the retrieval-quality gains (H3's exact prediction); the typed graph's one clean, provenance-attributable win (q2a) was offset by the seed starvation its own budget economics induced (q5b/q5c). The §12.8 decision is stop the graph, keep the filter, scoped to this corpus, budget, and expansion policy. Abstention discipline was perfect everywhere; the null-control was flat (no rigging smell); and the run surfaced a retrieval-independent fabrication pattern on partially answerable questions that generic evaluation metrics did not reliably flag — RAGAS dipped on those cells and the judge caught some conditions, but only the aspect-specific check identified every fabrication event. The negative result is the deliverable, as §2 committed.


7. Not registered: follow-up candidates

Nothing below was run, and no code, gold, scoring, or interpretation rule was changed after the run. These are ideas for future, separately labeled work.

  1. H4 diagnosis labels (proposal). Extend the FailureLabel vocabulary with mechanically derivable, provenance-dependent labels: prerequisite_not_retrieved (a requires edge exists from a packaged unit to an absent required unit), seed_starvation (required unit neither seeded nor reachable by any typed edge from any seed), and superseded_context (forbidden version present in the package). Each is computable from run.json + the manifest alone; the first two are inexpressible without the graph — H4's claim in concrete form.
  2. Hybrid seeding (exploratory re-run candidate). Use B's filtered topK-7 ranking as the seed pool and let typed expansion compete with ranks 5–7 for the remaining budget — directly targets the q5c failure without giving up the q2a win. Any such run is exploratory by §5 control 5.
  3. Generation-side partial guardrail. The §5.5 fabrication is a prompting/judging problem, orthogonal to retrieval; an answer-time "state what the evidence does not establish" instruction could be A/B-tested against the aspect booleans.
  4. Prose-degraded arm. Already named a follow-up in §12.10: condition A here is Markdown's well-migrated case; the "prose survives, semantics don't" story needs its own arm.

Source: experiments/findings-structure-aware-retrieval.md, committed with the registered run artifacts. This page renders that document without modification.