Executed research
KnowledgeGuard / EGB
typed evidence-deficiency diagnosis and repair for RAG
Controlled study — analysis complete
The question
Can a RAG system identify how its retrieved evidence is deficient — missing, insufficient, conflicting, outdated, or absent from the corpus — and does that diagnosis carry actionable information for selecting a repair action?
- Method
- Fully within-record 5 x 6 factorial: every source record is instantiated under every deficiency type and run under every repair action, including cells no router would choose. Deficiency type is oracle by design, so the factorial measures whether type carries information independently of whether it can be detected. Methods were frozen and the analysis script committed before the first result row existed.
- Data
- EGB, built on the HoH corpus (18,807 indexed passages) — the only available source with real superseded values, which OUTDATED requires. Four of six construction operators are purely subtractive; nothing is fabricated. HotpotQA is held as a separate replication factorial and is never pooled.
Result
Type and action interact strongly (partial eta-squared 0.32, permutation p = 1e-4). Oracle routing beats the best type-agnostic policy by 6.6 F1 points, 95% CI [2.6, 10.5]. But with a real detector the benefit reverses: predicted routing scores 0.064 F1 below type-agnostic. The headroom is real and, on this evidence, unreachable.
The 5 x 6 factorial
5 × 6, fully crossed. Every cell was run. Values are token F1 against the gold answer. Select a row, column or cell to see what it represents.
| Evidence-deficiency type ↓ / Repair action → | ||||||
|---|---|---|---|---|---|---|
// select a row, column or cell
Run 2026-09-15. 47 records x 5 types x 6 actions = 1,410 cells, every cell n = 47. 2,946 generator calls, no API spend. Floor check passed: SUFFICIENT x NONE = 0.831, so the reader can use good evidence and the cells are interpretable.
Measured values, released under the project's own Tier P policy, which clears per-cell scores with passage text and prompts removed. Absolute numbers reflect a 250M-parameter local reader and are not comparable with published RAG systems; the factorial is a within-instance contrast.
What the results support
Each number is paired with what it does and does not license.
Type and action interact
partial eta-squared 0.32, permutation p = 1e-4The action profile genuinely differs by deficiency type, corroborated by a mixed-model likelihood-ratio test. This says the cells differ; it does not by itself say that knowing the type is worth anything.
Oracle typing beats the best single action
+6.6 F1 points, 95% CI [2.6, 10.5]The pre-registered null is rejected, but by a margin whose lower bound sits exactly at the frozen practical threshold of three points rather than comfortably above it. The honest statement is that typing buys roughly six points and the data are consistent with as little as three.
The number that must not be quoted
+41.5 F1 points against fixed escalationAgainst a fixed-escalation policy typing looks enormous, but almost all of that is the action main effect: escalation is simply a poor universal policy for a small reader because it dilutes the context. Reporting this as the routing benefit would be the single easiest way to overstate the result.
With a real detector the benefit reverses
predicted routing 0.064 F1 below type-agnostic [-0.120, -0.012]This is the result that matters for anyone wanting to build on it. The headroom is real and, on this evidence, unreachable: routing on a diagnosed type is worse than just picking one good action and applying it everywhere.
Constructed conflict is far easier to detect than natural conflict
recall 0.957 against 0.574A pre-registered threat to validity, now measured rather than feared. Nearly a third of natural conflicts are called sufficient — the dangerous error, because the system then answers from evidence it has not noticed contradicts itself.
Knowing when to decline is the largest single effect
+0.96 selective utility for ABSENT to ABSTAINInvisible under answer correctness, where an abstention and a confident fabrication both score zero. It appears only because a second, abstention-sensitive dependent variable was pre-registered before the run.
Correction
The study's only cell surviving multiple-comparison correction was read as arbitration resolving conflict by corroboration. A forensic re-analysis of the frozen artifacts — no model loaded, nothing modified — found otherwise. ARBITRATE scores identically to four decimal places in the conflicting, outdated and sufficient cells, and produces the same answer string as the sufficient cell on 45 of 47 records. The injected counter-passages are the only passages absent from the index that arbitration corroborates against: 47 of 47, against 0 of 1,315 gold and distractor passages. The counter-passage also quotes the whole question, giving it a query-token overlap of 1.000 against gold's 0.676, making it the strict maximum-overlap passage in 47 of 47 delivered sets. Two trivial rules — pick the maximum-overlap passage, or the one absent from the index — identify the counter-side in 47 of 47 instances with no generation at all.
The number stands; the causal reading does not. The correction establishes that the artifact exists. Whether it explains the effect is what the pre-registered replication measures, and that replication has not been run.
Evidence, item by item
Ran and produced evidence6
Type and action interact
The action profile genuinely differs by deficiency type, corroborated by a mixed-model likelihood-ratio test. This says the cells differ; it does not by itself say that knowing the type is worth anything.
Oracle typing beats the best single action
The pre-registered null is rejected, but by a margin whose lower bound sits exactly at the frozen practical threshold of three points rather than comfortably above it. The honest statement is that typing buys roughly six points and the data are consistent with as little as three.
The number that must not be quoted
Against a fixed-escalation policy typing looks enormous, but almost all of that is the action main effect: escalation is simply a poor universal policy for a small reader because it dilutes the context. Reporting this as the routing benefit would be the single easiest way to overstate the result.
With a real detector the benefit reverses
This is the result that matters for anyone wanting to build on it. The headroom is real and, on this evidence, unreachable: routing on a diagnosed type is worse than just picking one good action and applying it everywhere.
Constructed conflict is far easier to detect than natural conflict
A pre-registered threat to validity, now measured rather than feared. Nearly a third of natural conflicts are called sufficient — the dangerous error, because the system then answers from evidence it has not noticed contradicts itself.
Knowing when to decline is the largest single effect
Invisible under answer correctness, where an abstention and a confident fabrication both score zero. It appears only because a second, abstention-sensitive dependent variable was pre-registered before the run.
Not built, or not run2
E6
E6 — the pre-registered replication that tests whether the construction artifact explains the CONFLICTING x ARBITRATE effect. Not yet run.
Completing the HotpotQA replication factorial.
Completing the HotpotQA replication factorial.
What this does not establish
- One corpus, one language, and a 250M-parameter local reader, so absolute numbers are not comparable with published RAG systems; the factorial is a within-instance contrast.
- Every conflict result is scoped to constructed, resolvable conflict. Natural, unresolvable conflict is in the benchmark for detection only.
- Constructed conflict is far easier to detect than natural conflict — recall 0.957 against 0.574 — so detection results do not generalise from the constructed set.
- The benchmark fails its own surface-leakage gate on detection for some types, and the failure is reported rather than repaired.
- The only cell surviving multiple-comparison correction was later found to be measuring a construction artifact rather than conflict resolution; the pre-registered replication that would settle it has not been run.
- The HotpotQA replication factorial has not been completed.
// Private repository