Research, separated from the plans that look like it

One study here ran, produced results, and then produced a correction against its own headline finding. The rest are at earlier stages, and this page says which is which before it says anything else — because a research page where a completed experiment and a proposal read the same is not reporting research, it is advertising.

How this work is categorised

  1. Executed research1

    A question was posed, an experiment was designed and frozen, it ran, and it produced results — including results that went against the hypothesis.

    // Did it run, and what did the numbers turn out to mean?

  2. Research-driven engineering1

    Software built to answer a research question. The infrastructure is real; the evaluation it exists to support has not produced results yet.

    // What is built, and what has it actually measured?

  3. Systems experiments1

    Protocol and systems work where the interesting question is what survives contact with real hardware rather than a simulator.

    // Which parts ran on a device, and which only in simulation?

  4. Current final-year project1

    University research at the stage it has actually reached. A proposal is not an implementation and is not presented as one.

    // What exists today, as opposed to what is proposed?

What exists, project by project

Each cell counts documented items — a measured finding, a protocol step, a verification rung, a pre-registered experiment — grouped by the strongest evidence that exists for them. Select a column to see what a state means, or a cell to see the items in it.

Research entries by evidence state. Counts are documented items, not a score, and are not comparable between rows.
Research ↓ / Evidence →
KnowledgeGuard / EGB
CortexWard
Emergency Mesh
SCAR-OS

// select a column or a cell

These are counts, not a score, and they are not comparable between rows: one project documents measured findings, another documents protocol steps, another verification rungs. A high number in one column says how much of that kind of item was recorded, not how good the work is.

13
Ran and produced evidence
3
Built, not yet evaluated
0
Simulated only
6
Not built, or not run

// documented items across all four entries — not a score, and not comparable between projects

Executed research

KnowledgeGuard / EGB

typed evidence-deficiency diagnosis and repair for RAG

Controlled study — analysis complete

The question

Can a RAG system identify how its retrieved evidence is deficient — missing, insufficient, conflicting, outdated, or absent from the corpus — and does that diagnosis carry actionable information for selecting a repair action?

Method
Fully within-record 5 x 6 factorial: every source record is instantiated under every deficiency type and run under every repair action, including cells no router would choose. Deficiency type is oracle by design, so the factorial measures whether type carries information independently of whether it can be detected. Methods were frozen and the analysis script committed before the first result row existed.
Data
EGB, built on the HoH corpus (18,807 indexed passages) — the only available source with real superseded values, which OUTDATED requires. Four of six construction operators are purely subtractive; nothing is fabricated. HotpotQA is held as a separate replication factorial and is never pooled.

Result

Type and action interact strongly (partial eta-squared 0.32, permutation p = 1e-4). Oracle routing beats the best type-agnostic policy by 6.6 F1 points, 95% CI [2.6, 10.5]. But with a real detector the benefit reverses: predicted routing scores 0.064 F1 below type-agnostic. The headroom is real and, on this evidence, unreachable.

Factorial executed1,410 balanced cellsAnalysis pre-registeredSelf-published correction

The 5 x 6 factorial

5 × 6, fully crossed. Every cell was run. Values are token F1 against the gold answer. Select a row, column or cell to see what it represents.

The 5 x 6 factorial. Evidence-deficiency type by Repair action, measured in token F1 against the gold answer.
Evidence-deficiency type ↓ / Repair action →

// select a row, column or cell

Run 2026-09-15. 47 records x 5 types x 6 actions = 1,410 cells, every cell n = 47. 2,946 generator calls, no API spend. Floor check passed: SUFFICIENT x NONE = 0.831, so the reader can use good evidence and the cells are interpretable.

Measured values, released under the project's own Tier P policy, which clears per-cell scores with passage text and prompts removed. Absolute numbers reflect a 250M-parameter local reader and are not comparable with published RAG systems; the factorial is a within-instance contrast.

What the results support

Each number is paired with what it does and does not license.

Type and action interact

partial eta-squared 0.32, permutation p = 1e-4

The action profile genuinely differs by deficiency type, corroborated by a mixed-model likelihood-ratio test. This says the cells differ; it does not by itself say that knowing the type is worth anything.

Oracle typing beats the best single action

+6.6 F1 points, 95% CI [2.6, 10.5]

The pre-registered null is rejected, but by a margin whose lower bound sits exactly at the frozen practical threshold of three points rather than comfortably above it. The honest statement is that typing buys roughly six points and the data are consistent with as little as three.

The number that must not be quoted

+41.5 F1 points against fixed escalation

Against a fixed-escalation policy typing looks enormous, but almost all of that is the action main effect: escalation is simply a poor universal policy for a small reader because it dilutes the context. Reporting this as the routing benefit would be the single easiest way to overstate the result.

With a real detector the benefit reverses

predicted routing 0.064 F1 below type-agnostic [-0.120, -0.012]

This is the result that matters for anyone wanting to build on it. The headroom is real and, on this evidence, unreachable: routing on a diagnosed type is worse than just picking one good action and applying it everywhere.

Constructed conflict is far easier to detect than natural conflict

recall 0.957 against 0.574

A pre-registered threat to validity, now measured rather than feared. Nearly a third of natural conflicts are called sufficient — the dangerous error, because the system then answers from evidence it has not noticed contradicts itself.

Knowing when to decline is the largest single effect

+0.96 selective utility for ABSENT to ABSTAIN

Invisible under answer correctness, where an abstention and a confident fabrication both score zero. It appears only because a second, abstention-sensitive dependent variable was pre-registered before the run.

Correction

What CONFLICTING x ARBITRATE actually measures

The study's only cell surviving multiple-comparison correction was read as arbitration resolving conflict by corroboration. A forensic re-analysis of the frozen artifacts — no model loaded, nothing modified — found otherwise. ARBITRATE scores identically to four decimal places in the conflicting, outdated and sufficient cells, and produces the same answer string as the sufficient cell on 45 of 47 records. The injected counter-passages are the only passages absent from the index that arbitration corroborates against: 47 of 47, against 0 of 1,315 gold and distractor passages. The counter-passage also quotes the whole question, giving it a query-token overlap of 1.000 against gold's 0.676, making it the strict maximum-overlap passage in 47 of 47 delivered sets. Two trivial rules — pick the maximum-overlap passage, or the one absent from the index — identify the counter-side in 47 of 47 instances with no generation at all.

The number stands; the causal reading does not. The correction establishes that the artifact exists. Whether it explains the effect is what the pre-registered replication measures, and that replication has not been run.

Evidence, item by item

Ran and produced evidence6

  • Type and action interact

    The action profile genuinely differs by deficiency type, corroborated by a mixed-model likelihood-ratio test. This says the cells differ; it does not by itself say that knowing the type is worth anything.

  • Oracle typing beats the best single action

    The pre-registered null is rejected, but by a margin whose lower bound sits exactly at the frozen practical threshold of three points rather than comfortably above it. The honest statement is that typing buys roughly six points and the data are consistent with as little as three.

  • The number that must not be quoted

    Against a fixed-escalation policy typing looks enormous, but almost all of that is the action main effect: escalation is simply a poor universal policy for a small reader because it dilutes the context. Reporting this as the routing benefit would be the single easiest way to overstate the result.

  • With a real detector the benefit reverses

    This is the result that matters for anyone wanting to build on it. The headroom is real and, on this evidence, unreachable: routing on a diagnosed type is worse than just picking one good action and applying it everywhere.

  • Constructed conflict is far easier to detect than natural conflict

    A pre-registered threat to validity, now measured rather than feared. Nearly a third of natural conflicts are called sufficient — the dangerous error, because the system then answers from evidence it has not noticed contradicts itself.

  • Knowing when to decline is the largest single effect

    Invisible under answer correctness, where an abstention and a confident fabrication both score zero. It appears only because a second, abstention-sensitive dependent variable was pre-registered before the run.

Not built, or not run2

  • E6

    E6 — the pre-registered replication that tests whether the construction artifact explains the CONFLICTING x ARBITRATE effect. Not yet run.

  • Completing the HotpotQA replication factorial.

    Completing the HotpotQA replication factorial.

What this does not establish

  • One corpus, one language, and a 250M-parameter local reader, so absolute numbers are not comparable with published RAG systems; the factorial is a within-instance contrast.
  • Every conflict result is scoped to constructed, resolvable conflict. Natural, unresolvable conflict is in the benchmark for detection only.
  • Constructed conflict is far easier to detect than natural conflict — recall 0.957 against 0.574 — so detection results do not generalise from the constructed set.
  • The benchmark fails its own surface-leakage gate on detection for some types, and the failure is reported rather than repaired.
  • The only cell surviving multiple-comparison correction was later found to be measuring a construction artifact rather than conflict resolution; the pre-registered replication that would settle it has not been run.
  • The HotpotQA replication factorial has not been completed.
Read the full case study →

// Private repository

Research-driven engineering

CortexWard

closing the verification loop on automated security findings

Pre-alpha — verification loop running

The question

Can an agent pipeline confirm that a reported security finding is real, and that a proposed fix removes it, instead of emitting unverified findings?

Method
Multi-scanner pipeline with agent-driven triage and a sandboxed verification step.

Evidence, item by item

Built, not yet evaluated3

  • 0 · NONE

    The weakest rung. It says a scanner matched something, and nothing more. Findings that never climb above this are exactly the noise the project exists to reduce.

  • 1 · STATIC_REACHABILITY

    The sink is reachable in the call graph. This is the first rung that requires real analysis rather than a match, and it comes from the Code Property Graph rather than from a model.

  • 2 · TAINT_CONFIRMED

    Attacker-controlled data actually reaches the dangerous operation. This is the highest rung the pipeline currently reaches, and a finding corroborated to here is marked verified while its exploitability verdict stays under investigation — because reachable and tainted is not the same as demonstrated.

Not built, or not run2

  • 3 · DYNAMIC_POC

    The Docker sandbox adapter is built and tested, but nothing in the agent pipeline calls it yet, and no component produces proof-of-concept evidence for it to replay. Both are documented as unbuilt rather than in progress.

  • 4 · DIFFERENTIAL_TEST

    Blocked behind the same gap as rung 3, and needs a design for invoking an arbitrary target project's test suite generically across ecosystems.

Stated gaps

  • Verification and patch-quality metrics are not populated

    Those metrics need evidence at rungs the pipeline does not yet produce, so the fields exist in the manifest and are deliberately left empty rather than estimated.

What this does not establish

  • Pre-alpha. Coverage and evaluation are incomplete, and results should not be treated as benchmarked.

Systems experiments

Emergency Mesh

store-and-forward messaging over Bluetooth Low Energy

Active development — transport and protocol built and tested

The question

Can phones relay messages for each other over BLE reliably enough to be useful when there is no internet, no cell service and no server?

Method
Custom protocol with routing, encryption and persistence over a native BLE transport.

Evidence, item by item

Ran and produced evidence7

  • Discovery

    Devices advertise and scan over Bluetooth Low Energy, rank the connections they find and retry the ones that drop. Verified between two phones in range.

  • Local identity

    Each device generates and holds its own keys. There is no server to register with, so identity is created on first run and never leaves the device.

  • Compose, seal and sign

    Announcements, SOS broadcasts and direct messages are composed, sealed and signed before they reach the radio, and verified on arrival.

  • Fragment and reassemble

    The Bluetooth MTU is far smaller than a message, so frames are fragmented and reassembled natively beneath the protocol layer.

  • Single-hop transfer

    Two phones directly in range exchange real signed frames: contacts populate from genuine announcements and a one-to-one conversation sends and receives over live Bluetooth.

  • Acknowledgement

    A received message is stored and acknowledged on first arrival, and the sender marks it delivered. This exists for direct messages only — an SOS broadcast still cannot claim it reached anyone.

  • Store and forward

    A message queued while its destination is unreachable is persisted, retried as soon as a peer connects and otherwise on a bounded poll, and expired once it passes its lifetime.

Not built, or not run1

  • Multi-hop relay

    Carrying a message for a third device that the sender cannot reach directly. This is the defining promise of a mesh and it is not implemented — the transport is single-hop only, which is why the far device above never receives anything.

Stated gaps

  • The native stack has not been verified on a device

    The final transport task requires two physical phones and has not been completed. Everything below the application layer is therefore verified against its own tests and contracts, not against hardware.

  • Mesh behaviour beyond one hop exists only in simulation

    The simulator drives the same routing code as the real path, but a passing simulation is not evidence of hardware behaviour and is not presented as such.

What this does not establish

  • Protocol, routing, cryptography, persistence and the native BLE transport are built and tested; the user interface is at an early stage.
  • Real multi-hop relay across physical devices is not yet fully validated.
  • Not yet installable as a finished product, and not suitable for real emergencies.
Read the full case study →

// Private repository

Current final-year project

SCAR-OS

voice and intent-driven interaction in a developer operating environment

Current FYP — Research & Architecture Stage

The question

How should voice and intent-driven interaction be integrated into a developer-oriented operating environment?

Evidence, item by item

Not built, or not run1

  • Not implemented

    No implementation yet. The repository contains a README only.

What this does not establish

  • No implementation yet. The repository contains a README only.

// Private repository

How this page is kept honest

  • Nothing on this page is written by hand about what exists. Every item in the ledger is derived from a canonical record — a measured finding, a protocol step with its own status, a verification rung, a pre-registered experiment — so a capability cannot be promoted by editing a sentence.
  • A result and an implementation are different states and are never merged. Code that exists but has measured nothing reads as built, not as evidence.
  • A simulation is evidence about the simulator. Where behaviour has only been simulated, it is labelled that way and is never counted with behaviour verified on hardware.
  • KnowledgeGuard’s correction is published at the same weight as the result it corrects, and the experiments that have not been run are listed rather than omitted.
  • KnowledgeGuard’s repository is private. Published here are measured scores, counts and statistics only — the categories its own release policy clears, with passage text and rendered prompts removed. No benchmark passage, prompt, question or gold answer is reproduced.

The engineering these questions produced is on the work page, and the changes that went upstream are on open source.