Pre-alphaSecurityApplied AI

CortexWard

An AI software-security engineer that understands, verifies, fixes, and secures code. Runs a multi-scanner and agent pipeline with a closed verification loop rather than reporting unverified findings.

PythonDockerStatic AnalysisAgents

The problem

A static analyser can tell you a dangerous pattern appears in your code. It cannot tell you whether that line is reachable, whether attacker-controlled data gets to it, or whether the finding is exploitable at all — so teams drown in findings they cannot triage. An LLM can reason about the code but cannot prove anything about it.

Context and constraints

Built as a security tool that has to earn its own conclusions. The design premise is that a finding is only as strong as the evidence attached to it, and that a language model — however fluent — is not evidence.

  • A model can be persuasive and wrong, so model judgement cannot be allowed to establish that a finding is real.
  • Running a proof-of-concept means executing potentially hostile code, which has to be isolated from the host.
  • Output has to be consumable by tools that already exist, which means standard formats rather than a bespoke report shape.
  • Every capability has to be honest about its own maturity, because a security tool that overstates its confidence is worse than one that says nothing.

What I built

A Python monorepo of independently versioned packages behind a hexagonal architecture. A Code Property Graph engine builds AST, control-flow, data-flow and call graphs over tree-sitter and answers reachability, taint and slice queries. Four scanner adapters feed a correlation layer. A seven-agent pipeline grounds its reasoning in the graph, and results are exported as SARIF, CycloneDX-VEX and a native JSON format.

Architecture

Ports and adapters throughout: the domain core is pure with no I/O, every external capability is a protocol-typed port, and adapters are discovered as plugins so a new scanner or model provider needs no core change.

invokedetectqueryreasonfindingsevidencehypotheseseventsverdictsward CLIscan, baseline, benchOrchestratorSeven-agent pipelineCode Property GraphAST, CFG, DFG, callsScanner adaptersFour, plus correlationLLM adaptersBounded, cannot verifyDomain corePure, no I/OEvent-sourced logSQLite storage portReportersSARIF, VEX, JSON
Architecture as documented in the public repository. Implemented components only — the packages listed here all exist and ship; see the ladder below for which verification capabilities are built.
ward CLI
scan, baseline, bench
Orchestrator
Seven-agent pipeline
Code Property Graph
AST, CFG, DFG, calls
Scanner adapters
Four, plus correlation
LLM adapters
Bounded, cannot verify
Domain core
Pure, no I/O
Event-sourced log
SQLite storage port
Reporters
SARIF, VEX, JSON

My contribution

Sole author and architect. Domain model, Code Property Graph engine, scanner adapters, agent pipeline, reporters and the evaluation harness.

The Verification Ladder

Rather than a binary "did an exploit run", each finding is assigned the strongest evidence that can actually be produced for it, and confidence is calibrated to that rung. Two rungs are built today and two are not — the distinction is the point, so it is stated on every rung rather than summarised.

Structural rules

  • A language model can never climb the ladder on its own. Model judgement is bounded by design and only concrete analysis produces rungs — this is enforced structurally rather than by convention.
  • Refutation is first-class. Evidence that a finding is not exploitable is captured and drives it toward a not-affected verdict, instead of being discarded as a non-result.

Rung status is taken from the repository's per-phase roadmap, which documents evidence for each line, rather than from the summary at the top of its README.

Technical decisions

What was chosen, why, and what it cost.

A model is never allowed to be the evidence

Decision
Language models participate in the pipeline as a source of hypotheses, but model judgement cannot raise a finding's verification rung. Only concrete analysis — graph reachability, taint tracing, sandboxed execution — can.
Why
A fluent, confident and wrong explanation is the characteristic failure of model-assisted security tooling. Making the ladder structurally unreachable by a model means that failure cannot silently become a verdict.
Trade-off
The tool is far more conservative than an LLM reviewer and will leave findings at a low rung that a model would happily call exploitable.

Hexagonal architecture with plugin-discovered adapters

Decision
The domain core is pure with no I/O, every external capability sits behind a protocol-typed port, and adapters register through entry-point discovery so adding one requires no core change.
Why
A security tool is mostly integrations — scanners, model providers, version control, sandboxes, report formats. Keeping them at the edge means the verification logic can be tested without any of them.
Trade-off
Considerably more indirection than a direct implementation, and a contributor has to understand the port catalogue before adding a capability.

Build a Code Property Graph rather than pattern-matching harder

Decision
A graph engine over tree-sitter unifies AST, control flow, data flow and the call graph, and answers reachability, taint and slice queries that the rest of the system treats as evidence.
Why
Reachability and taint are the two questions that separate a real finding from a match, and neither can be answered by a better regular expression.
Trade-off
A graph has to be built per language, so language coverage is bounded by the tree-sitter grammars and query sets actually written.

Emit SARIF and CycloneDX-VEX rather than a bespoke report

Decision
Findings export as SARIF 2.1.0, exploitability as CycloneDX-VEX, with a native JSON format alongside.
Why
A verdict that cannot be ingested by the tooling a team already runs does not change anything. Standard formats make the output actionable without adoption.
Trade-off
The mapping onto CycloneDX's own state enumeration is documented as one-directional rather than a lossless round-trip, so some internal nuance is lost on export.

Security and reliability

The failure modes that shaped the implementation.

Executing untrusted code

Dynamic verification means running code from the project under test. The sandbox adapter isolates that in Docker — built and tested, though not yet wired into the pipeline.

The model boundary is a trust boundary

Model output enters the system as a hypothesis, never as a fact. The ladder is the mechanism that enforces it.

Event-sourced finding log

Findings are stored as an append-only event log, so how a verdict was reached remains reconstructable rather than being overwritten by its latest state.

Verification

How the implementation was checked, and how much of that can be shown publicly.

  • Quality gate reproducible locallyevidenced

    The same gate CI runs is a single make target, alongside pre-commit hooks and a dev container.

  • Evaluation harness with a statistical protocolevidenced

    Detection metrics, a versioned golden dataset, run manifests, and a protocol specifying bootstrap confidence intervals and McNemar's test.

  • Verification and patch-quality metrics are not populatednot publicly evidenced

    Those metrics need evidence at rungs the pipeline does not yet produce, so the fields exist in the manifest and are deliberately left empty rather than estimated.

  • VEX export verified end to endevidenced

    The CycloneDX-VEX reporter is fully covered and was verified by a real scan producing a valid document.

Results

  • Phases 0 through 4 are complete: domain core, workspace and port contracts, the Code Property Graph engine, four scanners with SARIF output, and the seven-agent framework with multi-provider model support.
  • Six further phases are partially built, each with specific documented items still open, and the v1.0 phase has not started.
  • The verification pipeline reaches the taint rung today. The two rungs that would demonstrate rather than infer exploitability are not built.

Timeline

  1. Code Property Graph engine

    AST, control flow, data flow and call graph over tree-sitter.

  2. Scanners and SARIF

    Four scanner adapters with cross-tool correlation.

  3. Agent framework

    Seven agents, multi-provider models, graph-grounded reachability evidence.

Limitations and disclosure

What this project does not do, and what cannot be shown publicly.

  • Pre-alpha. Six of eleven phases are partially built and the v1.0 phase has not started.
  • The verification pipeline reaches the taint rung. Dynamic proof-of-concept and differential-test verification are not built: the sandbox adapter exists but nothing in the agent pipeline calls it, and no component produces proof-of-concept evidence yet.
  • Verification and patch-quality metrics are left unpopulated because they need evidence at rungs the pipeline does not yet produce.
  • A distinct false-positive-reduction capability, contamination-controlled evaluation splits and broader benchmark datasets remain unbuilt.
  • The summary at the top of the repository README describes the verification loop as closed to the dynamic rung; the per-phase roadmap records that wiring as not yet built. This page follows the roadmap.

The repository is public, so every claim here can be checked against it. Where the README summary and the per-phase roadmap disagree about what is built, this page follows the roadmap, which documents evidence for each line and is the more conservative of the two.

Other projects in the same engineering domains.