Appendix: Implementation Evidence Map

This appendix was withheld when this site was first published on 2026-08-12, and it is published now because it was re-graded rather than repaired. Four capabilities were graded Implemented/Tested on classes that a deliberate refactor removed on 2026-08-10 — two days before launch — and one citation named a file that has never existed in any repository. The gradings below record what survived that refactor, what did not, and what was never there. Publishing the correction rather than the original is the point of keeping the record.

Version: 0.4 (2026-08-12)

This appendix maps the five-thesis suite's major claims to current Consullo repository evidence. It is not a proof of capability. It is a review aid that distinguishes implemented code, tested support, design-only specification, proposed extensions, and evidence gaps.

Verification note (2026-08-12): every cited repository path and named test in this appendix was resolved against live files across the whole workspace, against full git history, and against the agentDescriptions registry. Resolution establishes only that a named artifact exists; it does not establish that the artifact supports the sentence citing it, and it does not convert component evidence into end-to-end capability evidence.

This replaces a note claiming the paths were "re-checked against the current repository on April 24, 2026" and that the check "confirms the cited component evidence." That claim had two failures worth recording rather than deleting. It went stale on 2026-08-10, when 199ea7fd removed the Bootstrapper subsystem and four cited files with it. And it was never true of LLMNonInteractiveSession.java, which has zero commits in the history of every repository searched and no registry entry — a path that never existed passed a check that claimed to confirm it. Both citations have been corrected below.

Evidence Status

Rows may use compound tags such as Implemented/Tested when implementation code exists and repository tests exercise at least part of that behavior. Compound tags do not imply full end-to-end deployment or benchmark validation.

Capability Status remains governed by 00-vocabulary-and-invariants.md. Evidence Status is narrower: it records what the Consullo codebase currently shows.

Master Claim

Claim:

Consullo Seed AI is a specified/proposed scaffold for governed recursive capability amplification, not a system that has reached ASI.

Evidence:

Evidence Status: Documented.

Gap:

No single repository-level benchmark report yet ties the five-thesis scaffold to observed recursive capability improvement.

Thesis 1: Validated Improvement Loop

Claim:

Candidate changes can be proposed, evaluated, validated, staged, monitored, rolled back, and recorded into method memory under explicit evidence gates.

Repository evidence:

Tests:

The five Bootstrapper*Test classes previously listed here were removed with the subsystem on 2026-08-10. No test currently exercises dependency validation, repair orchestration, or A2A method-call testing.

Evidence Status: Implemented/Tested for agent integration, compilation validation, and build reporting. Documented/Proposed for dependency validation, repair orchestration, and A2A method-call testing — all three implemented and tested until 2026-08-10, removed with the Bootstrapper subsystem, and owed to the agent-integration workstream. Documented/Proposed for full recursive improvement semantics, DGM-style population update, evaluator/validator separation across all agent types, and method-memory evolution.

Gaps:

Thesis 2: Cognitive Substrate

Claim:

Consullo can organize memory, knowledge access, reasoning, attention, metacognition, perception, creativity, and executive control into measurable cognitive workflows.

Repository evidence:

Tests:

Evidence Status: Documented for named cognitive agents and architecture. Simulated/Tested for cognitive-artifact composition, routing, validation allocation, trust transition, model-improvement, and portfolio dynamics.

Gaps:

Thesis 3: Causal-Decision Foundations

Claim:

Recursive improvement should use causal models, counterfactuals, robust intervention choice, prediction calibration, Goodhart checks, experiment portfolios, abstention, and escalation.

Repository evidence:

Tests:

Evidence Status: Documented/Proposed for causal-decision agents and robust causal intervention semantics. Simulated/Tested for selected business and AI-native commerce scenario modeling.

Gaps:

Thesis 4: Self-Modifying Software Substrate

Claim:

Consullo can constrain code generation, repair, testing, semantic validation, provenance, staged deployment, and permission through an executable software substrate.

Repository evidence:

Tests:

Evidence Status: Implemented/Tested for parts of compilation validation, A2A dispatch, agent utilities, and reporting. Documented/Proposed for dependency validation and repair-workflow coordination, which lost their implementation on 2026-08-10. Documented/Proposed for secure staged deployment, semantic invariants, validator non-regression, provenance proofs, and full repair-pipeline recurrence.

Gaps:

Thesis 5: Alignment Invariants And Scoped Trust

Claim:

Recursive modification must be constrained by alignment invariants, AAF dissent, scoped trust, containment, interruptability, evidence preservation, ADO reporting, and human authority.

Repository evidence:

Tests:

Evidence Status: Implemented/Tested for A2A ingress auth/policy gates, dispatch constraints, persona-template support, and selected trust/validation simulations. Documented/Proposed for AAF orchestration, the Friendship agent authority, ADO reporting, scoped trust estimates, containment, interruptability, and recursive alignment improvement.

Gaps:

Substrate Context

Claim:

Specialized LLM routing, rapid knowledge access, atomic prompts, internal economy, and A2A/server infrastructure support the five theses without proving capability or alignment.

Repository evidence:

The line previously here cited llm/hooks/LLMNonInteractiveSession.java as supporting session memory and hook behavior. It is deleted rather than repointed: that file has zero commits in the history of every repository searched and no registry entry, so it was never evidence for anything. The grading below rests on the remaining infrastructure and does not need a replacement citation.

Tests:

Evidence Status: Implemented/Tested for selected infrastructure. Documented/Proposed for specialized LLM ecosystem routing and rapid knowledge access as integrated Seed AI substrate.

Gaps:

Organizational Recursive Self-Improvement

Claim:

Consullo can be interpreted as an AI-native R&D organization whose agents, workflows, benchmarks, ledgers, method memories, validators, and governance gates produce validated improvement of research, engineering, evaluation, memory, and governance processes.

Evidence Status: Documented/Proposed.

Current evidence is concentrated in software-substrate utilities, benchmark/test-plan appendices, the canonical evidence-ledger schema, cognitive-artifact simulations, and the design-level organizational appendix. No implemented weekly organizational RSI loop, frozen V_ref_0, pre-registration ledger, external-evaluator sampling pipeline, kill-switch drill, or portfolio dashboard has been identified in this repository.

Gaps:

Highest-Priority Evidence Gaps Before Publication

Snapshot date: April 24, 2026.

  1. Implement or explicitly defer the three load-bearing Thesis 5 owner contracts specified in appendix-thesis-5-operational-contracts.md: Friendship, AdversarialAlignmentOrchestrator, and AbundanceDistributionMonitor. See risks-and-criticisms.md: Owner As Final AAF Adjudicator, AAF Decorative Failure, and Abundance Obligation Vagueness.
  2. Implement the canonical evidence-ledger schema from appendix-evidence-ledger-schema.md for improvement evidence, trust evidence, provenance, dissent, incidents, overrides, and rollback state. See risks-and-criticisms.md: Evidence-Ledger Schema Unimplemented Or Unenforced and Self-Modifying Provenance Graph Integrity.
  3. Produce an end-to-end improvement-loop demonstration with baseline, proposed change, evaluator evidence, validator results, staged deployment or simulation, cost/benefit report, and rollback semantics. Benchmark design is specified in appendix-thesis-1-improvement-loop-benchmarks.md; implementation evidence remains pending. See risks-and-criticisms.md: Recursive-Improvement Claim Without End-To-End Evidence.
  4. Bind Model 2's benchmark-family measurement conventions to at least one external or project-local benchmark report with declared baselines, units, and cost normalization. Benchmark design is specified in appendix-thesis-2-cognitive-workflow-benchmarks.md; implementation evidence remains pending. See risks-and-criticisms.md: Cognitive Architecture Inflation.
  5. Add causal-decision implementation evidence: causal graph representation, counterfactual procedure, calibration battery, and Goodhart checker. Benchmark design is specified in appendix-thesis-3-causal-decision-benchmarks.md; implementation evidence remains pending. See risks-and-criticisms.md: Causal Model Overreach and Goodhart And Validator Gaming.
  6. Add validator non-regression suites for Thesis 4 and operationalize the ValidatorStrength convention from appendix-formal-models.md. Benchmark design is specified in appendix-thesis-4-software-substrate-benchmarks.md; implementation evidence remains pending. See risks-and-criticisms.md: Software Repair Overclaim and Goodhart And Validator Gaming.
  7. Add AAF dissent aggregation and ADO reporting implementation evidence before treating those roles as more than specified/proposed. Benchmark design is specified in appendix-thesis-5-alignment-benchmarks.md; implementation evidence remains pending. See risks-and-criticisms.md: AAF Decorative Failure, Abundance Obligation Vagueness, and Cost Of AAF Risk.