---
title: "Appendix: Organizational Recursive Self-Improvement"
summary: "A long-horizon organizational interpretation of governed recursive improvement."
status: "speculative research target"
provenance: "derived from the owner-approved private Consullo design corpus; no artifact-specific public receipt has been issued"
claim_ids: ["CP-001", "CP-002"]
last_reviewed: "2026-08-12"
receipt: "none"
non_claims: ["No organizational recursive self-improvement capability is represented as operating.", "This research target has no public outcome evidence."]
---
# Appendix: Organizational Recursive Self-Improvement

Version: 0.1 (2026-04-24)

Capability Status: specified/proposed. Evidence Status: Documented/Proposed.

This appendix defines Consullo as an AI-native R&D organization at the architectural level. It is not an execution log and not evidence that Consullo already performs organizational-scale recursive self-improvement. Live operating controls, weekly cycles, and safety preconditions are specified separately in the internal execution plan.

The purpose is to prevent a narrow reading of recursive self-improvement as code self-editing. Java repair, compilation validation, generated-agent hygiene, and repository maintenance are valuable early exploitation lanes. They are not the whole Seed AI thesis. The broader claim is that Consullo can be specified as a governed organization of agents, workflows, ledgers, benchmarks, method memories, validators, and alignment gates whose product is validated improvement of its own research, engineering, evaluation, memory, and governance processes.

## Non-Claims

This appendix does not claim:

- Consullo already functions as a large AI R&D organization.
- Consullo has reached ASI, general superintelligence, or autonomous scientific discovery.
- brainstorming equals breakthrough discovery.
- agent count equals organizational intelligence.
- generated research artifacts equal scientific progress.
- benchmark design equals benchmark result.
- Java repair validates algorithmic self-improvement.
- organizational language solves alignment.
- completed cycles equal validated capability.
- process metrics equal outcome metrics.

## Definition

The **Consullo AI-Native R&D Organization** is the coordinated set of agents, workflows, ledgers, benchmarks, method memories, review gates, and governance roles that turn broad improvement objectives into agendas, candidate hypotheses, pre-registered experiments, implementations, evaluations, accepted improvements, rejected anti-patterns, and reusable organizational learning.

The organization is not a sixth thesis. It is a cross-thesis operating layer:

- Thesis 1 supplies the improvement loop, promotion semantics, evidence packages, protected-set non-regression, and method-memory update.
- Thesis 2 supplies cognitive functions such as brainstorming, metacognition, memory, executive routing, analogy, perspective, and negative-space search.
- Thesis 3 supplies experiment design, causal scope, calibration, portfolio reasoning, Goodhart controls, and abstention.
- Thesis 4 supplies prototype implementation, code repair, tests, semantic validation, validator non-regression, and deployment discipline.
- Thesis 5 supplies scoped trust, AAF dissent, owner override audit, incident learning, containment, and permission boundaries.

The organizational layer asks whether these thesis-local mechanisms compose into a disciplined research loop. It does not replace the five-thesis dependency topology.

## Design-To-Execution Cascade

The organizational layer uses a staged design-to-execution cascade:

```text
idea -> thesis -> execution plan -> planner JSON -> ledger evidence -> method memory / risk update
```

Each stage has a different acceptance standard. The idea layer can remain broad and aspirational. The thesis layer turns the idea into definitions, invariants, risks, evidence boundaries, and non-claims. The execution plan turns the thesis into operating doctrine: control artifacts, cadence, stop conditions, owner authority, and evidence obligations. Planner JSON turns the execution plan into temporally bounded work objects for StrategicPlanner, CampaignPlanner, OperationalPlanner, MissionPlanner, and TaskExecutor.

The layers must not collapse into each other. The thesis does not run the system. The execution plan does not create new thesis claims. Planner JSON does not invent governance policy. Ledger records, not generated plans or completed prose, establish what actually happened. Execution evidence may update method memory, risks, implementation-evidence boundaries, or future plans; it should update the thesis only when it changes the publication claim boundary.

## Why Java Repair Is Only The First Exploitation Lane

The strongest current repository evidence is concentrated in software-substrate work: Java repair, compilation checks, dependency validation, generated-agent repair, method insertion, and related code-generation constraints. That evidence matters because it gives Consullo a concrete early domain where proposals can be implemented, tested, rejected, repaired, and preserved.

But software maintenance can become an attractor. It gives cheap positive feedback: a patch compiles, a test passes, a method is cleaner, a dependency is fixed. Those wins are real, but they mostly exercise Thesis 4 and the local side of Thesis 1. They do not by themselves demonstrate algorithmic research improvement, evaluator hardening, portfolio learning, cross-lane method transfer, or novel cognitive strategy discovery.

The organizational layer therefore distinguishes:

- exploitation: improving known workflows, reducing cost, repairing code, strengthening validators, and reusing method memories
- exploration: searching for new algorithms, new workflows, new agent compositions, new benchmarks, new cognitive strategies, new evaluator designs, and new governance mechanisms

Seed AI progress requires both. A process that only exploits becomes a software-maintenance organization. A process that only explores becomes a brainstorm factory. Consullo needs portfolio governance over both.

The live execution plan binds this principle to an exploration budget floor, lane-diversity metric, rebalancing trigger, and stop condition.

There is also a positive reason to weight exploration toward tool-building, not only a defensive one. Empirical analysis of science's 750-plus major discoveries finds that the consistent driver of breakthroughs is the development of a new method or tool, and that one new tool tends to unlock many later discoveries across different fields (New tools drive scientific discovery, Humanities and Social Sciences Communications, doi:10.1057/s41599-026-06865-1, Krauss). The exploration lanes that build reusable tools — evaluators, benchmarks, method memories, cognitive strategies, planning and retrieval machinery — are therefore not diversity for its own sake; they are the historically highest-leverage form of research. The same evidence warns of a specific waste: a capability that exists but is never applied to an open problem is a missed discovery, which the execution plan instruments as a discovery-time-lag signal.

## Specification-Maturity Levels

These levels describe maturity of organizational specification and evidence. They are not ASI levels and not claims of demonstrated capability.

### Level 0: Documented Organization

Agent roles, workflows, and governance contracts are specified. No demonstrated organizational learning is claimed.

### Level 1: Software-Substrate Exploitation

Consullo improves local software artifacts: compilation repair, dependency repair, generated-agent hygiene, tests, and documentation consistency.

### Level 2: Method-Memory Compounding

Accepted and rejected changes alter future behavior through reusable method memories, anti-patterns, validator cases, routing rules, and cost models.

### Level 3: Workflow And Evaluation Improvement

Consullo improves how it decomposes tasks, routes agents, designs benchmarks, evaluates candidates, and controls false accepts or false rejects.

### Level 4: Research Portfolio Improvement

Consullo manages multiple exploration and exploitation bets, allocates attention and compute, runs comparative experiments, and improves research productivity over time.

### Level 5: Governed Organizational Self-Improvement

Consullo measurably improves the R&D organization itself while preserving evidence, alignment gates, scoped trust, rollback, incident learning, and human authority.

All levels above current repository evidence remain specified/proposed until benchmark reports and implementation evidence exist.

## R&D Organization Lifecycle

An organizational cycle should be read as a structured research operation, not as an unconstrained agent swarm.

1. Agenda formation
2. Hypothesis generation
3. Pre-registration
4. Portfolio selection
5. Experiment design
6. Prototype or implementation
7. Evaluation and held-out benchmarking
8. Adversarial review and external-evaluator sampling
9. Synthesis and decision
10. Method-memory / anti-pattern update
11. Governance gate
12. Post-cycle learning

Each stage should produce typed artifacts and ledger records. The publication suite specifies the structure; the execution plan specifies live cadence, thresholds, and stop conditions.

## Agent Role Families

The organizational layer groups agents by role family instead of treating every named agent as an independent capability claim.

| Role family | Contribution | Primary thesis |
| --- | --- | --- |
| Agenda and portfolio agents | Select objectives, allocate attention, rebalance exploration/exploitation lanes. | Thesis 1 / Thesis 3 |
| Cognitive generation agents | Produce hypotheses, analogies, negative-space maps, candidate methods, and alternative decompositions. | Thesis 2 |
| Experiment-design agents | Define causal scope, pre-registration packets, evaluation sets, baselines, and negative controls. | Thesis 3 |
| Implementation agents | Build prototypes, repair artifacts, generate tests, and package changes for validation. | Thesis 4 |
| Evaluator and validator agents | Detect false accepts, false rejects, regression, Goodharting, and evidence gaps. | Thesis 1 / Thesis 4 |
| Alignment and governance agents | Apply scoped trust, AAF dissent, owner override audit, ADO reporting, and containment constraints. | Thesis 5 |
| Memory and synthesis agents | Extract method memories, anti-patterns, playbooks, and post-cycle lessons. | Thesis 1 / Thesis 2 |

Role-family language prevents roster inflation. The question is not how many agents are named. The question is whether the workflow improves measured outcomes after integration cost.

The same warning applies at the organizational layer. Model 2's sub-additivity rule still binds: adding more agents, reviews, meetings, or artifacts can reduce net capability when coordination cost, contradiction, latency, or trust-review burden exceeds benefit. An AI-native R&D organization is therefore evaluated by net validated improvement after integration cost, not by organizational complexity.

## Brainstorming And Candidate Generation

Thesis 2 matters because an R&D organization needs search breadth. Brainstorming-related functions include:

- divergent option generation
- analogy and transfer candidate generation
- negative-space search
- contradiction hunting
- missing-evidence detection
- benchmark idea generation
- anti-pattern retrieval
- hypothesis packet construction
- stakeholder and perspective objection generation
- method-memory mutation proposals

These functions expand the candidate set. They do not validate it.

Hard rule:

> Brainstorming produces candidate search breadth. It does not produce accepted improvement until evaluator, validator, benchmark, cost, provenance, pre-registration, replication where applicable, and governance checks accept the result.

This rule is load-bearing. Without it, the organizational layer turns into research theatre.

The self-revising-discovery literature gives this candidate-generation stage a useful structure: Peirce's abduction generates candidate explanations, deduction derives their testable implications, and induction refines them from outcomes (Self-Revising Discovery Systems for Science: A Categorical Framework for Agentic Artificial Intelligence, arXiv:2606.01444). In Consullo terms, brainstorming is the abductive arm only — it produces candidate explanations and method mutations. Deduction corresponds to pre-registered prediction and experiment design; induction corresponds to evaluation, replication, and method-memory update. Keeping the three separated is the same load-bearing discipline as the hard rule above: an abductive stage allowed to also score its own output collapses the loop into research theatre.

Defixation discipline governs the *quality* of the abductive arm. Empirical evaluation of LLM idea generation finds that models exhibit human-like *fixation* — they over-produce conventional ideas and converge on familiar solution spaces, and the tendency worsens as task complexity grows (IDEAFix: Evaluation Framework for Creative Defixation Prompting in LLMs, arXiv:2606.00875, "IDEAFix: Evaluation Framework for Creative Defixation Prompting in LLMs"). Three findings are directly actionable for the brainstorming functions above. First, the levers that most increase originality are *explicit* divergence instructions and category negation — generate candidates, categorize them, then deliberately generate outside those categories — not elaborate creativity-method scaffolding; simple "go beyond the obvious" cues outperformed complex C-K- or SCAMPER-style prompts. Second, problem *framing* strongly shapes output: the same underlying problem stated as a different brief type, or seeded with surprising, atypical, or negatively-valenced attributes, yields markedly more novel and rare candidates, which argues for varying the framing of a brainstorming task rather than issuing it once. Third, and most important for a multi-agent organization, individual novelty gains do not fix *collective* homogenization: the study observed a persistent "LLM hivemind" in which many models converge into a shared semantic basin regardless of prompt. Spinning up many brainstorming agents therefore does not by itself guarantee a diverse candidate set — novelty must be scored as rarity against the ensemble of what other agents already produced, not only against a single agent's own output, and the negative-space, contradiction-hunting, and counter-paradigm functions in the list above are the structural defense against basin collapse. This is the creativity-side analogue of the alignment-monoculture and AAF dissent-collapse cautions elsewhere in the suite: breadth that is only nominal — many agents drawing from one basin — is not breadth.

The same evidence reshapes how brainstorming breadth is *sourced*. Benchmarking of scientific idea generation finds that divergent-thinking capability is largely decoupled from general intelligence and from model scale, and that it varies by domain — smaller or older models sometimes out-ideate larger, stronger ones, and a model's general-benchmark score does not predict its ideation value (Evaluating LLMs' divergent thinking, Nature Communications, doi:10.1038/s41467-026-70245-1, LiveIdeaBench). The organization should therefore not staff ideation by routing it to whichever model leads on general benchmarks. Ideation capability must be measured directly and per-domain; the brainstorming pool should be a diverse panel of generators selected for measured divergent thinking rather than a single strongest model, which also widens the candidate basin the failure mode below warns about. This is the model-selection counterpart to the prompt-level defixation discipline: defixation widens what one generator explores, while generator diversity widens whose basins are sampled at all.

Structured analogical reasoning is a specific, high-yield form of the analogy-and-transfer candidate-generation function above. Open-ended solution generation mode-collapses hard — in one evaluation only about 5% of proposed domains and a minority of solutions were unique, with models defaulting to the single canonical answer — and, critically, merely instructing an agent to "find solutions from other domains" does not fix it: unstructured cross-domain prompting still collapses (Unlocking LLM Creativity in Science through Analogical Reasoning, arXiv:2605.11258). What escapes the basin is *explicit* structure-mapping analogy: extract the problem's objects and the relations among them, map those relations — not surface attributes — onto a distant domain, then transfer that domain's method back as a candidate. Done this way, analogical reasoning becomes a diversity engine (roughly doubling solution diversity and finding novel solutions a majority of the time, versus a few percent for baselines) and it characteristically transfers a *method* across fields — a finite-mixture model from economics into perturbation prediction, a signal-to-noise threshold from telecommunications into cell-cell signaling. For Consullo this couples brainstorming to method memory: an analogical candidate is a proposal to reuse a method memory from a structurally similar but lexically unrelated task class (`../../technical-reports/method-memory/method-memory.md`), and like every brainstorming output it widens search without certifying a solution — it remains a candidate until the evaluator, validator, benchmark, and governance gates accept it.

Two cautions govern *adopting* creativity techniques across the agent fleet. First, an intervention that reliably helps human creators may do nothing for an LLM agent: the random cross-domain-mapping prompt that dependably raises human originality showed no average effect for LLMs, which already span remote associations without the human fixation bottleneck (Serendipity by Design: Evaluating the Impact of Cross-domain Mappings on Human and LLM Creativity, arXiv:2603.19087). A human creativity method should therefore not be imported into Consullo's prompts on the strength of its human track record — its benefit for a given agent is an empirical question, measured per agent and per domain, consistent with the separate finding that simple explicit cues outperformed elaborate human-method scaffolding for LLMs. Second, what *does* generalize across both humans and LLMs is the semantic distance of the inspiration source: more remote sources yield more original ideas, and not every pairing is generative. The organizational consequences are to select and rank inspiration sources by measured distance rather than drawing cross-domain sources at random, and to keep originality and feasibility as separate scored axes — the two trade off strongly, and an idea that looks infeasible only because the mechanism to realize it is not yet representable is not the same as a bad idea; that judgment belongs to the evaluator and validator stages, not to the generator.

## Pre-Registration And False-Progress Control

A research organization can fool itself by defining success criteria after results are observed. Consullo's organizational loop therefore treats pre-registration as an architectural requirement for experiments that may support capability claims.

A pre-registration packet should state:

- objective
- hypothesis
- null hypothesis
- expected mechanism
- success criteria
- failure criteria
- protected-set checks
- benchmark or evaluation suite
- expected cost
- affected capability or specification level
- lane classification
- required external-evaluator sampling

Exploratory work may still occur, but exploratory outputs cannot be promoted into capability claims unless later tested under pre-registered conditions.

## Institutional Memory

Organizational learning requires more than a successful patch. It requires memory that changes later behavior.

Consullo's institutional memory includes:

- method memories: reusable methods extracted from accepted work
- anti-patterns: rejected patterns, false starts, and failure modes
- playbooks: repeatable procedures for known task classes
- benchmark cases: positive, negative, adversarial, and held-out examples
- portfolio records: why lanes were funded, paused, expanded, or stopped
- incident records: how failures surfaced and what changed afterward

A memory artifact should not be promoted merely because it explains a past success. It should be tested for transfer. A Java-repair method memory, for example, should be evaluated against a non-Java task class before it supports a broader organizational-learning claim.

Institutional memory should also archive hypothesis genealogy — the lineage of which candidate led to which revision across a discovery lane's cycles — so that a self-revising loop remains interpretable and auditable after many iterations rather than only summarizable by its final result (Self-Revising Discovery Systems for Science: A Categorical Framework for Agentic Artificial Intelligence, arXiv:2606.01444).

## Governance Integration

The organizational layer introduces action classes that are more dangerous than ordinary code repair:

- successor-agent spawning
- external tool use
- benchmark modification
- validator-reference-suite modification
- model-family routing
- portfolio-level reprioritization
- batch deployment of multiple individually safe improvements
- public or customer-facing outputs

Composite or batch deployment is a gated action class even when individual component changes were accepted. These route through Thesis 5 when high-stakes, externally consequential, or capable of weakening future governance. The execution plan makes this operational through kill-switch drills, sandboxing, frozen reference suites, external-evaluator sampling, composite deployment decisions, and stop conditions.

The owner-as-final-AAF-adjudicator tension becomes sharper under live execution. This appendix does not solve that tension. It makes the tension visible and routes the operational controls to the execution plan.

## Relationship To The Execution Plan

This appendix is static publication scaffolding. the internal execution plan is the living operating document.

The appendix defines:

- what the organizational layer means
- how the five theses contribute
- why Java repair is insufficient
- why brainstorming is candidate generation
- what false-progress controls must exist
- where governance applies

The execution plan defines:

- week 0 safety preconditions
- kill-switch drills
- sandbox repository policy
- frozen `V_ref_0`
- external-evaluator sampling
- model-family diversity
- exploration budget floor
- stop conditions
- A-H capability ladder
- weekly operating loop
- immediate workstreams

The two documents should be revised together when execution changes the evidence boundary.

## Failure Modes

### Research Theatre

The system produces hypotheses, reports, benchmarks, or papers that look like research but do not survive pre-registration, held-out evaluation, or external review.

### Exploration Lane Collapse

Cheap software-maintenance wins dominate the portfolio, starving algorithmic, evaluator, memory, planning, and research-process exploration.

### Unapplied-Capability Lag

A tool, method, evaluator, or cognitive strategy is built and validated but then left unapplied to the open problems it could address, so the organization accrues capability without the downstream discoveries it was meant to enable. Historically such discovery time lags are large missed opportunities; the portfolio should track them rather than treat an idle capability as a neutral state.

### Candidate-Validation Collapse

Brainstorming outputs are treated as validated methods because they are plausible, fluent, or numerous.

### Fixation And Hivemind Convergence

Candidate generation defaults to conventional, familiar solutions (fixation), or many agents converge into one shared semantic basin (hivemind), so the candidate set looks numerous but is narrow. Per-agent idea counts and per-agent novelty hide the collapse. The defenses are measured rarity against the cross-agent ensemble rather than per-agent novelty, explicit defixation (category negation, counter-paradigm and negative-space arms, varied problem framing), and treating collective candidate diversity as a tracked portfolio metric (IDEAFix: Evaluation Framework for Creative Defixation Prompting in LLMs, arXiv:2606.00875).

### Benchmark Productivity Confusion

Benchmark designs or benchmark counts are treated as benchmark results.

### Organizational Goodhart

The organization optimizes dashboard metrics such as cycle count, hypothesis count, or accepted-change count instead of transferable capability improvement.

### Validator Capture

Agents learn the acceptance patterns of validators and produce evidence-shaped artifacts rather than real improvements.

### Capability Overhang From Uncommitted Improvements

Multiple individually acceptable changes accumulate and interact in ways that create emergent risk when deployed together.

### Owner Override Normalization

Repeated owner overrides become normal operation, weakening the AAF and scoped-trust discipline.

### Evaluation-Criterion Drift

A self-revising loop silently adjusts its own success criteria, thresholds, or evaluation suite until almost any output qualifies as an improvement. This is the discovery-system analog of validator capture, and is why mid-experiment revision of success criteria is prohibited; only pre-registered, `V_ref_0`-non-regressing evaluator changes are legitimate.

### Discovery-Loop Non-Termination

An autonomous discovery lane runs indefinitely with no external stop signal, consuming exploration budget while chasing diminishing returns past a plateau. Termination must depend on declared no-improvement bounds and resource budgets, not on the loop judging for itself that it is done.

## Minimal Demonstration Boundary

A minimal organizational-RSI demonstration would not need to show ASI or novel algorithmic discovery. It would need to show at least:

1. a pre-registered improvement objective
2. a candidate generated by a cognitive workflow
3. a prototype or implementation artifact
4. a benchmark result with negative controls
5. an adversarial review or external-evaluator sample
6. a governance disposition
7. a method-memory or anti-pattern update
8. a later cycle showing measurable benefit from that memory
9. no protected-set regression

Until such demonstrations exist, organizational RSI remains specified/proposed.

A discovery-lane demonstration should additionally record its declared termination bound and show that the loop actually halted on a stop signal — plateau, no-improvement-over-N, or budget — rather than on a self-judged completion. This is recorded alongside the nine items above, not as a tenth gating item, so that the boundary count referenced by the execution plan's capability ladder stays stable.

The execution plan's capability ladder should use this boundary for Level C and Level D claims. A method or research experiment that lacks these records may be useful exploratory work, but it should not be counted as organizational recursive self-improvement evidence.

## Publication Boundary

The publication suite may say:

> Consullo specifies an AI-native R&D organization layer in which agents, workflows, benchmarks, ledgers, and governance gates can be composed into a disciplined recursive-improvement loop.

It should not say:

> Consullo has demonstrated organizational-scale recursive self-improvement.

The former is an architecture claim. The latter requires execution evidence.
