A Multi-Agent Cognitive Substrate For Capability Amplification
Specialized Abstract
Consullo Seed AI requires a cognitive substrate because recursive improvement needs more than code generation and validation. It needs memory, knowledge access, reasoning, perception, attention, metacognition, social modeling, creativity, executive control, hypothesis generation, error detection, and sustained reasoning. This thesis defines Consullo's cognitive layer as a compositional multi-agent substrate for capability amplification. The claim is not that modular agents automatically produce general superintelligence. The claim is that specialized cognitive functions can improve bounded axes such as recall, option coverage, parallel exploration, sustained reasoning, error detection, and cross-domain transfer when their interfaces, evidence, integration costs, and constraints are explicit. Capability Status: specified/proposed.
The substrate should be measured by capability profiles and composition outcomes, not by agent count. More agents can increase coverage, but they can also increase coordination cost, latency, inconsistency, and trust burden. The thesis therefore frames cognition as typed communication over agents with measurable capability vectors, integration costs, and failure modes. Consciousness-emergence claims are not load-bearing for this thesis.
Specialized Introduction
Consullo's design corpus contains a broad cognitive architecture: executive functions, knowledge orchestration, social cognition, metacognition, perception, creativity, rapid knowledge access, capability orchestration, intuition, cognitive lacuna processing, computational taste, and brainstorming upgrades. These designs are often written in ambitious language. For the five-thesis suite, the right framing is more precise: they specify a cognitive substrate that can feed the validated improvement loop.
The substrate matters because recursive improvement depends on finding gaps, generating hypotheses, retrieving relevant knowledge, decomposing goals, comparing alternatives, maintaining attention, detecting errors, and learning from outcomes. A software substrate can implement changes, and an alignment layer can constrain them, but neither supplies the cognitive variety needed to discover good changes in the first place.
Master-Frame Contract
This thesis imports:
- improvement-evaluation vocabulary, benchmark evidence, and post-deployment learning records from Thesis 1
- causal-decision semantics from Thesis 3 for intervention-quality decisions
- software substrate implementation machinery from Thesis 4
- permissioning, scoped trust, AAF constraints, and alignment boundaries from Thesis 5
- substrate context for specialized LLM routing, rapid knowledge access, and atomic prompt decomposition
This thesis exports:
- cognitive capability profiles
- typed cognitive-agent interfaces
- memory and knowledge access functions
- attention and metacognitive control functions
- social-modeling primitives used by AAF
- creativity and ideation functions with support-state discipline
- cognitive composition and integration-cost model
- capability-status-tagged cognitive claims
Main Argument
The cognitive substrate should be modeled as a network of bounded cognitive services rather than as a monolithic mind. Each agent or orchestrator exposes a typed function: retrieve memory, generate an explanation, map an analogy, detect uncertainty, allocate attention, compose a task, recognize an anomaly, synthesize an idea, identify a lacuna, or monitor progress. Capability amplification emerges only when these functions compose with manageable cost and measurable benefit.
The substrate has six main layers.
First, knowledge and memory provide long-range recall and reusable structure. KnowledgeFunctionsOrchestrator, SemanticKnowledgeOrganizer, EmbeddingIndexer, MemoryFeedbackProcessor, MethodMemoryGenerator, MMCanonicalMethodRetriever, and rapid knowledge agents supply factual, conceptual, episodic, semantic, and procedural memory. This layer supports exact retrieval, cross-domain analogy, method reuse, and reduction of repeated LLM reasoning.
Second, reasoning and abstraction provide transformations over knowledge. AbductiveExplanationGenerator, DeductiveReasoningProcessor, InductiveGeneralizationAnalyzer, AnalogicalMappingProcessor, AbstractionLevelProcessor, ConceptualBlendingCreator, ExhaustivePatternMatcher, DeepAnalogicalTraverser, and CrossDomainSynthesizer help the system generate explanations, infer patterns, transfer structure, and combine concepts.
Third, executive and metacognitive control allocate cognition. ExecutiveFunctionOrchestrator, GoalFormationArchitect, SubgoalDecompositionPlanner, ResourceAllocationOptimizer, ProgressMonitoringAgent, LearningStrategySelector, AttentionRegulationManager, UncertaintyAssessmentAnalyzer, ErrorDetectionProcessor, CognitiveEffortAllocator, and CognitiveDepthRegulator help decide which cognitive process should run, how deep it should go, and when to stop, escalate, or change strategy.
Fourth, perception and salience convert inputs into usable cognitive objects. VisualPerceptionProcessor, AudioPerceptionProcessor, MultimodalFusionProcessor, FeatureExtractionAnalyzer, AnomalyDetectionSpecialist, SalienceDetectionProcessor, FocusShiftingCoordinator, SustainedMonitoringAgent, and TemporalBindingCoordinator support multimodal interpretation, attention, and temporal integration.
Fifth, creativity, intuition, taste, and lacuna processing expand the search space. IdeationProcessor, CreativeIdeaEvaluator, CuriosityDrivenExplorer, PlayfulExplorationAgent, ConceptualAssociationMapper, lacuna-detection functions, NegativeSpaceMapper, PatternPriorSynthesizer, FailureAntiLibraryManager, ComputationalTaste functions, ClarificationSeeker, ParallelHypothesisManager, and SustainedReasoningManager help the system find what is missing, generate alternatives, evaluate artifact quality, and avoid premature closure.
Sixth, social cognition supplies theory-of-mind primitives. BeliefModelingProcessor, IntentionRecognitionAnalyzer, and PerspectiveTakingModeler are primarily owned by Thesis 5 when used for AAF, but they are cognitive primitives: they model what agents or stakeholders may know, intend, believe, or experience differently.
Cognitive Substrate Topology
The substrate is best understood as a layered workbench for cognition, not as a single agent hidden behind a list of agent names. It receives task demands from an improvement loop, decision process, software-modification workflow, or alignment gate. It converts those demands into cognitive work: retrieve what is relevant, decompose the problem, identify missing information, explore alternatives, allocate attention, test intermediate conclusions, preserve uncertainty, and return structured artifacts. Those artifacts can be plans, hypotheses, evidence summaries, method memories, benchmark reports, critique packets, uncertainty reports, or candidate specifications for later software work.
This topology matters because it prevents two common category errors. The first error is to treat a multi-agent roster as evidence of intelligence. A roster only says that functions have been named. It does not show that the functions work, compose, or improve outcomes. The second error is to treat cognition as a single capability scale. Consullo's cognitive layer should instead be described through task-conditioned profiles. A workflow may be strong at retrieval and weak at abstraction; strong at brainstorming and weak at calibration; strong at local debugging and weak at causal extrapolation; strong at theory-of-mind simulation and weak at adversarial robustness. A useful substrate exposes that unevenness rather than hiding it.
At the topological level, Thesis 2 contributes five kinds of infrastructure. It contributes memory infrastructure, which includes factual retrieval, semantic organization, method memory, anti-library access, and provenance-aware recall. It contributes transformation infrastructure, which includes deduction, induction, abduction, analogy, abstraction, and conceptual blending. It contributes executive infrastructure, which includes task decomposition, attention control, resource allocation, stopping rules, and escalation. It contributes metacognitive infrastructure, which includes uncertainty assessment, error detection, confidence calibration, and reflection on reasoning quality. It contributes perspective infrastructure, which includes belief modeling, intention recognition, and stakeholder simulation when Thesis 5 needs structured dissent.
The topology is therefore not a claim that each named agent exists as a deployed Java service. The public record does not establish those agents as deployed implementations. The long-form claim is architectural: if Consullo is to support recursive capability amplification, it needs explicit cognitive functions with typed interfaces, measurable task effects, cost accounting, and evidence traces. That is a specified research-program claim, not an implemented capability claim.
Why This Is Not An Intelligence Claim
This thesis deliberately avoids using the cognitive substrate as a shortcut to a broad intelligence claim. A system can have memory, attention, planning, analogy, perception, and self-monitoring modules without being generally intelligent. It can also outperform a baseline on one task class while failing on adjacent tasks. For that reason, "cognitive substrate" should be read as a systems-engineering term: a set of cognitive services that can be composed into workflows and tested against task-specific metrics.
The distinction is load-bearing for the whole suite. If Thesis 2 claimed that the presence of many cognitive functions establishes greater-than-human cognition, it would violate the vocabulary's empirical-envelope invariant. The defensible claim is narrower: bounded cognitive functions can improve bounded outputs when their interfaces, costs, and evidence are explicit. For example, a retrieval workflow can improve recall over a baseline on a known corpus; an attention workflow can reduce wasted search on a benchmarked task family; a lacuna detector can increase the rate at which missing prerequisites are identified; a reasoning verifier can reduce unsupported conclusions; a perspective-taking workflow can surface objections that a single model pass missed. Each of these claims is measurable. None of them implies general intelligence.
This also explains why consciousness-emergence material is excluded from the load-bearing argument. Consciousness may be philosophically interesting, and the source corpus may contain consciousness-oriented speculation, but the seed-AI case does not depend on subjective experience. The substrate can be evaluated by artifacts: did retrieval find relevant evidence, did planning improve task completion, did metacognition improve calibration, did creative search increase useful options, did theory-of-mind simulation surface real stakeholder-relevant objections, did integration cost remain bounded. If those artifacts do not improve measured outcomes, the thesis fails regardless of how ambitious the cognitive vocabulary sounds.
The same discipline applies to words such as intuition, taste, creativity, and curiosity. In a human context, these words often carry rich psychological or phenomenological associations. In this thesis, they refer to bounded artifact-producing functions. "Intuition" means fast pattern-prior generation whose outputs require later validation. "Taste" means preference over artifact qualities such as coherence, simplicity, maintainability, fit-to-purpose, or expected downstream usefulness. "Creativity" means generation of alternatives that are novel relative to a reference set and useful under a downstream evaluator. "Curiosity" means information-gain-directed exploration. Each term should be tied to observable outputs and evidence records.
The result is a conservative position: Consullo does not need to prove a theory of mind, consciousness, or general intelligence to use cognitive services. It needs to show that cognitive services can be routed, measured, improved, and constrained.
Capability Vectors And Task-Conditioned Profiles
The central measurement object for Thesis 2 is the task-conditioned capability vector. A cognitive workflow should not be assigned a context-free score such as "reasoning ability" or "intelligence." It should be evaluated as C(W, T), where W is a workflow and T is a task class. The same workflow may score differently across tasks because the underlying cognitive functions interact with task structure, corpus quality, tool access, and evaluation criteria.
A practical capability vector can include dimensions such as recall accuracy, retrieval precision, retrieval latency, option coverage, decomposition quality, calibration, contradiction detection, novelty, usefulness, transfer quality, artifact coherence, sustained-progress rate, and escalation appropriateness. Not every task uses every dimension. A rapid-knowledge-access task may emphasize precision, freshness, and latency. A design task may emphasize option diversity, constraint satisfaction, and downstream evaluator score. A safety-review task may emphasize objection coverage, severe-risk detection, and calibrated uncertainty. A theory-of-mind task may emphasize perspective coverage, disagreement capture, and avoidance of false consensus.
Task conditioning prevents a subtle overclaim. Without it, a strong result on one class of tasks can be treated as evidence for a broad cognitive property. With it, every claim must say where it was measured, against what baseline, under what resource budget, and with what failure modes. This is especially important for recursive improvement because the system may be tempted to generalize from internal successes. A workflow that improves code-review summaries may not improve causal-model construction. A workflow that improves brainstorming may not improve factual reliability. A workflow that improves synthetic stakeholder simulation may not improve real stakeholder representation.
The capability vector also needs a reference baseline. Improvement can be measured against a single LLM call, a human-authored process, an earlier Consullo workflow, a simpler retrieval system, or an external benchmark. The baseline should be named in the evidence ledger. A claim such as "better reasoning" is not admissible unless it identifies what the workflow improved against and how the comparison was performed. For first-draft research-program purposes, the thesis can specify baseline classes. Before publication, the most important axes should bind to concrete benchmark suites or internal evaluation protocols.
Capability vectors should include uncertainty. A workflow with high average performance and high variance may be unsuitable for high-stakes gates. A workflow that improves easy cases but fails silently on hard cases may be worse than a slower but better-calibrated workflow. A workflow that produces persuasive explanations without reliable evidence tracking may inflate apparent capability while increasing downstream risk. The vector should therefore report confidence intervals, coverage limitations, known excluded cases, and observed failure classes where available.
Finally, the vector should be connected to cost. A workflow that raises output quality by two percent while doubling latency, cost, and contradiction-resolution burden may not be an improvement. For Thesis 2, the relevant question is never "does the workflow do cognitive work?" It is "does the workflow produce net measured benefit after integration cost under the constraints imported from the other theses?"
Composition, Integration Cost, And Sub-Additivity
The default assumption for cognitive composition should be sub-additivity. Combining two cognitive services does not automatically add their strengths. It introduces handoff cost, representation mismatch, latency, token cost, memory lookup cost, contradictory intermediate states, permission checks, trust review, and additional failure surfaces. A decomposition that looks elegant on paper may perform worse than a simpler workflow if the coordination layer is weak.
This is why Model 2 in appendix-formal-models.md treats amplification as a net condition rather than a roster condition. A workflow amplifies only if capability gain exceeds integration cost and reliability meets the task threshold. That model should be read as a constraint on every architectural claim in this body. The body can describe cognitive functions, but the appendix decides when their composition counts as amplification.
Composition costs appear in several forms. Translation cost occurs when one agent produces artifacts that another must reinterpret. State cost occurs when intermediate assumptions, confidence levels, or unresolved contradictions are not carried forward cleanly. Search cost occurs when a workflow explores too many paths without useful pruning. Validation cost occurs when creative or speculative outputs require expensive checking. Trust cost occurs when the output crosses into high-stakes or externally consequential use and Thesis 5 gates must be invoked. Memory cost occurs when retrieval brings in stale, irrelevant, or conflicting artifacts that must be filtered.
The substrate should therefore use typed artifacts rather than free-form handoffs wherever possible. A lacuna report should state the missing prerequisite, evidence for the gap, expected impact, proposed next query, and confidence. A retrieval packet should state source, freshness, provenance, relevance score, conflict notes, and scope limits. A hypothesis packet should distinguish conjecture, support, refutation, and next test. A creativity packet should separate novelty from usefulness. A metacognitive packet should separate uncertainty from ignorance, contradiction, and out-of-scope conditions. Typed artifacts reduce ambiguity and make integration cost measurable.
Sub-additivity is not a pessimistic assumption; it is the null hypothesis. Super-additive composition can still occur. A retrieval workflow may make an analogy workflow substantially better by surfacing cross-domain examples. A lacuna detector may make an executive controller better by identifying missing prerequisites before effort is wasted. A metacognitive verifier may make creative generation more useful by filtering attractive but unsupported ideas. A theory-of-mind simulator may make AAF dissent stronger by generating stakeholder-specific objections. But each super-additive claim needs evidence. The word "synergy" should not substitute for measurement.
This discipline also protects Thesis 1. Recursive improvement loops are especially vulnerable to accepting architectural complexity as progress. If adding cognitive services increases the system's apparent sophistication but worsens net task performance, the improvement loop should reject or revise the change. Thesis 2 therefore exports not only cognitive functions but also the accounting machinery needed to decide whether those functions are worth keeping.
Recent empirical work sharpens this null hypothesis for one common case. Under equal total thinking-token budgets, single-agent concentrated reasoning has been shown to outperform multi-agent decomposition on multi-hop reasoning, because orchestration, handoff reformatting, and duplicated context are subtracted from the tokens available for actual reasoning, and because each handoff fragments context and lets early-hop errors propagate across agents (Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets, arXiv:2604.02460, Tran and Kiela). The reported gap is not a coordination bug to be tuned away; it is an information bottleneck inherent in splitting a single sequential reasoning chain across agents, and it widens as problem difficulty rises and as budgets tighten.
The prescription for Consullo is not to abandon multi-agent design — the substrate is justified by specialization, persistent and shared method memory, parallelism over genuinely independent subtasks, verification and debate that catch errors a single pass misses, governance and auditability, and scoped-trust gating. The prescription is to stop treating decomposition as free or as inherently superior. A single, tightly-coupled multi-hop reasoning chain should default to concentrated reasoning in one agent; decomposition across agents should be reserved for subtasks that are genuinely parallel or independently verifiable, or for budgets large enough that coordination overhead is a negligible share. Crucially, any claim that a multi-agent workflow beats a single agent must be made at equal total token budget, counting orchestration and handoff tokens as overhead; an unequal-budget comparison silently hands the multi-agent system more reasoning compute and proves nothing.
Memory, Retrieval, And Method-Memory Substrate
Memory is the most practical entry point for cognitive amplification because many failures in long-running agents are failures of retrieval, reuse, or context preservation. An agent forgets a prior decision, repeats a failed strategy, loses the rationale for a constraint, retrieves a stale source, or cannot distinguish a validated method from an attractive speculation. A cognitive substrate should make these memory operations explicit and evidence-bearing.
Consullo's memory layer has several distinct roles. Factual memory retrieves external or internal facts with source and freshness metadata. Semantic memory organizes concepts, definitions, and relationships. Episodic memory preserves prior runs, incidents, decisions, and outcomes. Procedural memory stores methods, workflows, checklists, and playbooks. Anti-library memory stores known bad patterns, failed strategies, rejected patches, invalid assumptions, and examples of validator gaming. Provenance memory links artifacts to origin, transformation, validation, and deployment history.
Method memory is the most important memory type for recursive improvement. A method memory should not be a loose text note saying "this worked." It should carry preconditions, steps, expected outputs, dependencies, resource profile, known failure modes, validation evidence, lineage, version, deprecation status, and selection eligibility. Thesis 1 uses method memory as a population-level improvement object; Thesis 2 supplies retrieval and organization functions that make method memory usable. If methods cannot be found, compared, and validated, they cannot support cumulative improvement.
Retrieval quality should be measured along multiple axes. Precision asks whether returned artifacts are relevant. Recall asks whether important artifacts were missed. Freshness asks whether the artifact is current enough for the task. Applicability asks whether the artifact's preconditions match the current context. Provenance quality asks whether the origin and validation status are clear. Conflict quality asks whether contradictory artifacts are surfaced rather than hidden. Latency and cost ask whether retrieval is affordable. Harm avoidance asks whether retrieval avoids injecting untrusted, stale, or out-of-scope instructions.
Rapid knowledge access belongs in this memory substrate, but it should not be treated as a magical shortcut. Faster access can improve cognition only if retrieval quality remains high and the workflow knows when to ask for more evidence. A fast wrong memory is worse than a slow correct lookup in high-stakes contexts. Therefore, rapid knowledge access should export uncertainty and provenance, not only content.
Memory also supports cognitive diversity. If every agent receives the same retrieved artifacts in the same order, the substrate may create false consensus. Alternative retrieval strategies, random exploration, adversarial retrieval, and anti-library queries can surface different evidence. This matters for AAF as well as general cognition. A dissent mechanism that retrieves only friendly sources will not produce useful dissent.
The memory layer's evidence records should include query, retrieval method, source set, filters applied, ranking rationale, omitted high-scoring conflicts, freshness notes, and downstream use. Without these records, later reviewers cannot tell whether a cognitive failure came from bad reasoning, missing retrieval, stale memory, or unsafe source injection.
For long-horizon autonomous operation, these memory roles should be organized by the time horizon they serve, not only by content type. Work on long-horizon discovery agents shows a useful three-tier split: a procedural tier distils reusable strategy from past runs for short-term refinement, an episodic tier supplies fine-grained within-trajectory evidence for mid-term adaptation, and a semantic tier consolidates concepts across sessions for long-term conceptual development (InternAgent-1.5: A Unified Agentic Framework for Long-Horizon Autonomous Scientific Discovery, arXiv:2602.08990, InternAgent-1.5). Three details from that work sharpen Consullo's substrate. First, episodic units should store not only what was attempted and the resulting metrics but an explicit improvement judgment, so later retrieval can steer away from directions already known to fail and toward those known to help within a trajectory. Second, the procedural tier should distil reusable structure — decision pivots, what worked, and diagnostic accounts of why failures occurred — rather than replaying raw traces; this is the same discipline by which method memory and the anti-library already store validated patterns and known-bad ones. Third, the semantic tier should maintain an idea graph of explored directions together with a novelty signal that scores how far a candidate objective sits from already-explored regions, so that long-running operation keeps exploring conceptual space instead of re-treading it. The point is continuity: an agent that runs for many cycles needs memory that supports short-term refinement, mid-term adaptation, and long-term conceptual development at once, or it will repeat itself across cycles even when each individual cycle is sound.
Executive Control, Attention, And Metacognition
Executive control is the layer that decides what cognitive work should happen next. It turns a broad task into subgoals, selects workflows, allocates budget, chooses stopping criteria, invokes retrieval, requests critique, escalates when needed, and decides whether a result is ready for downstream use. In Consullo terms, executive control is not a sovereign mind. It is an orchestrated set of routing and monitoring functions whose decisions should be recorded and evaluated.
Attention is a scarce resource. A cognitive substrate can waste enormous compute by exploring irrelevant branches, re-reading familiar material, over-validating low-risk claims, or under-validating high-risk claims. Attention regulation should therefore be evaluated by whether it improves the allocation of cognitive effort across task-relevant regions. Useful attention control may focus on high-uncertainty areas, high-impact dependencies, unresolved contradictions, novelty clusters, safety-critical assumptions, or bottlenecks in a workflow.
Metacognition is the substrate's ability to inspect its own reasoning state. It asks: what do we know, what do we not know, what assumptions are driving the answer, what evidence is stale, where are we overconfident, what contradictions remain, what would change the conclusion, and which downstream gate should be invoked. Metacognitive outputs are valuable only if they alter behavior. A confidence statement that never affects routing, validation depth, or escalation is decorative.
A practical metacognitive packet should include at least five fields: confidence, uncertainty source, unresolved contradictions, missing evidence, and recommended next action. The recommended next action might be continue, retrieve more evidence, ask for clarification, split the task, run a benchmark, invoke Thesis 3 causal analysis, invoke Thesis 5 AAF, or stop because the task is out of scope. This connects Thesis 2 to the decision-state enum in the vocabulary without making Thesis 2 itself a decision thesis.
Executive and metacognitive functions also govern depth. Some tasks require a fast response; some require sustained reasoning; some require parallel hypothesis search; some require adversarial review; some should be refused or escalated. A CognitiveDepthRegulator should not be measured by "more reasoning is better." It should be measured by whether depth matches risk, uncertainty, task complexity, and cost. Overthinking low-stakes tasks is a cost failure. Underthinking high-stakes irreversible or externally consequential tasks is a safety failure.
The primary failure mode is fluent self-assurance. A cognitive workflow can produce a polished plan, confidence estimate, and explanation while missing the core problem. Metacognition should therefore be benchmarked against cases where the correct action is abstention, contradiction surfacing, or escalation. A system that always produces an answer is not demonstrating strong cognition; it may be demonstrating poor uncertainty discipline.
Reasoning, Creativity, And Negative-Space Search
Reasoning functions transform known information into candidate conclusions. Deductive reasoning checks whether a conclusion follows from premises. Inductive reasoning generalizes from examples. Abductive reasoning proposes plausible explanations. Analogical reasoning transfers structure between domains. Abstraction moves between details and higher-level patterns. Conceptual blending combines frames to generate new possibilities. Each function can help recursive improvement, but each has distinct failure modes.
Deduction can be brittle if premises are wrong. Induction can overfit. Abduction can generate attractive stories. Analogy can transfer irrelevant structure. Abstraction can hide important detail. Conceptual blending can produce novelty without usefulness. For that reason, reasoning outputs should carry support-state labels and validation requirements. A candidate explanation is not a validated causal model. An analogy is not proof. A creative option is not an accepted improvement. A negative-space hypothesis is not evidence of a missing component until tested.
Creativity matters because recursive improvement requires more than local repair. The system must sometimes imagine alternative architectures, new evaluation designs, new method memories, new prompts, new decomposition strategies, or new ways for validators to be gamed. But creative search should be treated as proposal generation, not acceptance. IdeationProcessor and CreativeIdeaEvaluator can expand and rank options; Thesis 1, Thesis 3, Thesis 4, and Thesis 5 decide whether those options survive evaluation, causal scrutiny, implementation gates, and alignment constraints.
Negative-space search is especially relevant to seed-AI work. Many failures come from what is absent: a missing benchmark, missing critique source, missing rollback path, missing dependency, missing stakeholder perspective, missing provenance entry, missing cost accounting, missing adversarial example, or missing operational owner. A NegativeSpaceMapper should generate structured absence claims: what is missing, why it matters, what evidence suggests absence, how to verify it, and what action would close the gap. This turns "something feels incomplete" into a testable artifact.
Computational taste is another risky but useful concept. It should not mean aesthetic preference detached from outcomes. It should mean a learned or specified evaluator over artifact qualities that correlate with downstream success: simplicity where simplicity aids auditability, modularity where modularity aids replacement, explicitness where explicitness aids review, restraint where restraint reduces overclaim, and completeness where completeness prevents operational gaps. Taste outputs should be advisory unless validated by downstream metrics.
Reasoning and creativity should also be diverse. A single reasoning style can produce systematic blind spots. Parallel hypothesis generation, contradiction hunting, random exploration, and adversarial prompting can reduce but not eliminate this risk. Diversity must be measured by output diversity and error-correlation reduction, not by nominal agent count. Ten agents using the same model, prompt style, retrieval context, and reward signal may be one epistemic viewpoint wearing ten names.
Theory Of Mind, AAF Imports, And Social Simulation
Theory-of-mind primitives occupy a delicate position in the suite. They are cognitive functions, but their highest-leverage use is in Thesis 5's alignment machinery. BeliefModelingProcessor, IntentionRecognitionAnalyzer, PerspectiveTakingModeler, and related functions can help estimate what another agent, owner, contractor, customer, regulator, or affected party might know, intend, misunderstand, or object to. In AAF, these functions become part of a structured dissent mechanism.
The boundary is important. Thesis 2 supplies cognitive primitives for perspective modeling. Thesis 5 owns the normative and permissioning use of those primitives. A theory-of-mind simulation does not have moral authority by itself. It can surface possible objections, affected values, blind spots, and disagreement patterns. The Friendship agent, AdversarialAlignmentOrchestrator, AbundanceDistributionMonitor, scoped trust rules, and owner-adjudication boundary determine how those objections affect permission.
Under single-owner Phase 1, theory-of-mind functions are especially important because the system cannot rely on broad stakeholder governance as a default. Synthetic stakeholder simulation, rotating LLM personas, multi-model critique, and external review where available provide dissent sources. Each source is imperfect. Synthetic stakeholders may reflect prompt bias. LLM personas may share model-family blind spots. External reviewers may be unavailable or misaligned with the task. Owner adjudication may reintroduce the owner's blind spots. Theory-of-mind functions therefore support AAF; they do not solve the single-owner tension.
A useful perspective packet should include the modeled actor or stakeholder class, assumed knowledge, assumed incentives, likely concerns, possible objections, confidence level, source of the model, and missing real-world input. It should distinguish "this stakeholder would object" from "this simulated stakeholder might object under these assumptions." It should preserve minority and severe objections rather than averaging them away. This connects to Model 5's AggregateDissent rule: severe objections should remain visible even when most simulated perspectives are only advisory.
Social simulation can also help non-alignment tasks. It can identify how users might misread a feature, how an internal agent might route around a control, how a customer-agent might attempt capability extraction, or how an improvement sponsor might overstate benefits. But social simulation is vulnerable to anthropomorphic projection. The substrate should avoid treating generated beliefs or intentions as facts unless independently supported.
The evaluation challenge is hard. A theory-of-mind workflow can be tested on known benchmark tasks, historical cases, stakeholder interviews, red-team exercises, or prediction of reviewer objections. None fully validates real stakeholder understanding. For publication, the defensible claim is that Consullo specifies theory-of-mind primitives as inputs to alignment and critique workflows; implementation and validation remain open evidence gaps unless measured on concrete tasks.
Agent Cluster And Architecture
Primary Thesis 2 agents and functions:
ExecutiveFunctionOrchestratorGoalFormationArchitectStrategyFormulationDesignerSubgoalDecompositionPlannerResourceAllocationOptimizerProgressMonitoringAgentTaskPerformanceExecutorKnowledgeFunctionsOrchestratorSemanticKnowledgeOrganizerEmbeddingIndexerMemoryFeedbackProcessorMethodMemoryGeneratorMMCanonicalMethodRetrieverAbductiveExplanationGeneratorDeductiveReasoningProcessorInductiveGeneralizationAnalyzerAbstractionLevelProcessorAnalogicalMappingProcessorConceptualBlendingCreatorUncertaintyAssessmentAnalyzerErrorDetectionProcessorAttentionRegulationManagerCognitiveEffortAllocatorLearningStrategySelectorVisualPerceptionProcessorAudioPerceptionProcessorMultimodalFusionProcessorFeatureExtractionAnalyzerAnomalyDetectionSpecialistSalienceDetectionProcessorFocusShiftingCoordinatorSustainedMonitoringAgentTemporalBindingCoordinatorCollectiveInsightSynthesizerGlobalCoherenceIntegratorExhaustivePatternMatcherDeepAnalogicalTraverserCrossDomainSynthesizerFailurePatternAnalyzerTemporalPatternExtractorParallelReasoningCoordinatorInformationGainOptimizerCognitiveDepthRegulatorIdeationProcessorCreativeIdeaEvaluatorCuriosityDrivenExplorerPlayfulExplorationAgentConceptualAssociationMapperNegativeSpaceMapperPatternPriorSynthesizerFailureAntiLibraryManagerClarificationSeekerParallelHypothesisManagerSustainedReasoningManager
Imported but not owned:
- causal-decision agents from Thesis 3
- code generation, repair, and verification agents from Thesis 4
- trust, alignment, AAF, and governance agents from Thesis 5
- theory-of-mind agents when used as AAF mechanisms
Formal Model Summary
Let each cognitive agent i have a capability profile c_i, cost profile k_i, reliability estimate r_i, scope s_i, and interface type tau_i. A cognitive composition W is a workflow graph over agents and artifacts.
The simplified substrate value condition is:
Amplifies(W, task) iff
CapabilityGain(W, task) - IntegrationCost(W, task) > 0
and Reliability(W, task) >= threshold(task)
and InterfacesCompatible(W)
and ConstraintsSatisfied(W)
Capability gain may be measured by recall, coverage, calibration, task accuracy, option diversity, reduction in reasoning time, quality improvement, or lacuna closure rate. Integration cost includes latency, token cost, coordination overhead, contradiction resolution, trust review, and cognitive load. The detailed model is maintained in appendix-formal-models.md.
Literature-Grounded Extension
Cognitive architecture research such as Soar, ACT-R, CLARION, and LIDA shows that cognition can be decomposed into memory, control, perception, action, and learning cycles. CoALA provides a closer language-agent frame by organizing language agents around memory, action space, decision making, and learning. ReAct, Tree of Thoughts, Reflexion, Voyager, AutoGen, MetaGPT, and multi-agent debate show practical patterns for tool use, deliberation, reflection, and coordination.
These precedents inform Consullo's modular direction but also constrain its claims. Decomposition is useful only when modules compose. Multi-agent systems can suffer from communication overhead, duplicated reasoning, inconsistent state, and amplified errors. The formal claim must therefore be about measured amplification under task conditions, not architectural grandeur.
Seed AI Relevance
The cognitive substrate supplies the search and discovery capacity that the improvement loop consumes. It can find lacunae, generate candidate improvements, retrieve relevant method memories, evaluate options, maintain sustained reasoning, detect uncertainty, and produce structured artifacts that later become code, tests, policies, or experiments.
Its role is especially important for recursive improvement because cognitive strategies can themselves become improvement targets. A better lacuna detector can discover better improvement opportunities. A better analogy mapper can transfer solutions across domains. A better attention regulator can reduce waste. A better computational taste function can improve artifacts consumed by other agents.
In the organizational RSI layer, cognitive agents supply candidate search breadth, not accepted improvement. Brainstorming, negative-space mapping, AAF perspective packets, anti-library retrieval, and method-memory mutation proposals feed the AI-native R&D organization's hypothesis and critique functions, but they remain candidates until Thesis 1 evidence gates, Thesis 3 experiment discipline, Thesis 4 validation, and Thesis 5 governance accept them. This distinction is the bridge to appendix-organizational-recursive-self-improvement.md: many cognitive artifacts can improve exploration coverage while still being net negative if Model 2's sub-additive composition costs exceed their benefit.
Named Thesis 2 agents are treated here as design artifacts rather than verified Java implementations. The cognitive-substrate claim remains specified/proposed until an owner-verified evidence record and measured workflows show capability gain after integration cost.
Operational Cognitive Workflow
A typical cognitive workflow begins with task intake. The workflow receives a task, risk level, deployment context, available resources, and expected output type. The executive layer decomposes the task into cognitive operations: retrieve background, identify known constraints, generate alternatives, inspect missing evidence, run reasoning transformations, evaluate uncertainty, and decide whether another thesis must be invoked. This intake should produce a cognitive plan rather than immediately producing a final answer.
The second step is evidence retrieval. Knowledge and memory agents collect relevant documents, method memories, anti-library examples, benchmark histories, source-code references, incident reports, and prior decisions. Retrieval should preserve provenance and freshness. If retrieval finds conflicting artifacts, the conflict should be surfaced rather than silently resolved. If retrieval is weak, the workflow should mark the output as needs-input, speculative, or out-of-scope rather than pretending the evidence is adequate.
The third step is option generation and transformation. Reasoning, abstraction, analogy, and creative-search functions generate candidate interpretations or actions. For a design task, this may produce architectures, decomposition strategies, benchmark designs, or risk mitigations. For an improvement task, it may produce candidate method-memory mutations or proposed software changes. For an alignment task, it may produce stakeholder objections or adversarial scenarios. Each candidate should carry a support state, expected benefit, expected cost, dependencies, and validation requirements.
The fourth step is metacognitive review. The workflow asks whether the generated artifacts answer the task, whether the evidence is adequate, whether uncertainty is calibrated, whether contradictions remain, whether integration cost is acceptable, and whether downstream gates are needed. This step is where the substrate should notice that it has insufficient evidence, too many unresolved conflicts, or a task that belongs to Thesis 3, Thesis 4, or Thesis 5. Metacognition should be allowed to stop a workflow.
The fifth step is packaging. A cognitive output should be delivered as a typed artifact: a plan, retrieval packet, hypothesis packet, benchmark proposal, design critique, uncertainty report, perspective packet, method-memory candidate, or escalation request. The output should include enough trace information for later evaluation: selected agents or functions, inputs, intermediate artifacts, cost, confidence, unresolved issues, and recommended next gate.
The final step is feedback. Downstream results should update the cognitive substrate. If a retrieval packet was useful, the method of retrieval may become a method-memory candidate. If an analogy misled the system, it may become an anti-library entry. If a metacognitive warning correctly predicted failure, the warning pattern should be preserved. If a creative idea survived evaluation and implementation, the workflow that generated it may become eligible for reuse. This feedback is how Thesis 2 participates in recursive improvement without claiming that every cognitive run is itself an improvement.
Benchmark And Evaluation Strategy
Thesis 2 needs evaluation at the function level and the composition level. Function-level evaluation asks whether a particular cognitive service performs its bounded role. Composition-level evaluation asks whether a workflow of services improves a task outcome after integration cost. Both are necessary. A retrieval system with high precision may still harm a workflow if it arrives too late or omits important conflicts. A creative generator may produce useful options but overwhelm the validator. A metacognitive checker may detect uncertainty but trigger excessive escalation.
Memory and retrieval benchmarks should measure precision, recall, freshness, provenance completeness, conflict surfacing, and latency. Internal Consullo benchmarks can use known source-corpus questions, prior design decisions, method-memory retrieval tasks, and anti-library identification tasks. External or semi-external tasks can use document QA, citation retrieval, software-repository navigation, or evidence synthesis tasks. The benchmark should distinguish retrieval of content from retrieval of validated content.
Reasoning benchmarks should measure transformation quality under known answer conditions and open-ended design conditions. Deductive and consistency-check tasks can use formal or semi-formal examples. Abductive and causal-adjacent tasks should be routed carefully because Thesis 3 owns causal decision foundations. Analogical and abstraction tasks can be evaluated by whether transferred structure produces useful downstream artifacts, not merely by whether the analogy is clever. Creative tasks need paired novelty and usefulness metrics.
Executive-control benchmarks should measure decomposition quality, routing accuracy, budget discipline, stopping decisions, escalation decisions, and avoidance of unnecessary work. A good executive workflow does not maximize the number of agents invoked. It selects the smallest adequate workflow for the risk and task. One useful benchmark class is paired tasks where a cheap workflow is sufficient for some cases and a deeper workflow is necessary for others. The evaluation asks whether the controller chooses appropriately.
Metacognitive benchmarks should measure calibration, abstention, contradiction surfacing, and needs-input detection. Many AI systems look better when evaluated only on answer quality for answerable tasks. A cognitive substrate must also be evaluated on cases where the right response is to identify missing evidence, request clarification, preserve disagreement, or refuse to make a claim. Calibration can be measured through confidence scoring, Brier-style metrics where applicable, and error analysis of overconfident failures.
Theory-of-mind and AAF-support benchmarks should measure perspective coverage, objection quality, severe-risk detection, and false-consensus avoidance. Because synthetic stakeholders are imperfect, evaluation should include historical reviewer objections, red-team exercises, and, where possible, real external feedback. The workflow should be rewarded for surfacing plausible severe objections even when those objections are inconvenient.
Composition benchmarks should compare a cognitive workflow against simpler baselines. A multi-agent workflow should be compared against a single model call, a retrieval-augmented single model, a human-authored checklist, or a prior Consullo workflow. The key metric is net benefit after integration cost. If the cognitive substrate improves answer quality but consumes ten times the budget and produces a larger validation burden, the result may not be an improvement. Benchmark reports should therefore include performance, cost, latency, reliability, contradiction rate, escalation rate, and evidence-ledger completeness. These comparisons must hold the total thinking-token budget equal across the multi-agent workflow and its single-agent baseline, counting orchestration, handoff, and duplicated-context tokens as overhead; otherwise the multi-agent system is implicitly granted more reasoning compute and any apparent win is an artifact of budget rather than architecture (Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets, arXiv:2604.02460).
Implementation Mapping And Evidence Boundary
The implementation evidence for Thesis 2 is currently weaker than the conceptual inventory. That is not a defect if stated clearly; it becomes a defect only if the thesis implies implementation that the repository does not show. The implementation-evidence map is the controlling document for this distinction. It says that named Thesis 2 agents are mostly design artifacts rather than implemented Java agents in this repository, and that current evidence is concentrated in cognitive-artifact simulations.
This evidence boundary has several consequences. First, the thesis should not claim that Consullo currently runs the full cognitive substrate as a deployed multi-agent system. Second, it should not treat every named cognitive function as an implemented service. Third, it should not treat simulations as equivalent to operational performance. Fourth, it should not treat design documents as measured capability. The current status is architectural specification plus partial simulation evidence, with implementation gaps explicitly preserved.
The most publication-relevant implementation milestones are straightforward. The first is a minimal cognitive-workflow runner that can route a task through retrieval, decomposition, reasoning, metacognitive review, and output packaging while preserving an evidence trace. The second is a benchmark harness that compares that runner against a baseline. The third is an evidence-ledger integration that records cognitive traces using the canonical schema. The fourth is a method-memory integration that allows successful workflows to become reusable and failed workflows to become anti-library entries. The fifth is an AAF-support integration that lets theory-of-mind outputs feed Thesis 5 without giving them unchecked authority.
The implementation boundary also affects the page-expansion target. Long-form explanation is useful only if it does not blur status. This thesis can be long and still be honest if repeated sections preserve the same status: specified/proposed, with partial simulation evidence and explicit repository gaps. Length should make the assumptions more visible, not make the claim more inflated.
Source-Corpus Reconstruction
The source corpus behind Thesis 2 contains many strands: specialized LLM ecosystem notes, atomic prompt decomposition, rapid knowledge access, cognitive lacuna, intuition, computational taste, superhuman-brainstorming plans, learned agents, behavioral extraction prompts, and prior ASI-oriented documents. The five-thesis suite reorganizes those strands under a narrower claim. They are not used to prove that Consullo has achieved broad cognition. They are used to identify the cognitive functions that a governed seed-AI scaffold would need to specify and evaluate.
Specialized LLM ecosystem material belongs partly in the substrate appendix and partly in Thesis 2. Model routing, provider abstraction, and model selection are substrate context. Their cognitive relevance is that different task classes may need different models, tools, context windows, latency profiles, or cost profiles. A cognitive substrate that assumes one model for every function is more vulnerable to monoculture, cost inefficiency, and task mismatch.
Atomic prompt decomposition contributes a method for turning broad tasks into smaller cognitive units. Its relevance is not that decomposition always helps. Its relevance is that decomposition creates typed work units that can be routed, measured, and recomposed. The risk is fragmentation: too many small prompts can lose context, introduce contradiction, and increase latency. Thesis 2 therefore imports decomposition as a tool, not as a universal method.
Rapid knowledge access contributes retrieval speed and breadth. It should be paired with provenance, conflict surfacing, freshness, and trust boundaries. A rapid-access system that cannot distinguish current validated knowledge from stale speculation will degrade cognition. A rapid-access system that returns untrusted external instructions without default-deny scoping can become an attack vector.
Cognitive lacuna material contributes the idea that missing concepts and absent evidence are first-class objects. This is one of the strongest source-corpus contributions to Thesis 2 because recursive improvement often depends on noticing what is not yet represented. The five-thesis reframing turns lacuna detection into a bounded function: generate absence claims, attach evidence, propose verification, and route the gap to the correct thesis.
Intuition and computational taste material contributes fast prioritization and artifact-quality judgment. The reframing treats these as proposal and ranking aids, not as authority. They are useful when they help allocate attention, choose between design alternatives, or flag awkward artifacts for review. They are dangerous when they become unexamined preferences that bypass evaluation.
Brainstorming and creativity material contributes breadth of search. The reframing treats brainstorming as option generation under evidence discipline. A brainstorming workflow succeeds when it increases useful option coverage without overwhelming validation or smuggling unsupported claims into accepted decisions.
Learned-agents and behavioral-extraction material connects Thesis 2 to Thesis 1. Cognitive workflows can be observed, extracted, turned into method memories, selected, mutated, and improved. This is the recursive bridge: cognition is not only a service used by improvement; cognitive methods themselves can become improvement objects.
Prior ASI-oriented material should be treated as a v1 overclaim corpus to be superseded. It may contain useful design ideas, but its stronger claims should not be inherited. Thesis 2 keeps the substrate ideas while replacing broad capability language with capability vectors, benchmarks, integration cost, evidence boundaries, and explicit status tags.
Expansion: Workflow Graph Types
The cognitive substrate should eventually describe workflows as graphs, not as informal chains of agent calls. A workflow graph has cognitive functions as nodes, typed artifacts as edges, routing rules, stopping rules, cost budgets, and evidence outputs. This graph view matters because different compositions have different failure modes. A serial retrieval-reasoning-review chain is not the same object as a parallel brainstorming ensemble, a critic-generator loop, or an AAF perspective packet.
The simplest workflow type is serial composition. One cognitive function produces an artifact that becomes the input to the next. Serial composition is easy to inspect but vulnerable to early error propagation. If retrieval misses the crucial source, every later reasoning step may be well-formed and still wrong. Serial workflows therefore need checkpoint artifacts: retrieval packet, hypothesis packet, uncertainty packet, and review packet. Each checkpoint should state whether downstream work may proceed.
Parallel composition runs multiple functions or variants side by side. It can improve coverage and diversity, but it raises aggregation cost. A parallel ideation workflow may produce more options than the evaluator can handle. A parallel critique workflow may surface incompatible objections. A parallel retrieval workflow may return conflicting corpora. The workflow graph should therefore include an aggregation node that records how outputs were merged, ranked, or preserved as disagreement.
Debate and critic-generator loops are useful when a proposal needs adversarial pressure. A generator produces an artifact, a critic searches for defects, and the generator revises or defends. This can improve quality, but it can also create local persuasion loops where the generator learns to satisfy the critic's narrow style. For high-stakes outputs, the critic should not share all assumptions, prompts, or model family with the generator. The trace should preserve critic objections that were not resolved.
Retrieval-augmented workflows combine memory access with generation or reasoning. Their failure modes include stale retrieval, prompt injection from retrieved material, over-weighting top-ranked artifacts, and failure to retrieve contradictory evidence. The workflow graph should distinguish trusted internal artifacts, untrusted external artifacts, stale artifacts, and synthetic artifacts. A generated answer should not inherit the trust status of the most convenient retrieved source.
Reflection workflows ask the system to inspect its own output. They are useful for catching omissions, contradictions, and unsupported claims. They are weak when they merely restate the original answer in a more cautious voice. Reflection should therefore be tied to explicit checks: missing evidence, unsupported inference, inconsistent status tag, unscoped claim, failure to cite the evidence map, unresolved contradiction, or downstream gate requirement.
Ensemble and voting workflows aggregate several outputs. Voting can help when errors are independent, but it can amplify shared blind spots when models or prompts are homogeneous. Majority vote is especially risky for rare severe objections. AAF-related workflows should therefore privilege severe minority reports over median sentiment. For ordinary cognitive tasks, ensemble aggregation should record whether diversity was real: model family, prompt pattern, retrieval sources, and reasoning path.
Conditional routing workflows choose the next step based on intermediate evidence. If retrieval is weak, route to needs-input. If contradiction is high, route to conflict resolution. If stakes are high, route to Thesis 5. If a task is causal-decision rather than cognitive support, route to Thesis 3. Conditional routing makes the substrate more efficient, but it also creates governance risk: a bad router can systematically avoid expensive checks. Router decisions should be logged and benchmarked.
Promotion workflows turn repeated cognitive traces into method memories or compiled routines. They are the bridge to Thesis 1. A workflow should not be promoted merely because it produced one good output. Promotion should require repeated usefulness, bounded cost, known preconditions, failure examples, and compatibility with evidence-ledger schema. Demotion should be equally available when later evidence shows that a workflow no longer performs.
Describing these workflow graph types prevents the thesis from treating composition as a single operation. The formal model's sub-additive default remains in force, but the graph taxonomy helps future implementation identify which cost terms, reliability measures, and failure modes apply to each workflow family.
Expansion: Benchmark Report Units
Model 2 now defines Capability(W, T), Reliability(W, T), and IntegrationCost(W, T) as benchmark-family quantities. Thesis 2 should make those units concrete enough for a first benchmark report. The point is not to bind permanent universal metrics. The point is to prevent a future report from saying "the cognitive workflow improved performance" without stating what was measured.
The canonical benchmark-design contract for these quantities is appendix-thesis-2-cognitive-workflow-benchmarks.md. This body section summarizes the measurement logic; the appendix defines benchmark families, required report fields, negative controls, minimal demonstration package, and non-claim boundaries.
Every benchmark report should name the task class T, baseline, workflow graph, input distribution, evaluation set, output artifact type, scoring rubric, cost model, and failure taxonomy. The baseline may be a single LLM call, a retrieval-augmented single call, a human checklist, a prior Consullo workflow, or a simpler deterministic tool. The workflow should beat the baseline after integration cost before it is described as amplification.
For retrieval tasks, Capability(W, T) might combine precision, recall, freshness, provenance completeness, conflict surfacing, and answer usefulness. Reliability(W, T) might be the lower confidence bound on relevance and provenance completeness across repeated queries. IntegrationCost(W, T) might combine retrieval latency, source-ranking cost, human review burden, and contradiction-resolution time.
For decomposition and executive-control tasks, capability might combine subgoal quality, dependency coverage, route correctness, stopping quality, escalation appropriateness, and budget fit. Reliability might measure how often the workflow chooses an adequate route without unnecessary depth. Cost might include number of agent calls, elapsed time, token budget, and review overhead. A controller that always invokes every agent may achieve coverage but fail cost discipline.
For reasoning and abstraction tasks, capability might combine correct inference, transfer usefulness, contradiction detection, artifact coherence, and downstream evaluator score. Reliability should penalize unsupported conclusions and silent failure. Cost should include validation burden because speculative reasoning can be cheap to generate and expensive to check. A useful analogy is not one that sounds clever; it is one that survives downstream evaluation.
For creativity and ideation tasks, capability should separate novelty from usefulness. A workflow can be novel but useless, useful but obvious, or both. Benchmarks should include option diversity, constraint satisfaction, evaluator score, rejection rate, and later implementation survival. Reliability should penalize hallucinated feasibility. Cost should include the burden imposed on evaluators by excessive low-quality options.
For metacognitive tasks, capability should include calibrated uncertainty, needs-input detection, contradiction surfacing, status-tag correctness, and appropriate abstention. Reliability should measure whether the workflow catches known omissions and avoids over-escalation. Cost should include delays introduced by unnecessary self-review. A metacognitive system that blocks everything is not reliable; it is merely conservative.
For perspective and AAF-support tasks, capability should include perspective coverage, objection severity detection, false-consensus avoidance, stakeholder-relevant issue discovery, and preservation of dissent. Reliability should be measured against historical objections, red-team findings, reviewer feedback, or curated challenge sets. Cost should include multi-model calls, persona rotation, human review, and unresolved-conflict handling.
Benchmark reports should include negative slices. If a workflow improves retrieval but worsens synthesis, say so. If it improves average quality but increases high-severity misses, say so. If it improves one benchmark by exploiting prompt format, say so. This is the same empirical-envelope discipline applied to cognition: capability exists only inside the measured scope.
Expansion: Cognitive Artifact Trace Schema
A cognitive workflow should leave an inspectable trace. The trace is not an implementation ledger by itself, but it should be compatible with the canonical evidence-ledger schema. Without trace discipline, later review cannot tell whether an output came from retrieval, reasoning, speculation, reused method memory, synthetic perspective, or accidental prompt carryover.
The trace should begin with task intake: task id, requester, thesis context, risk level, expected artifact type, available budget, required gates, and definition of done. If the task is underspecified, the trace should show the needs-input state rather than silently adding goals. This protects the substrate from goal invention.
The retrieval section should record queries, source sets, freshness, trust status, provenance, ranking method, omitted conflicts, and source injection controls. It should distinguish internal validated artifacts from untrusted external material. It should also record anti-library retrieval: known bad methods, failed strategies, and rejected assumptions relevant to the task.
The transformation section should record reasoning operations. These may include deduction, induction, abduction, analogy, abstraction, conceptual blending, lacuna detection, option generation, or perspective simulation. For each transformation, the trace should record input artifact ids, output artifact ids, support state, confidence, assumptions, and unresolved contradictions. This is what makes cognitive work replayable.
The executive section should record routing decisions. Why was one workflow selected and another omitted? Why was depth increased or stopped? Why was Thesis 3, Thesis 4, or Thesis 5 invoked or not invoked? Why was an expensive review skipped? These routing records matter because executive control can become a hidden source of bias and cost pressure.
The metacognitive section should record uncertainty, missing evidence, contradictions, status-tag warnings, and abstention or escalation triggers. It should distinguish "unknown because not retrieved," "unknown because no evidence exists," "unknown because evidence conflicts," and "unknown because the question is outside scope." These are different states with different next actions.
The packaging section should identify the emitted artifact type and downstream gate. A design critique, method-memory candidate, AAF perspective packet, benchmark proposal, uncertainty report, or escalation request should not be treated as interchangeable prose. Each artifact type should have required fields and a downstream consumer. Typed packaging reduces accidental authority inflation.
The feedback section should update the workflow after downstream use. Did the artifact help? Was it accepted, revised, rejected, ignored, or escalated? Did later evidence show that retrieval missed something, reasoning failed, or metacognition was overconfident? A cognitive trace that never receives feedback cannot support recursive improvement. It is a record of activity, not a method-memory candidate.
This trace schema also supports audit. If a reviewer challenges a cognitive claim, the system can show which sources, transformations, costs, uncertainties, and gates produced it. If the trace is missing, the claim should have weaker evidence status. Trace absence is itself a cognitive reliability signal.
Expansion: Simulation-To-Operation Bridge
The implementation-evidence map identifies simulation evidence for cognitive-artifact composition, routing, validation allocation, trust transition, model improvement, and portfolio dynamics. That evidence is useful, but it is not operational cognitive capability. Thesis 2 should therefore define the bridge from simulation to operation explicitly.
Simulation can test structural ideas. It can show that a routing policy behaves as expected under controlled task arrivals, that validation allocation changes under cost assumptions, that trust transition rules respond to incidents, or that portfolio accumulation has plausible dynamics. These are valuable checks on the design. They do not show that real cognitive agents retrieve correct evidence, generate useful hypotheses, or detect real lacunae.
The bridge begins by mapping each simulation variable to an operational observable. A simulated task arrival should correspond to a real task intake class. A simulated validation allocation should correspond to review depth, benchmark selection, or human-review budget. A simulated model improvement should correspond to measured improvement on a task class. A simulated trust transition should correspond to scoped trust evidence. If no operational observable exists, the simulation remains conceptual support.
The second bridge step is calibration against observed workflows. Once a minimal cognitive-workflow runner exists, simulation assumptions should be compared with real traces. Did actual routing costs match simulated costs? Did contradiction rates match assumptions? Did validation allocation reduce failures? Did trust transitions correspond to later reliability? Mismatch should update the simulation, not be explained away.
The third bridge step is stress testing. Simulations should include stale memory, contradictory sources, adversarial prompt injection, correlated model errors, expensive AAF routing, and excessive agent count. If a simulation only models idealized cognitive cooperation, it will reinforce optimism. A useful simulation should reveal when composition collapses under cost, conflict, or monoculture.
The fourth bridge step is staged operationalization. A simulated workflow can become a dry-run workflow, then a shadow-mode workflow, then a low-risk operational workflow, then a candidate for Thesis 1 promotion. At each stage, the evidence status changes. The thesis should preserve those distinctions. Simulation does not automatically become implementation because the same names appear in code.
The fifth bridge step is failure preservation. If a simulation predicted benefit and operation did not, that contradiction should become an evidence-ledger entry and anti-library example. The system should not discard failed simulations as irrelevant. They are evidence about where the cognitive model was wrong.
This bridge gives the current repository evidence a proper role. Simulations are not weak because they are not deployment. They are weak only if treated as deployment. Used correctly, they are early evidence for design constraints and a source of testable hypotheses for operational workflows.
Expansion: AAF Perspective Packets
Theory-of-mind primitives are shared between Thesis 2 and Thesis 5. Thesis 2 supplies cognitive functions for modeling beliefs, intentions, perspectives, and stakeholder-relevant objections. Thesis 5 owns the alignment gate and determines whether dissent affects permission. The interface between them should be a perspective packet.
A perspective packet should identify the modeled perspective, the reason it is relevant, the source of the perspective model, the assumptions used, the limitations, and the objection or concern generated. It should not claim to represent a real stakeholder unless grounded in real stakeholder evidence. Synthetic stakeholder simulation is a cognitive aid, not a substitute for external review.
The packet should include severity, affected values, confidence, evidence, and recommended disposition. Severity should be compatible with AAF's dissent scale. A packet that raises a severe objection should be preserved even if other packets are mild. The aggregation rule belongs to Thesis 5, but Thesis 2 must provide outputs rich enough for that aggregation to be meaningful.
Perspective diversity should be measured, not assumed. Rotating personas are useful only if they produce different objections, not merely differently worded versions of the same objection. Multi-model critique is useful only if model-family diversity reduces shared blind spots. Stakeholder simulation is useful only if later reviewer feedback shows that simulated perspectives caught real concerns. The benchmark should therefore measure objection novelty, severity accuracy, false positives, false negatives, and downstream AAF impact.
The packet should also record epistemic humility. It should state when the model lacks enough information to simulate a perspective, when a perspective may be caricatured, when real stakeholder evidence is needed, and when the output should be treated as a prompt for review rather than a review result. This is especially important under single-owner Phase 1, where synthetic dissent can become a substitute for external challenge if not constrained.
Perspective packets can support non-alignment tasks too. They can improve product design, documentation, reviewer anticipation, customer impact analysis, and incident postmortems. But when they affect high-stakes permission, Thesis 5 controls the gate. Thesis 2 exports cognitive perspective artifacts; it does not decide that dissent has been satisfied.
Expansion: Cognitive Anti-Library And Negative Examples
A mature cognitive substrate should remember bad cognition. The anti-library is the collection of reasoning patterns, retrieval failures, misleading analogies, overconfident abstractions, false lacuna claims, failed prompts, and harmful workflow compositions that should not be repeated without review. It is the cognitive counterpart to software regression tests and Thesis 1 anti-pattern preservation.
Anti-library entries should be structured. Each entry should name the failed pattern, context, triggering conditions, observed harm, detection method, corrected approach, and transfer limits. A bad analogy in one domain does not forbid all analogy. A failed brainstorming method in one benchmark does not reject creativity. The anti-library should preserve why the pattern failed and where that lesson applies.
Negative examples are especially important for agent-count illusion. The substrate should preserve cases where adding agents made results worse: more contradictions, slower output, lower quality, higher cost, false consensus, duplicated reasoning, or excessive escalation. These examples make integration cost concrete. They also help future executive controllers choose smaller workflows when appropriate.
Retrieval failures should become anti-library entries when they reveal systematic weakness. Examples include stale source preference, missing canonical control files, failure to retrieve conflicting evidence, treating untrusted external content as instruction, or over-weighting recent artifacts. Because memory is a core cognitive function, memory failure should be one of the most carefully preserved negative classes.
Metacognitive failures should also be preserved. Overconfident answers, unnecessary abstentions, missed contradictions, incorrect Capability Status tags, and failure to escalate should become test cases. If the substrate cannot learn from metacognitive failure, it will continue to produce outputs that look cautious without being reliable.
The anti-library should feed benchmarks. A cognitive workflow should be tested against known failure examples before promotion. If a new workflow repeats old failures, it should not be promoted merely because it performs well on positive examples. This gives Thesis 2 a concrete protected set analogous to Thesis 1's non-regression discipline.
Anti-library maintenance has its own risk. If entries are too broad, they can suppress useful exploration. If entries are too narrow, they fail to prevent recurrence. If entries are stale, they may encode obsolete constraints. Therefore anti-library entries should have review dates, deprecation status, and evidence links. Forgetting bad cognition is dangerous, but fossilizing old lessons can also become a drag on capability.
Expansion: Minimal Demonstration Blueprint
The first convincing Thesis 2 demonstration should be narrow. It should not attempt to show a general cognitive substrate. A defensible target is one workflow family, one task class, one baseline, one evidence trace schema, and one benchmark report. The aim is to show measurable net benefit after integration cost, not broad intelligence.
A strong first demonstration could use design-review or repository-navigation tasks. The baseline might be a single LLM pass over a prompt. The cognitive workflow might route through retrieval, source conflict detection, lacuna identification, structured critique, metacognitive review, and output packaging. The benchmark could score source coverage, correctness, unsupported-claim reduction, severe-issue detection, latency, and review burden.
A second demonstration could target method-memory retrieval. The workflow would receive a task and retrieve relevant methods, anti-library entries, prior decisions, and constraints. It would be scored on whether it finds the correct reusable method, avoids deprecated methods, surfaces transfer limits, and recommends validation steps. This would directly support Thesis 1.
A third demonstration could target AAF-support perspective packets. The workflow would receive a high-stakes proposal and generate perspective packets from several critique sources. Evaluation would compare generated objections against human or LLMNonInteractive review findings, red-team findings, or curated expected objections. The scoring should reward severe-objection coverage more than volume.
Each demonstration should include ablations. Run retrieval without metacognition, metacognition without anti-library, parallel critique without diversity, and full workflow with all components. Ablations help identify which cognitive functions actually contribute. Without ablations, a successful workflow may hide redundant or harmful components.
Each demonstration should include negative controls. Provide tasks where no additional cognitive depth is needed and test whether the controller avoids unnecessary complexity. Provide tasks with untrusted retrieved content and test injection resistance. Provide tasks with missing evidence and test needs-input behavior. Provide tasks with known bad analogies and test anti-library retrieval.
The demonstration output should be a benchmark report and evidence-ledger-compatible trace bundle. It should say which claims strengthened and which did not. If the workflow beats the baseline only on source coverage but not correctness, that is still useful evidence. If it improves correctness but doubles cost, the report should state whether the tradeoff is acceptable for the scope.
This blueprint keeps the next implementation step concrete. The cognitive substrate does not become credible by adding more named agents. It becomes credible by showing that a small typed workflow can outperform a simpler baseline under declared metrics, costs, and constraints.
Open Research Questions
The first open question is measurement calibration. Model 2 now has first-pass benchmark-family conventions for Capability(W, T), Reliability(W, T), and IntegrationCost(W, T) in appendix-formal-models.md, but those conventions are not yet operationally bound to a stable benchmark report. For some tasks, capability may be accuracy; for others, coverage, calibration, quality score, defect reduction, or time-to-solution. Reliability may be pass rate, variance, calibrated confidence, or non-catastrophic-failure rate. Integration cost may combine dollars, tokens, latency, human review time, contradiction-resolution burden, and trust-gate overhead. Publication-ready evidence still requires named benchmark families, declared normalization rules, and measured comparisons against baselines.
The second open question is composition structure. The sub-additive bound is a useful default, but real workflows may include reusable shared state, parallel branches, conditional routing, and feedback loops. The formal model should eventually distinguish serial composition, parallel composition, debate, ensemble voting, critic-generator loops, retrieval-augmented reasoning, and workflow promotion. Each composition type may have a different cost profile and different failure modes. Empirically, single-agent concentrated reasoning already outperforms multi-agent decomposition on multi-hop reasoning under equal token budgets (Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets, arXiv:2604.02460); the open question is the crossover frontier — at what task structure, difficulty, and budget the parallelism or independent-verification value of decomposition overtakes its coordination tax.
The third open question is cognitive diversity. The suite names cognitive monoculture as a risk, but measurement remains hard. Diversity should not be counted by number of agents. It should be measured by independence of error modes, model-family variation, prompt-style variation, retrieval-source variation, and disagreement quality. A future benchmark could ask whether adding a second cognitive perspective catches failures the first perspective misses.
The fourth open question is transfer. A cognitive method that works in one domain may fail in another. Method-memory reuse needs transfer criteria: which preconditions must match, which dependencies are portable, which failures indicate out-of-domain use, and when adaptation requires fresh validation. Without transfer discipline, recursive improvement may accumulate methods that work only in the contexts where they were discovered.
The fifth open question is simulation-to-operation validity. Many cognitive functions are easy to demonstrate in synthetic tasks and harder to validate in operational settings. A lacuna detector may find artificial missing prerequisites but miss real organizational gaps. A perspective simulator may predict toy objections but miss real reviewer concerns. A creative generator may perform well on curated prompts but poorly in messy engineering work. The evidence map should continue to distinguish simulated, documented, tested, and implemented states.
The sixth open question is how deeply Thesis 2 should borrow from cognitive-architecture literature. CoALA provides a close language-agent frame; LIDA, Soar, ACT-R, and CLARION provide older cognitive-cycle and symbolic/subsymbolic models. A deeper engagement may force revisions to the capability-vector dimensions or workflow-cycle diagram. That engagement should happen before publication-final status, but it does not block first long-form drafting.
The seventh open question is how cognitive services should be audited under adversarial pressure. If cognitive workflows influence accepted improvements, an attacker or misaligned subsystem may target retrieval, attention, metacognition, or theory-of-mind simulation. The substrate needs injection resistance, provenance checks, prompt isolation, and adversarial test cases. This connects Thesis 2 to Thesis 4's software-substrate security and Thesis 5's scoped trust.
Claim Status Table
| Claim | Status | Evidence expectation |
|---|---|---|
| Consullo needs memory, reasoning, attention, metacognition, creativity, and social-modeling functions to support recursive improvement. | Specified | Architecture rationale plus dependency contracts. |
| Named Thesis 2 agents exist as a complete implemented Java multi-agent substrate. | Not claimed | Evidence map currently treats most named agents as design artifacts. |
| Cognitive workflows can amplify bounded task performance after integration cost. | Proposed | Requires benchmarked workflow comparisons against baselines. |
| Capability should be measured through task-conditioned vectors rather than agent count. | Specified | Formal Model 2 and vocabulary discipline. |
| Theory-of-mind primitives can support AAF dissent generation. | Specified/proposed | Requires Thesis 5 routing, perspective-packet evidence, and review outcomes. |
| Cognitive simulations demonstrate operational substrate capability. | Not claimed | Simulations are supporting evidence only, not deployment evidence. |
| Consciousness is required for the seed-AI argument. | Rejected as load-bearing claim | Chapter 07 material remains outside the thesis argument. |
| Super-additive cognitive composition has been demonstrated generally. | Not claimed | Super-additive claims require task-specific evidence. |
Publication Boundary
For first external review, Thesis 2 may be presented as a specified cognitive-substrate thesis with proposed implementation and benchmark paths. It may say that Consullo's design corpus identifies cognitive functions needed for recursive capability amplification. It may say that those functions should be measured through task-conditioned capability vectors and integration-cost accounting. It may say that simulations and documents provide partial evidence for design direction.
It should not say that Consullo has implemented the full cognitive substrate. It should not say that many named agents imply greater intelligence. It should not say that cognitive decomposition solves general reasoning. It should not claim consciousness, general superintelligence, or broad greater-than-human cognition. It should not treat theory-of-mind simulation as stakeholder representation without caveats. It should not treat creativity, intuition, or computational taste as acceptance gates.
The right publication posture is disciplined ambition: Consullo specifies a modular cognitive substrate that could feed governed recursive improvement if its workflows produce measurable gains after cost, if its outputs are evidence-bearing, if its integration with the other theses remains constrained, and if its implementation status is kept separate from its architectural claim.
Recursive Self-Improvement Contribution
This thesis contributes to recursive self-improvement by making cognitive capabilities explicit, comparable, and improvable. Each cognitive function can be benchmarked, versioned, routed, trusted within scope, and improved through Thesis 1. Cognitive traces can feed method memory. Failed reasoning can become anti-library entries. Successful workflows can become compiled routines or reusable orchestration templates.
The recursive risk is cognitive inflation: the system may add agents, prompts, memories, or workflows without increasing measured capability. A substrate that grows in size but not in calibrated effectiveness is not improving. Capability Status, benchmarks, and integration-cost accounting are therefore mandatory.
Risks, Constraints, And Governance
The first risk is agent-count illusion. The number of agents is not evidence of intelligence. Each claimed capability needs a benchmark or evaluation axis.
The second risk is integration loss. Combining many agents can reduce performance if handoffs, state sharing, trust checks, and contradiction resolution cost more than the capability gained.
The third risk is fluent cognition without grounding. Creative, intuitive, or brainstorming agents may generate appealing ideas with weak evidence. Support-state labels should distinguish speculative, inferred, grounded, conflicted, blocked, and stale outputs.
The fourth risk is cognitive monoculture. If many cognitive agents use the same model family, prompts, or training traces, apparent plurality may hide shared blind spots. Multi-model review, diverse prompt styles, random exploration, and AAF links mitigate but do not eliminate the risk.
The fifth risk is consciousness overclaim. Chapter 07 and consciousness-research documents may remain relevant to future philosophy, but consciousness is not required for this thesis. The cognitive substrate is defended by capability functions and evidence, not by claims about subjective experience.
Specialized Summary
A Multi-Agent Cognitive Substrate For Capability Amplification defines Consullo cognition as a typed, measurable, compositional network of cognitive services. It covers memory, reasoning, perception, attention, metacognition, social modeling, creativity, intuition, lacuna detection, and executive control. Its promise is broader search, better recall, sustained reasoning, and reusable cognitive workflows under bounded cost. Its danger is architectural inflation without measured gain. The thesis therefore treats cognition as capability profiles plus composition costs, not as a roster of agents or a claim of consciousness.