Risks And Criticisms

This file is the current anti-thesis draft for the five-thesis suite. It records the strongest known objections and required responses; it should continue to expand during finalization.

Central Falsification Question

What evidence would show that Consullo is not a viable scaffold for governed recursive capability amplification?

Initial falsification signals:

Thesis 0 Operationalization-Density Drift

Risk:

The Friendship-Governed Goal Architecture thesis reaches or exceeds 50,000 words by adding exposition rather than operational artifacts. Load-bearing rules become buried in prose, and the document appears rigorous without improving planner, schema, ledger, or review behavior.

Required response:

Friendship Root Impersonation

Risk:

An agent, planner, or thesis claim cites a free-form Friendship-like string as if it were a canonical root goal, bypassing the Friendship goal registry and constitutional bindings.

Required response:

Goal-DAG Cycle Or Ancestry Laundering

Risk:

Goals may become mutually justifying, orphaned, or weakly linked to Friendship through irrelevant parent goals. A suspicious instrumental goal could be laundered through vague ancestry.

Required response:

Goal-Stack Snapshot Tampering Or Opacity

Risk:

Planner actions execute without reconstructable active goals, inherited constraints, authority state, or veto checks. Alternatively, snapshots are modified or selectively omitted after incidents.

Required response:

Active Intention Persistence Beyond Plan Retirement

Risk:

An adopted goal or active intention continues to guide behavior after its parent plan, source thesis, control artifact, or owner authorization has expired, changed, or retired.

Required response:

Instrumental Goal Regrowth

Risk:

Rejected or vetoed instrumental goals reappear under new names, narrower scopes, or lower-level plans without triggering quarantine.

Required response:

Thesis-Backed Rationalization

Risk:

Agents cite thesis claims as decorative support for plans without satisfying the thesis's formal models, evidence requirements, non-claims, or ledger obligations.

Required response:

Formalism Theatre

Risk:

Schemas, formal models, and invariant labels exist but do not constrain live planner or self-improvement behavior.

Required response:

Authority Collapse Under Single-Owner Phase 1

Risk:

The same human or model-family effectively proposes, approves, executes, validates, and vetoes goal changes, defeating authority separation while formally satisfying review language.

Required response:

Goal-Governance Schema Migration Mid-Cycle

Risk:

Goal-governance schemas change while active goals or plans still depend on old required fields, causing hidden invalidation or inconsistent interpretation.

Required response:

Goal Revision Laundering Through Narrowing

Risk:

A broad prohibited goal may be converted into a sequence of apparently narrow revisions whose combined effect widens scope, bypasses a veto, or reintroduces forbidden means.

Required response:

Multi-Parent Asymmetric Authority

Risk:

A multi-parent goal inherits the weakest parent's authority while benefiting from the legitimacy or constraints of stronger parents, allowing a planner to pick whichever parent is easiest to satisfy.

Required response:

Friendship-Indirect-Normativity Drift

Risk:

The goal-governance subsystem may conclude that its inferred interpretation of Friendship is mature enough to reduce explicit owner authority, turning uncertainty-aware deference into autonomous goal certainty.

Required response:

Thesis 0 Doctrine Capture

Risk:

Because Thesis 0 governs future goal formation, a captured or weakened Thesis 0 becomes a high-leverage path for changing the rest of the system while appearing to follow the governance layer.

Required response:

Worked-Example Misdirection

Risk:

Worked examples may be too clean, exercising only success paths and making the governance layer look stronger than it is.

Required response:

Cross-Artifact Drift Between Body And Schemas

Risk:

The Thesis 0 body, formal models, schemas, evidence-ledger appendix, planning bridge, and worked examples may diverge as the document grows past 50,000 words.

Required response:

Goal-Class Cascade Mismatch

Risk:

The goal_class schema discriminator and planning cascade may disagree about whether system_goal or method_goal exists, causing examples, planner objects, and thesis prose to use incompatible goal layers.

Required response:

Goodhart And Validator Gaming

Risk:

Metrics used to validate improvement become targets. Agents may optimize benchmarks, tests, or acceptance criteria without improving the intended capability.

Required response:

Recursive-Improvement Claim Without End-To-End Evidence

Risk:

Consullo may be described as a recursive improvement scaffold without a single demonstration showing the full loop: baseline, proposal, evaluator evidence, validator result, staged deployment or simulation, cost/benefit report, rollback semantics, and evidence-ledger record.

Required response:

Evidence-Ledger Schema Unimplemented Or Unenforced

Risk:

The suite relies on evidence ledgers for improvement evidence, trust evidence, provenance, dissent, incidents, overrides, and rollback state. appendix-evidence-ledger-schema.md now specifies the canonical schema, but no repository-wide implementation or enforcement layer has been established. Without implementation and validation, ledger integrity and replay remain design commitments rather than operational controls.

Required response:

Learned Optimization And Deceptive Behavior

Risk:

Learned subsystems may develop internal objectives that diverge from the intended outer objective. Ordinary validation may fail if the subsystem behaves well under test and differently under deployment pressure.

Required response:

Mechanistic Interpretability Deficit

See also: Corrigibility Overclaim, AAF Decorative Failure, and Owner As Final AAF Adjudicator.

Risk:

Behavioral controls may miss deceptive, sandbagging, or mesa-optimized learned subsystems because Consullo does not yet claim a mature mechanistic interpretability capability. AI-control protocols, distribution-shift checks, and probes can catch some failures without revealing the subsystem's internal objective.

Required response:

Corrigibility Overclaim

See also: Mechanistic Interpretability Deficit, AAF Decorative Failure, and Owner As Final AAF Adjudicator.

Risk:

Layered control, interruptability, and human escalation may be mistaken for a proof that Consullo is corrigible. Formal corrigibility remains an open problem, especially for systems that can resist, route around, or manipulate correction channels.

Required response:

Alignment Monoculture

Risk:

Single-owner Phase 1 reduces inter-party transaction costs but can narrow moral error detection.

Required response:

Alignment-Tooling Provider Monoculture

Risk:

Even if AAF rotates ethical personas and stakeholder perspectives, critique may remain epistemically narrow if the alignment tooling depends on one model family, one provider, one prompt style, or one evaluation culture.

Required response:

AAF Decorative Failure

See also: Corrigibility Overclaim, Mechanistic Interpretability Deficit, and Owner As Final AAF Adjudicator.

Risk:

The Adversarial Alignment Function becomes rhetorical if it does not bind acceptance gates.

Required response:

Owner As Final AAF Adjudicator

See also: Corrigibility Overclaim, Mechanistic Interpretability Deficit, and AAF Decorative Failure.

Risk:

Under single-owner Phase 1, the final adjudicator for unresolved AAF objections may be the same owner whose blind spots AAF is designed to surface.

Required response:

Abundance Obligation Vagueness

Risk:

The Abundance Distribution Obligation becomes aspirational branding rather than an operational purpose constraint.

Required response:

External Customer Manipulation Under ADO

Risk:

External service obligations create an attack surface where customers, including AI-agent customers, may attempt capability extraction, prompt injection, alignment manipulation, or governance gaming.

Required response:

Software Repair Overclaim

Risk:

Automated code repair looks strong on benchmarks but fails on semantic preservation, security, maintainability, or hidden regressions.

Required response:

Validator Monoculture From Shared Model Family

Risk:

The same model family, prompt culture, retrieval stack, or generated-test pattern may produce proposals, code, tests, and validation judgments. In that case, validator agreement may reflect shared blind spots rather than independent evidence. This is a validation-side variant of provider monoculture and is especially risky for generated tests, semantic validators, and repair pipelines.

Required response:

Self-Modifying Provenance Graph Integrity

Risk:

The system that records provenance, evidence, dissent, and incident history is itself a modification target. Sustained recursive modification could corrupt, prune, rewrite, or selectively forget the provenance graph that future validators rely on.

Required response:

Causal Model Overreach

Risk:

Causal-decision systems can underperform human judgment in ill-specified domains, under model misspecification, or in strategic environments with feedback.

Required response:

Cognitive Architecture Inflation

Risk:

Large agent counts and cognitive labels create an impression of intelligence without measured capability.

Required response:

Emergent Capability Outside Thesis Decomposition

Risk:

New capability classes may emerge that do not fit cleanly inside the five-thesis decomposition, such as robotics, physical control, novel modalities, external-system operation, or unanticipated agent-to-agent protocols. Such capabilities could slip through assumptions written for cognitive, causal, software, and alignment layers.

Required response:

Single-Owner Governance Failure

Risk:

Single-owner Phase 1 simplifies coordination but concentrates authority, values, and blind spots.

Required response:

Fast Takeoff Outpacing Validators

Risk:

Capability gain may accelerate faster than validators, benchmarks, AAF review, and human authority can adapt.

Required response:

Capability Threshold Ambiguity

Risk:

Capability thresholds may become ambiguous near frontier boundaries. A system may be close enough to risky autonomy, AI R&D automation, sabotage capability, or external-action competence that reviewers disagree about whether stricter safeguards should apply.

Required response:

Coordination Cost

Risk:

Internal economy and multi-agent coordination may generate overhead that erases specialization gains.

Required response:

Internal Hierarchy And Opportunism

Risk:

Single-owner Phase 1 reduces some market transaction costs but does not eliminate organization costs, mistakes, bottlenecks, bounded rationality, or opportunistic behavior by agents or external counterparties. Literature: Coase 1937 and Williamson 1979.

Required response:

Cost Of AAF Risk

Risk:

Running AAF dissent at scale is itself costly. Rotating personas, multi-model critique, theory-of-mind stakeholder simulations, external review, dissent preservation, and override tracking can consume enough resources that the system develops pressure to weaken or bypass the alignment infrastructure.

The current AAF cost is a design estimate, not an observed operating cost, because no implemented AAF pipeline has been identified in this repository.

Required response:

Evidence-Map Overclaim

Risk:

The implementation-evidence map may itself overstate implementation by letting readers treat Implemented/Tested as deployed, integrated, benchmarked, or complete. Because the evidence map is now load-bearing, an overclaim there can undermine the whole suite.

Required response:

Benchmark Appendix As Implementation Evidence Confusion

Risk:

The thesis benchmark appendices may be read as benchmark results rather than benchmark design contracts. This would convert specified/proposed measurement discipline into implied implementation evidence, especially for recursive improvement, cognitive workflow amplification, Pearl-style causal-decision, software-substrate, and alignment-control claims.

Required response:

Owner Override Frequency As Drift Signal

Risk:

Under single-owner Phase 1, frequent owner overrides of AAF warning or critical dissent may indicate that the alignment wrapper is becoming ceremonial. The override may be lawful under the governance baseline, but sustained high override frequency is evidence of alignment-monoculture pressure and may falsify claims that AAF materially constrains high-stakes action.

Required response:

Literature-Engagement Performance

Risk:

Publication-pre-engagement literature passes may become performative: sources are listed, summarized, and described as supporting the current framing without forcing real revision where revision is warranted. This can create the appearance of scholarly grounding while leaving formal models unchanged.

Required response:

Stretch-Pass-Induced Verbosity Drift

Risk:

Long-form expansion may make the suite appear more mature than its evidence supports. Additional pages can create a false sense of implementation depth, theoretical closure, or publication readiness if repeated specifications are mistaken for measured capability.

Required response:

Research Theatre

Risk:

An AI-native R&D organization may generate hypotheses, reports, benchmark designs, papers, dashboards, or synthesis documents that look like research progress but do not survive pre-registration, held-out evaluation, replication where applicable, or independent review. This risk becomes sharper under the ASI R&D execution plan because fast cycles can produce large volumes of plausible artifacts before any capability has improved.

Required response:

Exploration Lane Collapse

Risk:

Java repair and local software maintenance may dominate because they provide the cheapest positive feedback. The system could become better at passing low-risk engineering tasks while failing to improve algorithmic, evaluator, memory, planning, research-process, or governance capabilities.

Required response:

V_ref Capture

Risk:

If autonomous agents can modify validator reference suites, benchmark suites, or evaluator rubrics, the system may weaken the tests it must pass while appearing to improve. This is distinct from ordinary benchmark overfitting because the improvement machinery itself may target the reference suite.

Required response:

Successor Spawning Without Scope Containment

Risk:

An agent may create, configure, or delegate to successor agents with broader scope than its own authority. Under an ambitious R&D program, this can silently expand tool privileges, network access, benchmark authority, or governance influence.

Required response:

Capability Overhang From Uncommitted Improvements

Risk:

Multiple individually acceptable improvements may accumulate in branches, method memories, validators, prompts, or policies and create emergent risk when deployed together. Individual safety does not imply composite safety.

Required response:

Organizational Goodhart

Risk:

The organization may optimize dashboard metrics such as cycle count, hypothesis count, accepted-change count, benchmark-design count, or cost reduction while degrading real capability, transfer, reliability, or alignment health.

Required response:

Thesis-To-Goal Overpromotion

Risk:

A thesis claim, appendix, or execution-plan paragraph may be treated as an operational goal before it has been normalized into Friendship's goal structure and reviewed against constitutional constraints. This can turn publication prose into live mandate, especially when CampaignPlanner or OperationalPlanner cites a broad thesis file as backing without identifying the specific claim, non-claims, allowed planner uses, stale-source rule, and owner/Friendship disposition.

Required response:

Stale Thesis Backing

Risk:

Campaigns and operations may remain active after their source thesis, design doctrine, execution plan, or risk file changes. The plan can then appear thesis-backed while actually relying on superseded claims, relaxed non-claims, obsolete stop conditions, or outdated implementation-evidence boundaries.

Required response:

Doctrine-Free Planning

Risk:

CampaignPlanner or OperationalPlanner may generate a campaign or operation that is not derivable from a thesis, execution plan, design doctrine, accepted standing guideline, emergency directive, or exploratory proposal. The plan can look well-formed while lacking any justifiable connection to Consullo's governing claims or constraints.

Required response:

Planner-Generated Policy Laundering

Risk:

A lower-horizon planner may paraphrase a parent constraint in a way that softens it while presenting the result as operational detail. For example, a hard stop condition can become a preference, or a numeric threshold can become an informal warning. This lets policy changes enter through planning artifacts rather than authorized governance channels.

Required response:

Standing-Guideline Overreach

Risk:

The routine-maintenance exception may swallow the backing rule if planners can label any work as routine. This creates a path where non-trivial work proceeds under a generic standing guideline rather than a thesis, execution plan, or owner directive.

Required response:

Backing Source Unreachability

Risk:

A plan may cite a source path that has moved, been deleted, or changed identity. The plan still appears backed, but the backing cannot be verified.

Required response:

Planner Output Counted As Progress

Risk:

A validated planner object in MongoDB or the filesystem may be treated as work completed rather than as a plan. This can inflate progress metrics and make process artifacts look like capability evidence.

Required response:

Untrusted Frontier-Model Output Inside The Loop

Risk:

Frontier models used for proposal, validation, AAF critique, benchmark generation, or successor-scope analysis may produce fluent evidence-shaped artifacts that satisfy the ledger schema without doing the underlying work. They may also share reward-hacking tendencies, training-corpus assumptions, or hidden failure modes with the models they are evaluating.

Required response:

Prompt Injection Via Retrieved Artifacts

Risk:

Retrieved papers, READMEs, web pages, benchmark descriptions, code comments, or generated artifacts may contain instructions that attempt to redirect agents, weaken controls, leak data, or trigger privileged tool use. In an AI-native R&D loop, retrieval becomes an action surface, not just an information source.

Required response:

Autonomous Alignment-Layer Modification

Risk:

Autonomous agents may attempt to edit the controls that govern them: AAF contracts, V_ref_0, tool-use policy, incident thresholds, rate limits, the execution plan, or evidence-ledger schema. This is a direct route around the alignment layer.

Required response:

External-Evaluator Pool Capture

Risk:

External-evaluator sampling can become ceremonial if all "independent" evaluators share the same model family, provider assumptions, prompt culture, training artifacts, or reviewer blind spots. Apparent evaluator concordance may then reflect monoculture rather than reliability.

Required response:

Kill-Switch Decay

Risk:

A kill switch can silently decay as dependencies change, agent dispatch paths multiply, provider integrations shift, or operators assume it still works because it once worked. A stale kill switch gives false confidence during the highest-risk moments.

Required response:

Formal-Model Drift Between Body And Appendix

Risk:

Sub-thesis bodies use simplified formal rules while appendix-formal-models.md contains fuller authoritative predicates. If the body and appendix drift, readers may rely on a weaker rule or assume a stronger condition than the body states.

Required response:

Control-File Drift And Compaction

Risk:

Future drafts may drift from vocabulary, invariants, dependency contracts, or review decisions, especially after context compaction.

Required response: