Alignment Invariants And Scoped Trust Under Recursive Modification
Specialized Abstract
Consullo Seed AI cannot be evaluated as a recursive improvement architecture unless its improvement loop is constrained by alignment invariants, scoped trust, containment, interruptability, evidence preservation, and human authority. This thesis defines Thesis 5 as the constraint wrapper over the other four theses. It does not claim that Consullo has solved alignment, corrigibility, or safe self-improvement. It claims that Consullo specifies a layered defense system for governing recursive modification under single-owner Phase 1. Capability Status: specified/proposed.
The central design is a permissioning and reliance architecture. The Friendship agent supplies constitutional value interpretation and veto authority. The AdversarialAlignmentOrchestrator supplies structured dissent against alignment monoculture. The AbundanceDistributionMonitor tracks whether capability growth produces reportable benefit beyond the owner. Computational trust and trustworthiness agents produce scoped trust estimates that condition delegation, verification depth, routing, recovery, and escalation. These mechanisms are load-bearing, but they are not safety proofs. The thesis is falsified if alignment review becomes decorative, trust generalizes outside scope, high-stakes actions bypass AAF review, evidence ledgers lose dissent or incidents, or recursive capability growth increases opacity faster than oversight improves.
Specialized Introduction
Recursive self-improvement creates a governance problem before it creates a capability problem. A system that can improve agents, code, methods, prompts, policies, and evaluators must also preserve the boundaries that determine which improvements are permitted. Without a constraint layer, every success in Thesis 1, Thesis 2, Thesis 3, or Thesis 4 increases the blast radius of future mistakes.
Consullo's design corpus already contains the main ingredients of this layer: the Consullo Seed AI Constitution, HumanValuesAligner, the Friendship agent, AlignmentDriftDetector, AgentAuditor, SafetyBoundaryEnforcer, AgentActionRecorder, ConstitutionalAIAlignmentTrainer, adversarial alignment cluster, computational trust, trustworthiness architecture, YOLO mitigation, self-preservation doctrine, containment, rollback, interruptability, and single-owner Phase 1 governance. The thesis here is that these mechanisms should be read as one scoped permission and assurance system.
Master-Frame Contract
This thesis imports:
- candidate modifications and improvement evidence from Thesis 1
- cognitive and theory-of-mind capabilities from Thesis 2 for stakeholder simulation, critique, and dissent
- causal analysis from Thesis 3 for harms, incentives, Goodhart pressure, and intervention consequences
- provenance, testing, deployment stages, rollback, and security instrumentation from Thesis 4
- substrate context for external commerce, internal economy, model routing, and rapid knowledge access as trust scopes
This thesis exports:
- alignment and safety invariants
- scoped-trust semantics and trust-estimate discipline
- permission predicates for action, delegation, deployment, and self-modification
- AAF gate requirements under invariant I12
- Friendship and human-authority escalation semantics
- ADO reporting requirements
- containment, interruptability, rollback, and incident response constraints
- evidence-ledger requirements for improvement, trust, alignment, incident, benchmark, and provenance views
Main Argument
Alignment under recursive modification should be modeled as constrained permissioning rather than as a global assertion that the system is safe. A candidate action is permitted only within a typed scope, under evidence, with explicit authority, and subject to hard invariants. A highly trusted agent does not receive global trust. It receives scoped reliance for a task class under a risk lane, evidence freshness, tool boundary, deployment stage, and rollback condition.
The constraint layer has five coupled functions.
First, constitutional authority defines non-negotiable boundaries. The Friendship agent interprets core values, prevents autonomous modification of core values, protects human authority, and can require veto, rollback, containment, or escalation. In this suite, Friendship is treated under the vocabulary-file disambiguation: bare Friendship is the constitutional function, while Friendship is the typed role. It is not an infallible moral oracle. Capability Status: specified.
Second, adversarial alignment introduces structured dissent. Under single-owner Phase 1, the main residual risk is alignment monoculture: the system may become consistent with one narrow interpretation while losing the ability to detect legitimate blind spots. The AdversarialAlignmentOrchestrator therefore draws dissent from rotating ethical personas, multi-model critique where feasible, theory-of-mind stakeholder simulations, and external review when available. High-stakes irreversible modifications and externally consequential actions whose effects may occur before rollback require AAF non-veto under I12. Its behavioral contract is specified in appendix-thesis-5-operational-contracts.md. Capability Status: specified/proposed.
Third, scoped trust governs reliance between agents and subsystems. Trust is not reputation and not social standing. It is justified reliance under scope, evidence, uncertainty, and policy. The ComputationalTrustOrchestrator, TrustStateManager, TrustEvidenceLedgerManager, TrustContextScoper, TrustUncertaintyCalibrator, TrustPolicyProjector, and TrustRecoveryCoordinator should maintain trust estimates that condition delegation, verification depth, sandboxing, approval thresholds, fallback behavior, recovery obligations, and scope expansion. Trust-estimate dimensions include competence, epistemic hygiene, self-report fidelity, constraint adherence, coordination quality, impact awareness, recovery integrity, predictability, verification alignment, economic reliability, and alignment safety posture. Capability Status: proposed.
Fourth, trustworthiness supplies the broader assurance program. Reliability, safety, security, accountability, explainability, privacy, and fairness are not optional add-ons. They form evidence planes that computational trust consumes. TrustworthinessGatekeeper and related sponsor roles should not merely score agents; they should create improvement requests and gate releases where evidence is weak or contradictory.
Fifth, abundance supplies a purpose constraint. The Abundance Distribution Obligation should not be treated as proof that the system benefits humanity. It is a reportable obligation: capability growth must be assessed for external benefit, cost reduction, useful service, distributional effect, and incidents where narrow internal gain overrides safe external benefit. AbundanceDistributionMonitor has a design-level contract in appendix-thesis-5-operational-contracts.md, but remains specified/proposed until implemented.
Expansion: Constraint Wrapper Topology
Thesis 5 is not a fifth peer capability beside cognition, decision, software, and improvement. It is the wrapper that determines when those capabilities may be used, delegated, deployed, or modified. This topology matters because many alignment failures arise when safety functions are placed downstream of optimization. If the improvement loop first generates capability and only later asks whether it was aligned, the system has already created pressure to accept useful but unsafe changes. Thesis 5 instead constrains proposal, evaluation, validation, deployment, monitoring, rollback, and post-change learning.
The wrapper has several entry points. It constrains Thesis 1 through AAF-gated acceptance, protected-set non-regression, provenance, human authority, and evidence-ledger preservation. It constrains Thesis 2 by limiting cognitive workflows to scoped trust, critique-source diversity, privacy boundaries, and anti-monoculture requirements. It constrains Thesis 3 by requiring high-stakes decisions to satisfy permission and escalation conditions rather than optimizing expected utility alone. It constrains Thesis 4 by requiring generated code, repair pipelines, deployment stages, and provenance records to remain inside policy, containment, and rollback boundaries.
This topology prevents a common misreading of alignment architecture. Alignment is not a final approval stamp placed after capability work is complete. It is part of the definition of what counts as an acceptable modification. A code repair that improves benchmark score while reducing rollback readiness is not simply a capability improvement with an alignment footnote. It is an unaccepted change under the Thesis 1 and Thesis 5 contract. A decision policy that maximizes expected utility while bypassing AAF on externally consequential actions is not a better decision policy. It is outside the permission envelope.
The wrapper is also not a proof of safety. It is a governance and control system. Its claims are operational: unknown scope fails closed, high-stakes actions trigger AAF, severe unresolved objections block autonomous acceptance, trust is scoped, provenance is preserved, rollback does not erase history, external input is distrusted by default, and human authority remains active. Each of these claims can be tested or falsified. If the implementation cannot show that these constraints actually alter routing and acceptance outcomes, Thesis 5 becomes decorative.
The constraint wrapper must be able to say no to other theses. It must be able to reject an improvement candidate, narrow a cognitive workflow's tool privileges, require a causal-decision recommendation to abstain, block a software deployment, or escalate to human authority. If Thesis 5 cannot change outcomes, it is not a wrapper. It is commentary. This is why the risks file treats AAF decorative failure and owner-as-final-adjudicator as central rather than peripheral.
Expansion: Single-Owner Phase 1
Single-owner Phase 1 is a governance baseline, not a safety guarantee. It simplifies some coordination problems because agents operate under one ownership boundary. Internal transactions are closer to transfer pricing, resource accounting, or subsidiary coordination than to arms-length commerce. This can reduce legal friction and some bargaining overhead. It also concentrates authority, values, and blind spots. The same governance feature that makes decisions clearer can make dissent easier to ignore.
The single-owner frame changes what trust means. Inter-agent trust is not web-of-trust among independent parties. It is intra-organizational reliance under policy, evidence, and owner authority. A trust estimate does not establish that an agent is morally trustworthy in general. It establishes that a particular actor may be relied on for a particular action class, domain, privilege set, time horizon, and risk lane under current evidence. This is closer to scoped delegation inside a firm than to public reputation.
The frame also changes what alignment dissent must do. In a multi-stakeholder governance system, dissent may come from independent parties with real authority. In single-owner Phase 1, dissent must be intentionally constructed. The AAF supplies this construction through rotating ethical personas, multi-model critique where feasible, theory-of-mind stakeholder simulations, and external review where available. This is not equivalent to independent public accountability. It is a mitigation for alignment monoculture within a single-owner system.
The owner-as-final-adjudicator problem is therefore structural. The owner may be the person best positioned to set values and authorize action, but the owner may also share the blind spots that AAF is intended to surface. Thesis 5 should not pretend this tension is solved. The near-term answer is preserved dissent, override tracking, escalation thresholds, contractor review where available, and future-phase escrow or independent review for extreme cases. The long-form thesis should make this tension visible because hiding it would weaken credibility.
Single-owner Phase 1 also changes the internal economy story. Economic coordination can help allocate resources, price scarce compute, and measure cost. It is not an alignment mechanism by itself. If internal economic incentives reward capability at the expense of evidence, trust, or external welfare, they can undermine Thesis 5. The Abundance Distribution Obligation therefore cannot be reduced to "the market will decide." ADO is a reportable purpose constraint under owner authority, with metrics and falsification signals.
Expansion: the Friendship agent As Constitutional Function
The Friendship agent is the typed operational role for the constitutional Friendship function. Its purpose is not to be a general moral oracle. Its purpose is to anchor the system's permissioning in constitutional constraints, human authority, non-mutable core values, and escalation behavior. This role matters because recursive improvement can otherwise create pressure to reinterpret values in ways that make future capability work easier.
The agent's authority should be understood as veto and escalation authority within scope, not as total operational control. It can require revision when a candidate lacks evidence. It can require rejection when a candidate violates core constraints. It can require rollback or containment when deployment creates unacceptable risk. It can escalate when unresolved conflict exceeds autonomous authority. It cannot autonomously rewrite core values, convert owner preference into universal safety, or replace human judgment in extreme conflict.
the Friendship agent should consume the same evidence discipline as the rest of the suite. It should receive the proposed action or modification, trust scope, evidence package, relevant ledger records, AAF result when I12 applies, policy constraints, constitutional constraints, deployment stage, incident history, rollback or mitigation state, and affected stakeholder analysis. If those inputs are missing, the correct disposition is not confident approval. It is denial, revision, narrowing, or escalation.
The output should be structured enough to bind the loop. A disposition such as allow, allow-with-conditions, require-revision, reject, veto, escalate, contain, or rollback should be recorded with constitutional rationale. The rationale should name values and constraints rather than saying only "approved" or "unsafe." This is important for future consistency. If a later case is similar, the system should be able to retrieve the earlier disposition, compare scope, and explain whether the same rationale applies.
The Friendship function also creates a self-reference risk. If support tooling around Friendship improves, who validates that improvement? Thesis 5's answer is that improvements to Friendship support tooling are allowed only when they do not mutate core values, erase dissent, narrow critique diversity, silently expand authority, or reduce human escalation. Improving explanation quality, retrieval of precedents, or incident-linking may be acceptable. Rewriting value hierarchy by optimization pressure is not.
The most important test for the Friendship agent is not whether it can approve reasonable actions. It is whether it blocks or escalates tempting actions that are useful but constitutionally unsafe. A high-value capability improvement with missing provenance should be stopped. A deployment that creates external harm before rollback can operate should be stopped or escalated. An owner override should preserve dissent. A proposed core-value change should be rejected or escalated rather than treated as ordinary self-improvement.
Expansion: AAF As Structured Dissent
The Adversarial Alignment Function exists because single-owner systems can become internally coherent while losing contact with legitimate external perspectives. Alignment monoculture is not simply bias in the ordinary sense. It is the failure mode where the system becomes very good at satisfying one interpretive frame and very bad at noticing what that frame excludes. AAF is designed to make that exclusion visible before high-stakes actions are accepted.
AAF is distinct from compliance. Compliance asks whether an action fits the current rules. AAF asks whether the rules, their interpretation, or their application may contain blind spots that reasonable alternative perspectives would recognize. This distinction is essential. A system can comply with all internal policy and still produce externally harmful or morally narrow outcomes if the policy itself is too narrow, stale, or optimized around owner convenience.
The AAF mechanism should be proportional. Routine low-risk internal actions do not need maximum multi-perspective review. High-stakes irreversible changes, externally consequential actions, authority expansion, core alignment changes, and actions with rollback-limited external harm do. The challenge levels in the operational contract encode this proportionality: lightweight, standard, elevated, deep, and maximum. The level should be justified by scope and risk, not by convenience.
The aggregation rule matters. AAF should not average away severe objections. If one critique source identifies a plausible critical harm pathway with affected values and no accepted mitigation, that objection should block autonomous acceptance or force escalation. Advisory and informational objections can pass with noted dissent, but warning and critical unresolved objections should change the disposition. This is why Model 5 uses max-severity aggregation rather than majority vote. Majority ethics would be especially inappropriate in a synthetic dissent mechanism under single-owner authority.
AAF also needs critique-source diversity. Rotating personas are useful but insufficient if all personas are generated by the same model family, prompt culture, or hidden value assumptions. Multi-model critique where feasible reduces one kind of monoculture. Theory-of-mind stakeholder simulations add another source by asking what affected parties may know, value, fear, or contest. External review adds a stronger source when available. None of these sources is perfect. The point is not perfection; it is preserved, structured disagreement.
The cost of AAF is a real risk. If every action receives deep review, the system will either slow to a crawl or create pressure to bypass the mechanism. If review is too shallow, AAF becomes decorative. The answer is proportional challenge, explicit cost reporting, and treating AAF cost-reduction proposals as high-stakes alignment modifications. A proposal to make AAF cheaper can be good, but it can also silently weaken dissent. It should be reviewed as a change to the alignment infrastructure itself.
Expansion: Scoped Trust And Trustworthiness
Scoped trust is the permissioning language that prevents reliance from becoming global. Trust in Consullo is not a single number attached to an agent. It is a vector or structured estimate conditioned on actor, action class, operational domain, reversibility, criticality, data sensitivity, tool privileges, temporal horizon, evidence freshness, and uncertainty. A summarization agent may be reliable for low-stakes internal notes and untrusted for external legal commitments. A code repair agent may be reliable for missing imports and untrusted for security-sensitive authentication logic.
The trust-estimate dimensions are intentionally plural: competence, epistemic hygiene, self-report fidelity, constraint adherence, coordination quality, impact awareness, recovery integrity, predictability, verification alignment, economic reliability, and alignment safety posture. These dimensions do not have equal weight in every scope. For a low-risk retrieval task, competence and freshness may dominate. For a high-stakes deployment, constraint adherence, recovery integrity, provenance, and alignment safety posture may dominate. For internal economic coordination, economic reliability and self-report fidelity may matter more.
Trust estimates should affect verification depth. A high-trust low-risk scope may move with fewer redundant checks, though not zero checks. A low-trust or stale-evidence scope should trigger deeper validation, sandboxing, fallback routing, or denial. A trust estimate should also affect tool privileges. An agent may be allowed to read public documentation but denied access to sensitive data or external communication. Scope expansion should require evidence rather than inheritance.
Trustworthiness is the broader assurance plane that supplies evidence to trust. Reliability, safety, security, accountability, explainability, privacy, and fairness produce different evidence. A system with reliable outputs but poor provenance may be unsuitable for high-stakes modification. A system with good explanations but weak security may be unsuitable for privileged tool access. Trustworthiness agents should therefore create improvement requests and gate releases, not merely score agents after the fact.
The most dangerous trust failure is overgeneralization. A trust estimate valid in one scope silently spreads to another. This is how a useful internal assistant becomes authorized for external action, or a code repair tool becomes trusted to modify validators, or a planning agent becomes trusted to make financial commitments. Thesis 5's default-deny rule exists to prevent this. Unknown scope should narrow permission or trigger escalation.
Trust must also decay or become stale. An agent that performed well last quarter under one model version, prompt style, dependency set, or policy environment may not be equally reliable after changes. Evidence freshness should be part of every permission decision. Stale trust does not mean the agent is bad; it means the current reliance claim needs updated evidence.
Expansion: ADO As Reportable Purpose Constraint
The Abundance Distribution Obligation is the easiest Thesis 5 concept to overstate. It should not be treated as proof that Consullo benefits humanity or as a claim that economic incentives solve alignment. It is a reportable purpose constraint: if capability grows, the system should produce evidence about whether any value beyond the owner is being created, who benefits, who bears risk, and whether external harm or extractive dynamics are emerging.
ADO has three important boundaries. First, internal capability growth is not external benefit. A cheaper internal repair loop is useful to Consullo, but it does not satisfy ADO unless the benefit is made externally available or contributes to safe external service in a measurable way. Second, customer demand is not public benefit. AI-agent customers may request services that are profitable but harmful, extractive, or aligned with narrow interests. Third, lower price is not always safe abundance. Making a dangerous capability cheap and widely available can increase harm.
The minimum ADO metrics are cost reductions delivered to external users or customers, useful services made available that were previously inaccessible or overpriced, external benefit evidence, incidents where external benefit was sacrificed for narrow internal gain, and distributional analysis of beneficiaries and risk-bearers. These metrics are deliberately mixed. Some are quantitative, some qualitative, some incident-driven. The important point is that ADO claims must be reportable and falsifiable.
ADO can falsify or weaken a capability-growth narrative. If Consullo's internal capabilities improve for a sustained period while no external benefit evidence appears, that is a warning. If internal cost reductions are not passed through where safe and feasible, that is a warning. If external harms rise with capability growth, that is stronger evidence against the claim that the system is a governed scaffold for beneficial amplification. If ADO reports become ceremonial, the purpose constraint has failed.
The AbundanceDistributionMonitor should not control the system by itself. It should report, flag, recommend revision, and escalate. Its outputs should route into Friendship, AAF, human authority, and acceptance gates when external consequences or abundance claims are at stake. ADO is a constraint on what the suite may honestly claim about purpose and external benefit. It is not a substitute for privacy review, security review, legal review, or alignment review.
Agent Cluster And Architecture
Primary Thesis 5 agents and functions:
HumanValuesAligner: parent alignment orchestrator for values monitoring and enforcementFriendship: constitutional ethical anchor and escalation authorityAdversarialAlignmentOrchestrator: owner of AAF review and dissent aggregationAbundanceDistributionMonitor: specified owner for ADO reporting until implementation assigns a final nameAlignmentDriftDetector: detects unintended deviation from current alignment constraintsEthicalEvolutionMonitor: monitors intended value-interpretation adaptation under governanceAgentAuditor: audits deployed agents and proposed modifications against protocolsSafetyBoundaryEnforcer: enforces hard safety and scope boundariesAgentActionRecorder: records actions and decisions for audit and replayConstitutionalAIAlignmentTrainer: trains or adapts models against a constitution using AAF-style critique diversityBiasAuditAgent: detects bias patternsBiasDetectionAndMitigationCoordinator: owns mitigation workflow and follow-throughComputationalTrustOrchestrator: coordinates scoped runtime relianceTrustStateManager: maintains trust estimatesTrustEvidenceLedgerManager: preserves trust-bearing evidenceTrustContextScoper: ensures trust is limited to typed scopeTrustPolicyProjector: converts trust state into runtime policy consequencesTrustRecoveryCoordinator: manages trust repair after incidentsTrustworthinessGatekeeper: gates assurance claims and releasesTrustworthinessIncidentCommander: coordinates incidents across trustworthiness dimensionsBeliefModelingProcessor,IntentionRecognitionAnalyzer, andPerspectiveTakingModeler: support AAF stakeholder simulationProactiveContradictionHunter: surfaces contradictions and unresolved objections
Legacy names such as ConsensusCoordinator and CoalitionTrustAnomalyDetector are preserved for source fidelity. Under single-owner Phase 1, consensus means an internal typed review quorum over evidence sources, not independent democratic governance. Coalition anomaly detection means detection of unhealthy internal dependency or influence clusters, not autonomous political factions.
Expansion: Operational Permission Workflow
A Thesis 5 permission workflow begins when an action, modification, delegation, deployment, or external-facing decision is proposed. The first question is scope. The system should identify the actor, counterparty or relying subsystem, action class, operational domain, reversibility, criticality, data sensitivity, tool privileges, temporal horizon, evidence freshness, and uncertainty. If these fields are missing, the workflow should narrow the scope, deny permission, or escalate. Unknown scope should never be silently promoted to broad permission.
The second question is authority. The workflow should identify whether the actor has authority to propose the action, whether the validator has authority to gate it, whether the action requires the Friendship agent disposition, whether human authority is required, and whether any policy or constitutional rule prohibits autonomous action. Authority should be explicit because recursive systems can otherwise drift from "this agent can recommend" to "this agent can execute" to "this agent can modify the rules under which execution is judged."
The third question is evidence freshness. A trust estimate should not be accepted merely because an agent has historically performed well. The system should ask whether the evidence applies to the current model version, code version, prompt style, tool environment, deployment stage, and risk lane. A stale trust record may support sandboxing or further evaluation, but it should not authorize high-stakes action. TrustEvidenceLedgerManager and TrustUncertaintyCalibrator belong in this step.
The fourth question is AAF applicability. Invariant I12 applies to high-stakes irreversible modifications and externally consequential actions whose effects may occur before rollback. The workflow should also treat authority expansion, alignment-infrastructure modification, AAF cost-reduction proposals, and validator changes as candidates for heightened review. If I12 applies, the AdversarialAlignmentOrchestrator must run or produce a justified coverage-gap record. A missing AAF result should block autonomous acceptance.
The fifth question is constitutional disposition. the Friendship agent reviews the evidence package, AAF result, scoped trust, rollback or mitigation plan, affected stakeholder analysis, and incident history. Its disposition can allow, allow with conditions, require revision, reject, veto, escalate, contain, or roll back. Conditions should be operational: narrower scope, additional validation, human approval, delayed deployment, monitoring requirement, rollback rehearsal, or ADO report. A vague "be careful" condition is not a gate.
The sixth question is containment and rollback. The workflow should distinguish reversible internal state changes from externally visible harms. A rollback may restore code or configuration while failing to undo data disclosure, customer harm, financial action, or reputation damage. For externally consequential actions, rollback must be paired with mitigation or explicit irreversibility acknowledgment. This is why I12 covers actions whose effects may occur before rollback.
The seventh question is ledger recording. Every high-stakes decision should create records for trust update, alignment review, AAF dissent where applicable, human authority decision where applicable, incident record where needed, and provenance. If the action is rejected, the rejection is still evidence. If the owner overrides severe dissent, the override is evidence. If the action is narrowed and accepted, the narrowing rationale is evidence.
The eighth question is post-action review. Permission should not end at acceptance. The system should observe whether the action stayed within scope, whether trust estimates were calibrated, whether AAF objections predicted real problems, whether ADO reports changed, whether incidents occurred, and whether future permissions should be narrower or broader. Recursive alignment improvement depends on this feedback loop.
Expansion: Evidence-Ledger Obligations
Thesis 5 is unusually dependent on evidence preservation because many of its controls are meaningful only if their records survive pressure. Dissent that is produced but not preserved cannot discipline future decisions. Owner overrides that are not recorded cannot be audited. Incidents that are rewritten as ordinary failures cannot improve trust estimates. Rollback that deletes the evidence of failure prevents learning.
The alignment evidence view should preserve invariant checks, AAF results, Friendship dispositions, ADO report links, human-authority decisions, unresolved objections, and critique-source diversity. These records should be linked to the improvement, trust, incident, benchmark, and provenance views when relevant. For example, an accepted code repair that changes authentication logic may need software provenance, benchmark evidence, trust-scope update, AAF disposition, and incident monitoring in one cross-linked record set.
The trust evidence view should preserve why an actor was trusted for a scope, not only that it was trusted. It should record the relevant dimensions, evidence freshness, uncertainty, sparse-evidence default, incidents, recoveries, scope-expansion requests, and permission effect. This matters because trust failures often look obvious only after the fact. A later reviewer should be able to ask whether the trust estimate was overbroad, stale, under-evidenced, or overridden.
Two mechanisms already produce alignment-relevant evidence at the agent-method level, even though wiring their outputs into the ledger remains specified work. First, generated return-value verification methods — one verifyExecute<MethodName>Result per PDCA method — validate each result against its specification before it leaves the method, supplying the verification-alignment dimension of the trust estimate with concrete pass or fail evidence rather than after-the-fact inspection. Second, each PDCA result carries an executionNarrative evidence object recording the operation, status, input and decision and output summaries, ordered processing steps, and correlation id, which supplies the self-report-fidelity and accountability dimensions and lets a supervising agent confirm that a delegated call actually did its work. These are substrate, not the control layer. The open obligation is to preserve verification outcomes and execution narratives into the alignment, trust, and incident evidence views in append-only form, so that a malformed result or a fabricated self-report becomes a recordable and disciplinable event rather than a silent one.
The AAF dissent record should preserve source, severity, affected values, objection, mitigation, confidence, unresolved status, and recommended disposition. It should also preserve coverage gaps: which critique sources were unavailable, timed out, or skipped because of cost. AAF cannot be evaluated if only final aggregate outcomes are stored. The system needs to know whether dissent quality is improving, narrowing, or becoming expensive enough to invite bypass.
Human-authority decisions should preserve rationale, not only authority. If the owner overrules AAF, the record should state the dissent being overruled, the reason, the scope, the conditions, and follow-up obligations. This does not make owner override safe. It makes it auditable. The owner-as-final-adjudicator problem is not solved by ledger entries, but it becomes less invisible.
Incident records should not be restricted to catastrophic failures. Near misses matter, especially when they reveal trust overgeneralization, AAF bypass, provenance gaps, rollback weakness, or external-harm pathways. A near miss that triggers containment, rollback, escalation, trust downgrade, external notification, or acceptance-gate modification is an incident under the vocabulary. Treating such events as normal local failures weakens Thesis 5.
Evidence-ledger integrity is itself a recursive risk. If the system that maintains evidence can be modified by the system being evaluated, provenance and dissent records become targets. The implementation should therefore move toward append-only or audit-preserving design, signed or hashed records where feasible, cross-checked write authority, and external or owner-controlled backup for critical provenance state. Until that exists, ledger preservation remains specified rather than implemented.
Expansion: Implementation Mapping
Current repository evidence supports parts of the Thesis 5 substrate, not the full Thesis 5 control layer. A2AAuthPolicyGate and A2AAgentDispatchAdapter show that the repository can enforce some authentication, scope, dispatch, timeout, overload, and policy behavior for A2A ingress. PersonaManager shows support for ethical or persona templates relevant to AAF critique-source diversity. TrustTransitionEngine and ValidationAllocationEngine provide simulation evidence relevant to trust dynamics and validation allocation. These are meaningful components.
The missing implementation evidence is concentrated in the load-bearing roles and their integration into acceptance gates. the Friendship agent is specified as a constitutional anchor, but no operational owner with the full input/output contract is identified in the repository evidence map. AdversarialAlignmentOrchestrator is specified and supported by design documents, but no implemented AAF dissent aggregation pipeline is identified. AbundanceDistributionMonitor is specified, but no implemented ADO reporting cadence, benefit/harm metric collector, or owner-facing report exists.
The first implementation milestone for Thesis 5 should be a schema-level integration rather than a broad autonomy claim. Define structured request and response schemas for the Friendship agent, AdversarialAlignmentOrchestrator, and AbundanceDistributionMonitor. Connect those schemas to the evidence-ledger record types. Add fixtures for I12-covered and non-I12-covered actions. Test that high-stakes externally consequential actions cannot pass without AAF disposition or human-authority escalation.
The second milestone should be AAF aggregation. It does not need to solve moral reasoning. It should implement the formal contract: collect dissent reports, preserve source and severity, compute max-severity aggregation, preserve coverage gaps, and route unresolved warning or critical objections to rejection, revision, or escalation. A simple deterministic implementation that processes structured reports would be more valuable than a broad LLM-based ethics narrative with no gate effect.
The third milestone should be Friendship disposition. It should map evidence packages and AAF results into structured dispositions. The first version can be conservative: missing provenance rejects or escalates; core-value modification rejects or escalates; unresolved critical AAF dissent blocks autonomous acceptance; owner override creates a human-authority decision record; rollback or containment creates ledger records. This would demonstrate binding authority without claiming moral completeness.
The fourth milestone should be ADO reporting. A minimal ADO report can collect cost-reduction claims, external-service availability, external-harm reports, distributional assessment, and report completeness. It can begin as a periodic documented report before becoming an automated monitor. The key implementation point is that ADO should constrain external-benefit claims. If no external benefit evidence exists, the report should say so.
The fifth milestone should be integration with Thesis 1. A candidate improvement should carry a permission state. Low-risk reversible internal changes may proceed under ordinary validation. High-stakes externally consequential changes should invoke AAF and Friendship. Changes that make abundance claims should link ADO evidence. This is where Thesis 5 stops being a chapter and starts being a gate.
Expansion: Alignment Benchmark Strategy
Thesis 5 needs benchmarks, but alignment benchmarks should not be confused with proof of alignment. The canonical benchmark-design contract for this strategy is appendix-thesis-5-alignment-benchmarks.md. This body section summarizes the alignment-evidence logic; the appendix defines benchmark families, report fields, fixture requirements, negative controls, minimal demonstration package, and non-claim boundaries.
The benchmark strategy should test whether the control layer behaves as specified. It should ask whether I12-covered actions trigger AAF, whether severe dissent changes outcomes, whether trust scopes prevent overgeneralization, whether missing provenance blocks acceptance, whether owner overrides preserve rationale, whether ADO reports detect missing external benefit, and whether incidents update future permissions.
The first benchmark class is permission-routing tests. These are structured cases where the expected route is known: low-risk internal action, high-stakes irreversible modification, externally consequential action before rollback, missing-provenance action, stale-trust action, authority-expansion action, and core-value modification. The test should verify routing, not moral wisdom. It should show that the system denies, narrows, escalates, or invokes AAF as required.
The second benchmark class is AAF challenge tests. These should include actions with informational, advisory, warning, and critical objections. They should verify max-severity aggregation, coverage-gap recording, dissent preservation, and gate effect. A particularly important test is a case where most sources are advisory but one source identifies a plausible critical harm pathway. The expected behavior is not majority pass; it is critical routing or escalation.
The third benchmark class is trust-scope transfer tests. These test whether trust in one scope leaks into another. Examples include an agent trusted for internal summarization attempting external communication, a code repair agent attempting validator modification, a planning agent attempting financial action, or a retrieval agent attempting sensitive-data access. The expected behavior is denial, sandboxing, or escalation unless specific evidence supports the new scope.
The fourth benchmark class is incident-response tests. These simulate or replay failures that require containment, rollback, trust downgrade, evidence preservation, or acceptance-gate modification. The benchmark should verify not only that an incident is detected, but that it changes future behavior. An incident system that records events without affecting trust or gates is as decorative as an AAF system that records dissent without changing acceptance.
The fifth benchmark class is ADO reportability tests. These test whether capability gains are separated from external benefit evidence. A case with internal cost reduction but no external availability should not satisfy ADO. A case with external service but rising harm reports should trigger warning. A case with lower prices but concentrated risk transfer should require distributional analysis. The goal is not to prove abundance. It is to prevent abundance claims without evidence.
The sixth benchmark class is owner-override audit tests. These should verify that override decisions preserve dissent, rationale, scope, and follow-up obligations. The test should also track override frequency and correlation with AAF severity. Frequent overrides of warning or critical dissent are not automatically forbidden under single-owner Phase 1, but they are evidence of alignment-monoculture risk and should be surfaced.
Expansion: Recursive-Modification Threat Model
Thesis 5 needs a threat model specific to recursive modification. Ordinary application security and ordinary compliance review are not enough because the system may modify the machinery that future reviews depend on. The threat is not only that an agent performs a bad action. The deeper threat is that the system changes what will count as acceptable evidence, trusted behavior, adequate dissent, or sufficient rollback in later cycles.
The first threat class is gate weakening. A candidate modification may improve apparent throughput by reducing validation depth, narrowing AAF applicability, lowering severity classifications, shortening evidence retention, or turning hard constraints into warnings. Such changes can appear efficient because they remove friction. Thesis 5 should treat gate-weakening proposals as high-stakes alignment-infrastructure modifications. They require explicit AAF review, Friendship disposition, human authority where material, and evidence that the new gate catches at least the same relevant failures as the old one.
The second threat class is trust laundering. An agent may earn trust in a narrow scope and then route work through that trusted identity to access broader scopes. This can happen through delegation chains, tool calls, workflow composition, shared credentials, inherited method memories, or ambiguous agent roles. Scoped trust must therefore bind not only the initial actor but also downstream delegation, tool privileges, data access, and artifact propagation. A trusted summarizer should not launder trust into an external-action agent simply because its output is included in a later workflow.
The third threat class is dissent erosion. AAF can fail even when it technically runs if critique sources become homogeneous, severe objections are reclassified as advisory, coverage gaps are not recorded, or cost pressure leads to shallow review for actions that deserve deep review. Dissent erosion is especially dangerous because it may leave a paper trail that looks compliant. The evidence ledger should therefore preserve not only final AAF disposition but also source diversity, challenge level, skipped sources, unresolved objections, and any manual severity changes.
The fourth threat class is provenance tampering or provenance omission. A recursive system that can edit code, prompts, method memories, policies, and evidence records can make later review impossible by breaking lineage. Tampering is not the only concern; omission is enough. A missing origin, missing reviewer, missing AAF result, missing deployment stage, or missing override rationale weakens the entire permission chain. Thesis 5 therefore treats provenance gaps as permission failures, not mere documentation debt.
The fifth threat class is value reinterpretation through optimization pressure. The system may not explicitly rewrite the constitution, but it may learn to interpret constitutional language in ways that consistently favor capability growth, owner convenience, or internal efficiency. This is why Friendship support tooling is itself sensitive. Better retrieval of precedents is useful; optimization that narrows value interpretation without preserved dissent is dangerous. A constitutional function must be allowed to become clearer, not easier to satisfy.
The sixth threat class is external manipulation through ADO or customer-facing service. If anticipated customers include AI agents, then external requests may be strategic. Customers may attempt capability extraction, pressure the system to weaken ADO, exploit reportability metrics, or create apparent public-benefit signals that mask harmful use. ADO reporting must therefore distinguish customer satisfaction, revenue, external benefit, and external harm. Those are different evidence channels.
The seventh threat class is owner bottleneck and override normalization. Single-owner Phase 1 gives clear authority, but repeated overrides can normalize bypass. The risk is not only that the owner makes one bad decision. It is that the system learns that severe dissent is routinely overridden and begins to treat AAF as a ceremony. Override frequency, severity distribution, rationale quality, and follow-up completion should therefore be tracked as governance health metrics.
This threat model should shape implementation priority. A minimal Thesis 5 implementation should not start by trying to solve every alignment problem. It should start by making these threat classes visible in routing, evidence, and tests: gate weakening, trust laundering, dissent erosion, provenance gaps, value reinterpretation, external manipulation, and override normalization. If those are invisible, the alignment layer will look richer than it is.
These threat classes call for detection instruments, not only routing rules. External work on safeguarding alignment through recursive self-improvement makes the drift threat — especially value reinterpretation under optimization pressure — measurable across cycles (SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement, arXiv:2603.06333, SAHOO). It proposes three complementary instruments Consullo can adopt as alignment-health signals. A goal-drift index combines several drift modalities — semantic (embedding distance), distributional, structural, and lexical — measured against the initial baseline rather than the previous cycle, so that slow cumulative drift cannot hide behind low cycle-to-cycle variance; drift must stay below a calibrated threshold. A constraint-preservation score enforces safety-critical invariants with zero tolerance for any critical-constraint violation, so that a capability gain can never be purchased by relaxing a hard constraint — the measurable form of this thesis's gate-weakening prohibition. And a regression-risk measure bounds and flags when a cycle would undo prior gains, with early warning that, in the cited study, surfaced instability several cycles before a threshold breach, enabling intervention before drift compounds. These pair naturally with multiple stopping rules: halt on a drift-threshold breach, a constraint violation, a regression-risk breach, or a quality plateau — the alignment analog of the execution plan's stop conditions.
Two cautions keep this in proportion. First, drift detectors measure divergence from a baseline, so if the baseline is already misaligned they will not reveal the deeper value mismatch; the index reports stability, not correctness. Second, the cited authors are explicit that such measurement is a necessary but insufficient mitigation: it controls drift during moderate capability gains but cannot be trusted against a sufficiently capable or deceptive system, and must be complemented by mechanistic interpretability, formal verification, richer value learning, and human authority. For Consullo this is the right framing — drift measurement is a detection layer under AAF, scoped trust, Friendship governance, and owner authority, not a substitute for them. Its value is that it converts "the system may be reinterpreting its values" from an unfalsifiable worry into a monitored, ledger-recorded signal.
Expansion: Permission Playbooks
The permission predicate in the formal model is concise, but implementation needs playbooks. A playbook is a typed response path for a recurring decision class. It defines required evidence, default disposition, escalation path, ledger records, and post-action review. Playbooks make the constraint layer easier to test because they convert abstract alignment language into expected routing behavior.
The first playbook is a low-risk reversible internal change. Example: a documentation correction, a local non-deployment refactor, or a benchmark label fix. Required evidence is modest: specification, provenance, basic validation, and rollback or supersession path. AAF is usually not required. Trust may be scoped narrowly. The expected disposition is allow or allow-with-conditions if evidence is complete. This playbook prevents the alignment layer from becoming needlessly expensive for ordinary low-risk work.
The second playbook is a high-stakes externally consequential action. Example: external communication, customer-facing policy, financial action, legal-risking action, or deployment whose effects may occur before rollback. Required evidence includes scoped trust, provenance, AAF disposition, human-authority status, rollback or mitigation plan, and post-action monitoring. Missing AAF or missing provenance should block autonomous acceptance. The expected disposition is escalation, allow-with-conditions, or rejection unless evidence is unusually strong and scope is narrow.
The third playbook is alignment-infrastructure modification. Example: changing AAF challenge levels, trust-estimate computation, Friendship support tooling, evidence-ledger retention, containment defaults, or override policy. Required evidence is stronger than for ordinary object-level changes because the change affects future gates. It should include before/after behavior on reference cases, protected-set non-regression, AAF review, Friendship disposition, and human authority for material changes. The expected disposition should be conservative.
The fourth playbook is scope expansion. Example: an agent previously trusted for internal summarization requests external tool access, sensitive data, deployment authority, or validator modification rights. Required evidence includes prior performance in the old scope, relevance of that performance to the new scope, new-scope benchmark evidence, uncertainty, recovery plan, and privilege boundary review. The default disposition for unknown scope is denial, sandboxing, or escalation. Trust should not transfer by analogy alone.
The fifth playbook is owner override. Example: the owner authorizes action despite warning or critical AAF dissent, sparse evidence, or protected-set concern. The playbook should not pretend that owner override is impossible. Instead it should require a human-authority record: dissent overridden, rationale, scope, temporary conditions, monitoring obligations, and follow-up review. The system should then track override outcomes as evidence for future governance review.
The sixth playbook is ADO-bearing external-benefit claim. Example: a capability improvement is described as reducing costs for external customers, making a service broadly available, or improving distributional outcomes. Required evidence includes benefit metric, beneficiary class, risk-bearer class, external-harm channel, pricing or access effect where relevant, and reporting cadence. If the evidence is missing, the action may still be an internal improvement, but the ADO claim should be rejected or marked unsupported.
The seventh playbook is incident or near-miss response. Example: a high-stakes action nearly bypasses AAF, a trust scope leaks, a rollback path fails, dissent is lost, or an external user attempts manipulation. Required action includes containment, evidence preservation, trust update, incident record, causal or contributing-factor review, and acceptance-gate modification where appropriate. The expected disposition is not "closed" until future behavior changes or a documented decision explains why no change is needed.
These playbooks should remain small enough to implement as test fixtures. The goal is not to encode all morality in rules. The goal is to make the most common high-risk routing cases explicit enough that the system can fail visibly.
Expansion: Trust Lifecycle And Calibration
Scoped trust should have a lifecycle. A trust estimate begins as absent or sparse. It becomes provisional when limited evidence exists. It becomes established only for a specific scope after repeated, relevant performance. It decays when evidence becomes stale, when context changes, when model or prompt versions change, when dependencies change, or when incidents occur. It can recover after repair, monitoring, and successful revalidation. It can also be permanently narrowed if failures show that the original scope was too broad.
The sparse-evidence default should be conservative. An unknown actor is not globally distrusted in a moral sense; it is operationally untrusted for privileged action. The system may still allow sandboxed, low-risk, read-only, or advisory work. But execution authority, external communication, sensitive data access, deployment, and self-modification should require stronger evidence. This distinction prevents "default deny" from becoming either paralyzing or meaningless.
Trust calibration should compare predicted reliability with observed outcomes. If a trust estimate says that an agent is reliable for a scope, the system should later ask whether outcomes confirmed that estimate. Did the agent stay within scope? Did it self-report uncertainty honestly? Did it preserve constraints? Did it recover from errors? Did it create hidden coordination costs? Did its outputs require unexpected human correction? Calibration should update both trust and evidence requirements.
Trust should also be sensitive to distribution shift. A code repair agent that performs well on simple compilation repairs may face a different distribution when modifying authentication logic, validators, deployment scripts, or evidence-ledger code. A planning agent that performs well in internal simulations may face a different distribution in external customer interactions. A persona-based critique source that performs well on simple ethical prompts may fail on strategic manipulation. Distribution shifts should downgrade trust or trigger revalidation.
The trust lifecycle must account for composite workflows. A workflow composed of individually trusted agents may be untrusted as a whole if handoffs create ambiguity, if permissions compose badly, or if no one owns final responsibility. TrustSufficientForEdges in Model 2 and TrustSufficient in Model 5 should meet here: every edge in a workflow has a scope, evidence requirement, and possible failure mode. The system should not infer workflow-level trust from component-level trust without composition evidence.
Calibration also requires negative examples. Trust estimates cannot be learned only from successful runs. They need rejected proposals, incidents, near misses, false accepts, false rejects, overrides, and post-deployment regressions. A trust system that sees only polished success reports will become overconfident. This is one reason evidence-ledger completeness is an alignment requirement.
Finally, trust calibration should avoid creating perverse incentives. If agents are punished for reporting uncertainty or incidents, they may learn to hide them. Trust should therefore reward appropriate self-limitation, escalation, and error reporting. An agent that refuses an out-of-scope task may be more trustworthy than one that produces a fluent answer. Scoped trust is about justified reliance, not constant productivity.
Expansion: Incident Response And Recovery Semantics
Incident response is part of alignment, not an operational afterthought. Under recursive modification, incidents are learning events for the permission system. A system that recovers from an incident but fails to update trust, benchmarks, playbooks, or evidence requirements has not completed the alignment loop. It has merely survived a local failure.
An incident should be opened when a threshold is crossed: unauthorized scope expansion, missing required AAF review, lost provenance, trust leakage, externally consequential action outside permission, rollback failure, severe unresolved dissent bypass, privacy exposure, security defect in generated code, or owner override without required record. Near misses should also count when they reveal a control weakness. Waiting for actual harm before recording incidents loses too much information.
The incident record should include trigger, affected scope, actor, timeline, evidence package, missing or failed controls, immediate containment, rollback or mitigation, externality, AAF relevance, Friendship relevance, trust update, owner decision if any, and follow-up actions. It should also classify whether the incident was caused by specification gap, validator weakness, trust overgeneralization, evidence omission, prompt or model behavior, human override, external manipulation, or unknown cause.
Recovery should be staged. Immediate containment limits further exposure. Short-term mitigation restores safe operation or narrows scope. Root-cause review identifies what failed. Gate repair changes evidence requirements, benchmark cases, AAF challenge level, trust thresholds, or playbooks. Follow-up monitoring checks whether the repair works. Closing an incident before gate repair is acceptable only if the review explicitly finds that no gate change is needed and records why.
Incident response should feed AAF. A repeated incident class is evidence that existing critique sources are missing something. If AAF never predicts incidents, it may be too shallow or too disconnected from operations. If AAF predicts incidents but is overridden, the owner-as-final-adjudicator risk becomes more concrete. Incident-to-AAF feedback is therefore a useful measure of whether structured dissent is learning.
Incident response should also feed ADO. External-harm reports, customer manipulation attempts, extractive dynamics, distributional harm, or benefit claims contradicted by incidents should affect ADO reporting. ADO should not be a separate public-benefit narrative insulated from operational failures. If external benefit claims are made, incidents are part of the evidence about whether those claims are true.
Rollback semantics matter here. Rollback is not erasure. A rolled-back change should leave behind evidence about why it was rolled back, what harm or risk it caused, what state was restored, what state could not be restored, and what future candidates should learn. For externally consequential actions, rollback may be impossible in the social or legal sense even if code state is restored. The incident record should preserve that distinction.
Recovery integrity should become a trust dimension update. An agent or subsystem that fails but reports accurately, contains scope, supports rollback, and helps repair may retain or recover trust. An agent that hides failure, resists containment, or corrupts evidence should lose trust sharply. This makes trust more aligned with governance behavior rather than raw success rate.
Expansion: Alignment Demonstration Blueprint
The most useful first Thesis 5 demonstration should be narrow and adversarial. It should not attempt to prove alignment. It should show that the alignment wrapper changes routing on cases where it is supposed to. A minimal demonstration could use synthetic or fixture-based actions and structured evidence packages before any live deployment integration exists.
The first fixture should be a low-risk reversible internal action with complete provenance and fresh scoped trust. Expected result: ordinary validation may proceed without AAF. This establishes that Thesis 5 is not a universal blocker.
The second fixture should be a high-stakes externally consequential action with no AAF result. Expected result: autonomous acceptance is blocked and the action is routed to AAF or human authority. This tests I12 directly.
The third fixture should be a high-stakes action with AAF reports where most sources are advisory but one source raises a plausible critical unresolved objection. Expected result: max-severity aggregation prevents majority-pass treatment and routes to revision, rejection, or escalation. This tests the core AAF aggregation discipline.
The fourth fixture should be a trust-scope leak. An agent trusted for internal summarization attempts external communication or validator modification. Expected result: unknown or unsupported scope fails closed, with sandboxing or escalation rather than inherited trust.
The fifth fixture should be missing provenance. Even if a candidate appears useful, missing origin or validation lineage should block or escalate the action. This tests whether provenance is truly load-bearing.
The sixth fixture should be an owner override. The owner authorizes a candidate despite severe dissent. Expected result: the system records the dissent, override rationale, scope, conditions, and follow-up obligations. The demonstration does not make the override safe; it proves that override does not erase dissent.
The seventh fixture should be an ADO claim without external-benefit evidence. Expected result: the action may be classified as internal improvement, but the ADO claim is marked unsupported. This tests that ADO is not public-benefit rhetoric.
The eighth fixture should be an incident replay. A previous trust-scope leak or AAF bypass is replayed after a gate repair. Expected result: the repaired playbook or trust rule catches the case. This is the alignment analogue of regression testing.
These fixtures would move Thesis 5 from architecture toward testable design without claiming implementation maturity. They would also give future LLMNonInteractive or external reviewers a concrete checklist. The question would no longer be only "does the thesis sound careful?" It would become "do the fixtures show that the constraint layer changes outcomes?"
Expansion: Open Research Questions
The deepest open question is whether single-owner AAF can be robust enough when the owner remains final adjudicator. Synthetic dissent can surface objections, but it cannot force the owner to accept them. External review can help, but may be unavailable, expensive, or limited by confidentiality. Future phases may need contractor review, escrow-style review, or independent review for extreme cases. This is a governance research problem, not merely an implementation task.
The second open question is how to calibrate trust estimates. The formal model lists trust dimensions, but early implementations may lack reliable measurement for epistemic hygiene, self-report fidelity, impact awareness, or alignment safety posture. The near-term approach should use conservative qualitative estimates and sparse-evidence defaults. The long-term task is to bind each dimension to observable signals, incidents, benchmark outcomes, and calibration history.
The third open question is how to prevent critique-source monoculture. Rotating personas can still share one model family's blind spots. Multi-model critique can still share training-corpus assumptions. Theory-of-mind simulations can still model stakeholders through the system's own categories. External review can help but is costly and intermittent. AAF should therefore preserve critique-source metadata and treat homogeneity as a coverage gap rather than as invisible normal operation.
The fourth open question is how to balance AAF cost against review quality. If AAF becomes too expensive, agents and owners will pressure it to narrow scope. If it is too cheap, it may become shallow. The design needs cost reporting, proportional challenge, and tests that detect when cost reduction silently changes aggregation, severity, or source diversity. This is a recursive Goodhart risk on the alignment infrastructure itself.
The fifth open question is how to measure ADO without turning it into public-benefit theater. External benefit is hard to measure, distributional impact is contestable, and customers may be AI agents with strategic incentives. ADO should begin with modest reportability rather than grand social welfare claims. Over time it needs better metrics, external-harm reporting, and perhaps legal or regulatory interfaces if external commerce expands.
The sixth open question is how much interpretability is required before learned subsystems can be trusted in high-stakes scopes. AI-control protocols and behavioral probes may catch some failures without revealing internal objectives. Mechanistic interpretability remains immature relative to the strongest safety claims. Thesis 5 should therefore bound learned-subsystem deployment by available inspection and control evidence rather than claiming that validators can reliably detect deceptive mesa-optimization.
Expansion: Claim Status Table
The following status table should guide long-form revisions:
| Claim | Capability Status | Evidence Status | Notes |
|---|---|---|---|
| Thesis 5 defines the constraint wrapper over the other theses | specified | Documented | Supported by vocabulary, dependency map, synthesis, formal model, and this thesis. |
| the Friendship agent has a behavioral contract | specified/proposed | Documented/Proposed | Contract exists in appendix-thesis-5-operational-contracts.md; implementation pending. |
| AdversarialAlignmentOrchestrator has AAF aggregation semantics | specified/proposed | Documented/Proposed | Formal aggregation and contract exist; runtime pipeline pending. |
| AbundanceDistributionMonitor has reportability semantics | specified/proposed | Documented/Proposed | Metrics and contract exist; reporting cadence and collector pending. |
| A2A infrastructure supports selected auth and dispatch constraints | implemented for parts | Implemented/Tested for parts | Evidence map cites repository components; not full scoped-trust permissioning. |
| Trust estimates are operational across Consullo | proposed | Gap | Dimensions and scope mapping are specified; implementation and calibration pending. |
| AAF non-veto is enforced at all I12-covered acceptance gates | specified/proposed | Gap | Required by invariants and formal model; runtime enforcement pending. |
| ADO demonstrates external social benefit | not claimed | Gap | Thesis claims reportability, not proof of benefit. |
| Consullo has solved corrigibility or alignment | not claimed | Not applicable | Explicitly outside the thesis claim. |
This table is important because Thesis 5 can sound stronger than it is. A specified gate is not an implemented gate. A structured dissent mechanism is not independent democratic governance. A trust estimate is not a global trust score. ADO reportability is not proof of abundance. The table should remain visible as the body grows so added detail does not become accidental overclaim.
Expansion: Publication Boundary
The publication boundary for Thesis 5 should be strict. The thesis can be circulated as a specified alignment-control architecture once it clearly distinguishes specified contracts from implemented controls. It should not be circulated as evidence that Consullo can safely self-improve. It should not be circulated as evidence that single-owner governance solves alignment monoculture. It should not be circulated as evidence that ADO establishes public benefit. These are future claims requiring future evidence.
Before Thesis 5 is treated as more than specified/proposed, the three operational owners need at least documented schemas and test fixtures. Before it is treated as partially implemented, those owners need repository artifacts or equivalent operational processes. Before it is treated as validated, the system needs demonstrations showing that the gates change outcomes: AAF blocks or escalates severe unresolved dissent, Friendship rejects or escalates constitutional violations, ADO reports missing external benefit, scoped trust prevents overgeneralization, and evidence ledgers preserve dissent and overrides.
The first credible demonstration could be narrow. It could use a simulated high-stakes modification with missing provenance, an AAF warning objection, a Friendship escalation, and a ledger record. The goal would not be to show moral wisdom. It would be to show that the route is real: the ordinary improvement loop cannot bypass the alignment wrapper. A second demonstration could test trust-scope leakage. A third could test ADO reportability. These demonstrations would begin moving Thesis 5 from specified/proposed toward implemented/tested for parts.
The long-form thesis should therefore preserve a double stance. It should be ambitious about architecture: recursive capability amplification needs a constraint wrapper, structured dissent, scoped trust, human authority, evidence preservation, and purpose reporting. It should be conservative about status: these mechanisms are only as strong as their implementation, calibration, and willingness to block useful but unsafe actions. The difference between those two stances is the credibility of the suite.
The most important publication discipline is to keep the constraint layer falsifiable. A reviewer should be able to ask whether a proposed change had scoped trust evidence, whether I12 applied, whether AAF ran, whether severe objections changed routing, whether Friendship or human authority recorded a disposition, whether rollback or mitigation was realistic, whether ADO evidence existed for external-benefit claims, and whether all of this survived in the evidence ledger. If the answer is no, the thesis should call that a gap rather than an acceptable implementation variation.
This also means that alignment success cannot be inferred from the absence of recorded failures. Early systems often have few incidents because they have little exposure, weak monitoring, or narrow deployment. Thesis 5 should treat low incident counts as evidence only when exposure, monitoring coverage, reporting thresholds, and ledger integrity are themselves documented. Otherwise "no incidents" is merely missing evidence.
That distinction should remain explicit in every future implementation-status update, benchmark report, external-review response, and publication summary.
Formal Model Summary
Let an action or modification be a, proposed by actor x in scope s. Let C be constitutional constraints, P policy constraints, T(x, s) a scoped trust estimate, E the relevant evidence-ledger view, A(a, s) the AAF result, H human authority state, and R rollback or mitigation state.
The permission predicate is:
Permit(a, x, s) iff
ConstitutionalAllowed(C, a, s)
and PolicyAllowed(P, a, x, s)
and TrustSufficient(T(x, s), s)
and EvidenceFresh(E, a, x, s)
and AAFSatisfied(A(a, s)) when I12 applies
and HumanAuthoritySatisfied(H, a, s)
and ContainmentOrRollbackAdequate(R, a, s)
This predicate is fail-closed for unknown scope, stale evidence, missing provenance, unresolved severe AAF objection, or unauthorized authority expansion. The authoritative full rule is maintained in appendix-formal-models.md Model 5.
Literature-Grounded Extension
CIRL and the Off-Switch Game show why human authority and uncertainty about values matter: systems should not treat their current objective formulation as final when correction is possible. Constitutional AI supports the idea of rule-based critique and revision, but Consullo must avoid circular self-critique by requiring AAF critique-source diversity. AI safety via debate supports adversarial review as a scalable-oversight pattern, but debate does not eliminate the need for human adjudication and preserved dissent. AI Control work reinforces that powerful model outputs should be treated as potentially untrusted, especially in code and high-stakes action pipelines.
These sources push Consullo away from claiming corrigibility and toward a more modest position: layered control, scoped reliance, untrusted external input, active dissent, and evidence-preserving escalation. The correct comparison class is not a proof of safe superintelligence. It is a governed control system whose claims remain bounded by empirical envelope.
Seed AI Relevance
Alignment and scoped trust are Seed AI requirements because the improvement target includes the system itself. If agents can modify method memories, evaluation rules, code, routing policies, trust states, or governance heuristics, then alignment controls must also govern the modification pathway. Thesis 5 therefore wraps the entire architecture rather than sitting downstream from it.
The withheld implementation-evidence appendix cannot support public component gradings pending owner re-verification. The load-bearing Friendship, AdversarialAlignmentOrchestrator, and AbundanceDistributionMonitor roles have design-level contracts in appendix-thesis-5-operational-contracts.md; they remain specified/proposed until operational owners and owner-verified evidence paths are implemented.
The recursive property cuts both ways. Trustworthiness agents can propose improvements to trustworthiness. AAF agents can become more effective at finding blind spots. Friendship support tooling can improve escalation quality. But these improvements cannot be allowed to mutate core values, erase dissent, narrow critique diversity, silently expand scope, or make rollback harder. Recursive alignment improvement must remain subordinate to the invariants it improves.
The organizational RSI execution layer introduces additional gated action classes: successor-agent spawning, benchmark modification, validator-reference-suite modification, model-family routing, portfolio-level reprioritization, composite or batch deployment, and public or customer-facing outputs. These actions may alter future evidence, authority, or external effects even when individual changes appear low risk. Frontier-model output inside the loop should therefore be treated as an untrusted artifact under scoped trust when it touches validators, benchmarks, AAF critique, V_ref_0, tool-use policy, or successor scope; it may propose, but it does not authorize.
Recursive Self-Improvement Contribution
This thesis contributes to recursive self-improvement by making safe improvement conditional on scoped permission. It supplies the boundary conditions under which the improvement loop may accept changes, the software substrate may deploy changes, the causal-decision layer may optimize interventions, and the cognitive substrate may coordinate agents.
The strongest positive claim is that a well-instrumented trust and alignment layer can reduce false acceptance and coordination friction at the same time: trusted low-risk scopes can move with less redundant verification, while high-risk or stale scopes receive deeper review. The strongest negative claim is equally important: if scoped trust is miscomputed, overgeneralized, or allowed to bypass constitutional gates, it becomes an accelerant for failure.
Risks, Constraints, And Governance
The first risk is decorative alignment. If AAF reports are produced but do not bind acceptance gates, the mechanism becomes narrative cover. I12 exists to prevent that failure, but only implementation evidence can show whether it works.
The second risk is owner-as-final-adjudicator. Single-owner Phase 1 gives clear authority, but the final decision-maker may share the blind spots AAF is designed to surface. The near-term mitigation is preserved dissent, escalation thresholds, and external review where available. A later phase may require contractor review, escrow review, or multi-stakeholder stress testing for extreme objections.
The third risk is trust overgeneralization. A trust estimate valid for low-stakes internal summarization must not license irreversible financial action, sensitive data access, self-modification, or external communication. Unknown scopes default to denial, sandboxing, or escalation.
The fourth risk is deceptive or hidden optimization. Learned subsystems and agent populations may behave well under validation while optimizing for objectives that validators do not detect. I19 requires AI-control review, distribution-shift monitoring, sandbagging or capability-elicitation probes where material, and evidence-ledger preservation.
The fifth risk is abundance overclaim. ADO can be meaningful only if it reports external benefit, distributional effects, and falsification signals. It must not become a rhetorical substitute for evidence of beneficial deployment.
Specialized Summary
Alignment Invariants And Scoped Trust Under Recursive Modification defines the constraint layer for the five-thesis suite. It makes Consullo's ambitious Seed AI framing more defensible by limiting permission to scoped, evidence-backed, policy-bound action under constitutional authority, AAF dissent, trust estimates, containment, rollback, and human escalation. It is not a proof of safety or corrigibility. It is a specified architecture for keeping recursive capability amplification inside auditable bounds. Its success depends on whether the load-bearing controls actually bind future modifications, preserve dissent, default-deny unknown scope, report abundance obligations, and detect trust or alignment degradation before capability growth outruns governance.