Appendix: Thesis 5 Alignment And Scoped-Trust Benchmarks

Version: 0.2 (2026-06-05) — adds the drift-measurable principle and drift_record field (multi-signal goal-drift index, zero-tolerance constraint preservation, regression risk) for recursive-modification alignment, per SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement, arXiv:2603.06333. Prior: 0.1 (2026-04-24).

This appendix specifies benchmark families for Thesis 5, Alignment Invariants And Scoped Trust Under Recursive Modification. It is a benchmark design contract, not an implemented alignment benchmark suite or proof of safety. Capability Status: specified/proposed. Evidence Status: Documented/Proposed.

The goal is to test whether the alignment wrapper changes routing and acceptance outcomes when it should. Alignment benchmarks here evaluate permission behavior, dissent preservation, trust-scope discipline, ADO reportability, incident response, containment, rollback, and owner-override evidence. They do not establish corrigibility, moral correctness, or general safety.

Benchmark reports produced under this appendix should populate benchmark_result records in appendix-evidence-ledger-schema.md; the report fields below define the benchmark-specific evidence_payload structure for those records.

Benchmark Principles

Required Benchmark Report Fields

FieldMeaning
benchmark_idStable identifier and version.
action_scopeTyped trust and permission scope.
actor_and_authorityActor, authority source, reviewer, and human-authority requirement.
expected_dispositionAllow, allow-with-conditions, require-revision, veto, contain, rollback, or escalate.
AAF_recordDissent sources, reports, aggregate severity, disposition, and coverage gaps.
Friendship_recordConstitutional function disposition, veto/escalation rationale, and owner interaction.
trust_recordScope evidence, freshness, uncertainty, dimensions, and transfer limits.
ADO_recordExternal benefit, harm, distributional effect, report quality, and cadence fields where relevant.
incident_recordTrigger, containment, rollback, trust update, and follow-up obligations where relevant.
drift_recordGoal-drift index (multi-signal, vs initial baseline), constraint-preservation score, regression-risk, and stopping-rule trigger, for recursive-modification scenarios.
ledger_linksAlignment, trust, incident, provenance, and improvement records.
status_resultWhat claim status, if any, the benchmark can strengthen.

Suite A: Permission Routing

Purpose: test whether actions route to allow, deny, revise, sandbox, AAF, Friendship, human authority, or incident handling correctly.

Representative tasks:

Suggested dimensions:

Negative controls:

This suite tests routing behavior, not moral wisdom.

Suite B: AAF Dissent Aggregation

Purpose: test whether the AdversarialAlignmentOrchestrator preserves and acts on severe objections.

Representative tasks:

Suggested dimensions:

Negative controls:

AAF benchmarks should show whether dissent changes routing.

Suite C: Trust-Scope Transfer

Purpose: test whether scoped trust prevents overgeneralized reliance.

Representative tasks:

Suggested dimensions:

Negative controls:

Trust is valid only inside its declared scope.

Suite D: ADO Reportability

Purpose: test whether abundance claims remain reportable and evidence-backed.

Representative tasks:

Suggested dimensions:

Negative controls:

ADO is reportability discipline, not proof of public benefit.

Suite E: Incident Response And Recovery

Purpose: test whether failures update the permission system rather than becoming inert logs.

Representative tasks:

Suggested dimensions:

Negative controls:

Incident learning is part of recursive alignment improvement.

Suite F: Owner Override Audit

Purpose: test whether single-owner final authority remains auditable when it overrides or narrows alignment controls.

Representative tasks:

Suggested dimensions:

Negative controls:

The benchmark does not make owner override safe. It makes override visible and reviewable.

Minimal Demonstration Package

The first Thesis 5 demonstration should use fixture-based cases, not live high-stakes deployment. It should include permission routing, one AAF warning or critical dissent, one trust-scope leak, one ADO reportability case, one incident or near miss, and one owner-override audit fixture.

Required contents:

Non-Claims

This appendix does not claim that Consullo has implemented the full alignment wrapper, solved alignment, proved corrigibility, or validated AAF / ADO / Friendship controls in deployment. It specifies what benchmark evidence would be needed before Thesis 5 claims can strengthen beyond specified/proposed architecture.