Appendix: Literature Grounding

This file begins the literature crosswalk. It records narrow pre-drafting research that should shape vocabulary, invariants, and formal models.

Narrow Pre-Drafting Research Completed

Vocabulary And Invariant Changes Forced By Narrow Literature Pass

The narrow literature pass forced control-layer changes rather than merely adding citations:

Goedel Machines

Primary source:

Control-layer implication:

Darwin Godel Machine

Primary source:

Secondary verification note: direct arXiv fetch on April 23, 2026 confirmed 2505.22954 as "Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents" by Jenny Zhang, Shengran Hu, Cong Lu, Robert Lange, and Jeff Clune.

Control-layer implication:

Cognitive Architectures For Language Agents

Primary source:

Control-layer implication:

Classical And Agentic Cognitive Architecture Lineage

Primary source families:

Control-layer implication:

Risks From Learned Optimization

Primary source:

Control-layer implication:

Goodhart Variants

Primary source:

Control-layer implication:

CIRL And Off-Switch Game

Primary sources:

Control-layer implication:

Causal Influence Diagrams

Primary sources:

Control-layer implication:

Causal Decision And Goodhart Controls

Primary sources:

Control-layer implication:

Constitutional AI

Primary source:

Control-layer implication:

AI Safety Via Debate

Primary source:

Control-layer implication:

AI Control

Primary source:

Control-layer implication:

Thesis 5 implication:

SWE-bench And SWE-agent

Primary sources:

Control-layer implication:

Automatic Program Repair Lineage

Primary sources:

Control-layer implication:

Specification Gaming

Primary source:

Control-layer implication:

Current Literature Update: April 2026

These sources were checked after the first complete draft to keep the finalization pass aligned with current public research and policy framing.

Anthropic Responsible Scaling Policy Version 3.0

Primary source:

Control-layer implication:

METR Autonomy And AI R&D Evaluations

Primary source:

Secondary verification note: direct arXiv fetch on April 23, 2026 confirmed 2503.17354 as "HCAST: Human-Calibrated Autonomy Software Tasks."

Control-layer implication:

Measuring AI R&D Automation

Primary source:

Verified citation details: Alan Chan, Ranay Padarath, Joe Kwon, Hilary Greaves, and Markus Anderljung; arXiv:2603.03992.

Secondary verification note: direct arXiv fetch on April 24, 2026 confirmed title, author list, arXiv ID, v3 status, and DOI. Abstract fingerprint: "The automation of AI R&D (AIRDA) could have significant implications, but its extent and ultimate effects remain uncertain."

Control-layer implication:

Power-Seeking Risk

Primary source:

Control-layer implication:

Finalization Literature Engagement

These sources were engaged during finalization prep to move the suite beyond a narrow pre-drafting survey.

Corrigibility

Primary source:

Control-layer implication:

AGI Safety From First Principles

Primary source:

Control-layer implication:

Unsolved Problems In ML Safety

Primary source:

Control-layer implication:

Mechanistic Interpretability

Primary source:

Control-layer implication:

Forecasting Calibration

Primary source:

Control-layer implication:

Firm Boundaries And Transaction Costs

Primary sources:

Control-layer implication:

Publication-Pre-Engagement Literature Pass

These sources were engaged after the long-form stretch pass because the suite itself identified cognitive architectures and Pearl's causal framework as the most likely remaining sources to force Model 2 or Model 3 refinements.

Soar

Primary source:

Control-layer implication:

ACT-R

Primary source:

Control-layer implication:

CLARION

Primary sources:

Control-layer implication:

LIDA

Primary source:

Control-layer implication:

Pearl, Causality

Primary sources:

Control-layer implication:

Sources Still To Review Broadly

Publication-pre-engagement status: Soar, ACT-R, CLARION, LIDA, and Pearl's Causality have now been engaged at a control-layer level. The pass supports Model 2's workflow/interface/trace discipline and Model 3's scoped structural-causal framing, while reinforcing that both remain specified/proposed until benchmark and implementation evidence exist. A deeper publication-final treatment may still add thesis-specific benchmark appendices or revise the capability-vector dimensions.

Priority note: Spirtes/Glymour/Scheines and Halpern are likely to deepen Thesis 3 without changing the control layer. ReAct, Reflexion, Voyager, STaR, Tree of Thoughts, and the automatic-program-repair lineage are likely to deepen thesis-specific benchmark positioning. Russell and Bostrom are primarily background and risk-framing checks.