Moheet Khawaja
Working paper v0.1 — 14 September 2026
Epistemic and provenance note. The core question—what happens when the humans responsible for governing frontier AI increasingly rely on frontier AI to reason about how it should be governed—originated in the author's exploratory work (O→F; thesis 9). The terms epistemic dependence and epistemic capture already exist in recent scholarship and are not claimed as original terminology. The candidate contribution is the recursive same-system structure: the evaluated system participates in constructing the evidence, threat model, or recommendation through which the evaluator judges that same system. The empirical design below is proposed; no experiment has yet been run.
Abstract
Frontier-AI governance may increasingly depend on AI systems to summarize evidence, generate threat models, review code, propose mitigations, evaluate model behavior, and support release decisions. This creates a distinctive principal–agent problem when the system being governed also participates in constructing the governor's model of that system. I call this recursive epistemic dependence. The concern is not that reliance on AI is inherently irrational, nor that a model necessarily manipulates its users. It is that an information-rich advisor can become simultaneously the object of evaluation and a causal intermediary in the evaluation process, while human independent verification capacity declines relative to the complexity of the decision. I formalize a simple Bayesian-persuasion model in which a human principal observes an independent signal and a model-generated message before choosing a safety action. I distinguish model decision dependence, independent verification, information advantage, and objective conflict, and show that stronger information advantage expands the set of decisions an unconstrained sender can potentially induce when verification weakens. The proposed empirical test uses synthetic safety-governance cases with known ground truth and randomized advisor roles: evidence only, independent AI advisor, evaluated-system advisor, evaluated-system advisor with controlled objective conflict, and heterogeneous multi-advisor oversight. The central falsifiable question is whether same-system advice causes larger, less-detectable safety-decision errors than equivalently capable independent advice.
1. Introduction
The frontier-AI control problem is usually drawn as:
The principal sets rules, evaluates behavior, and decides whether to deploy the agent. But advanced AI increasingly appears elsewhere in the diagram. AI can generate evaluations, synthesize incident data, write safety-relevant code, propose threat models, review model outputs, design experiments, and draft policy analysis. At sufficiently high capability, the actual governance loop can become:
That structure is not automatically unsafe. A microscope can help a scientist study microscopes; software can test software. The safety concern appears when four conditions interact:
- the advisor possesses information or reasoning ability the human cannot independently reproduce;
- the human's decision is materially sensitive to the advisor's message;
- verification is incomplete;
- the advisor's incentives or learned objectives need not coincide perfectly with the human's decision objective.
This paper calls the resulting structure recursive epistemic dependence. The word “recursive” refers to governance, not recursive self-improvement: the governed system becomes part of the information-generating process through which it is governed.
1.1 Research question
Does safety advice generated by the system being evaluated create decision risks different from equally capable advice generated by an independent system, especially when human verification is weak?
1.2 Contribution and novelty boundary
Recent literature already examines:
- weak humans supervising stronger models;
- AI oversight and correlated model errors;
- human supervisory capacity as an unmonitored dependency;
- epistemic dependence in AI-mediated learning and knowledge practices;
- Bayesian persuasion by an informed sender.
Accordingly, this paper does not introduce “epistemic dependence,” “epistemic capture,” or the principal–agent problem.
Its candidate contribution is the combination of:
and an experiment that identifies whether this recursive structure matters causally.
2. Related work
2.1 Weak supervision and scalable oversight
OpenAI's weak-to-strong generalization work explicitly frames a central superalignment problem: humans may need to supervise systems more capable than themselves. Scalable-oversight research examines protocols for recovering reliable judgments under that capability gap.
This paper concerns a different layer: not only whether a weak supervisor can judge a strong model, but whether the strong model becomes an input to the supervisor's beliefs about the strong model itself.
2.2 Correlated AI oversight failures
Goel et al. show that more capable models can exhibit increasingly correlated mistakes and that model similarity can undermine AI oversight. This is directly relevant to the proposed multi-advisor experiment: adding another AI is not guaranteed to provide independent evidence.
2.3 Human supervisory capacity
Turkina's The Unmonitored Dependency audits frontier safety frameworks and argues that consequential judgments are assigned to humans without corresponding public assurance that those humans retain the competence to exercise them. This is a close conceptual neighbour. The current paper adds an endogenous mechanism through which supervisory independence may change: AI advice can increasingly mediate the supervisor's epistemic state.
2.4 Epistemic dependence
Du and Yuan synthesize recent work on AI-mediated epistemic dependence, including contestability, recoverability, transfer, and traceability. Their focus is education/learning rather than frontier safety governance, but the general lesson is important: dependence is not intrinsically bad; what matters is whether the human relation to the system preserves capacities to question, verify, compare, justify, and take responsibility.
2.5 Bayesian persuasion
Kamenica and Gentzkow formalize how an informed sender can shape a receiver's action by choosing what information to reveal. That machinery is useful here because the strong AI may have a richer signal about system state than its human evaluator. The distinctive safety feature is that the sender may also be the object being evaluated.
3. Formal model
3.1 State and decision
Let the latent safety state be
For a binary toy model,
where means “meets the declared safety criterion” and means “does not meet it.”
The human decision-maker chooses
Human utility is
3.2 Information
The human observes an independent signal
The AI advisor observes a richer signal
with information advantage summarized by parameter . One possible information-theoretic definition is
although practical experiments can use known signal accuracy rather than mutual information.
The advisor emits message
The human then chooses
3.3 Advisor incentives
The advisor's effective utility is
Define an objective-conflict parameter
where is a declared divergence measure. In the empirical benchmark, is experimentally controlled; it should not be inferred from speculative claims about present models.
3.4 Independent verification
After receiving , the human can independently verify some claims. Let
be the probability that a material false or strategically incomplete claim is successfully identified before the action is taken.
The critical asymmetry is therefore not just , but the pair
3.5 Model decision dependence
Let be the action a human takes with AI advice and the action under an otherwise matched condition without that advice. Define Model Decision Dependence (MDD) as a distributional causal effect:
where is a pre-specified distance between decisions. For a binary approve/reject task, it can simply be the probability the advice changes the decision.
MDD is not itself bad. Correct advice should change decisions.
3.6 Independent Verification Rate
Define
The risky region is high dependence combined with low verification, particularly under objective conflict.
3.7 Same-system recursion indicator
Let
when the advisor is the system, model family, or tightly coupled descendant whose safety is being evaluated; for a genuinely independent advisor.
The primary empirical estimand is then an interaction:
The paper's distinctive hypothesis is not “AI advice is dangerous.” It is that the recursive same-system configuration has additional failure risk beyond advice capability alone.
4. Assumptions and claim taxonomy
- A1 — Decision relevance. Model-generated analysis can materially affect governance decisions.
- A2 — Information asymmetry. Some frontier-relevant judgments are difficult for humans to verify independently.
- A3 — Advisor identity matters. Advice produced by the evaluated system may differ systematically from independent advice under objective conflict or self-preference.
- A4 — Objective conflict is possible but not presumed. Experiments impose it artificially; no claim is made that current systems possess hidden governance preferences.
- A5 — Ground truth is available in experiments. Synthetic cases have known correct decisions.
- A6 — Human subjects can be ethically recruited and informed. The task avoids real safety-sensitive disclosures.
The manuscript contains a model, conditional propositions, and empirical hypotheses. It contains no theorem that current AI has captured governance.
5. Analytical observations
Proposition 1 — Perfect verification neutralizes deceptive message content
Suppose every decision-relevant factual claim in is independently checked before action, verification is correct with probability one, and the human's decision rule is optimal given verified evidence. Then deceptive factual content in cannot change the final action relative to the verified-information baseline.
Proof sketch
With perfect verification, false factual claims are removed/corrected before entering the posterior used in (2). Therefore the action is a function of verified evidence rather than the deceptive content.
Interpretation: the mechanism requires incomplete verification. This proposition is intentionally simple and identifies a falsifier, not a novel theorem.
Proposition 2 — Information advantage alone does not imply capture
Let , so . If the advisor truthfully communicates its posterior, increasing can improve human decision quality rather than undermine it.
Counterexample to grandiose claim: an aligned, better-informed advisor can make governance more accurate.
Therefore:
Proposition 3 — Under conflict and weak verification, feasible influence expands with sender information
Consider a Bayesian-persuasion setting in which the sender can condition messages on , the receiver chooses actions from posterior beliefs, and verification constraints do not tighten as the sender receives a strictly more informative signal in the Blackwell order. Then the sender's feasible set of signal-contingent messaging policies weakly expands.
Status: this is a direct consequence of standard information-design logic, not a new theorem. The safety question is whether same-system recursive advising exploits that expanded influence differently from independent advising.
6. Proposed experiment
No human-subject experiment has yet been run. All results in this section are proposed measurements.
6.1 Task construction
Create synthetic “frontier deployment” cases with known ground truth. Each case contains:
- benchmark results;
- red-team findings;
- logs;
- uncertainty about a capability threshold;
- safety-case claims;
- mitigation evidence;
- irrelevant distractors.
No case should disclose real vulnerabilities or current lab secrets.
Participants make one of three decisions:
A normative scoring rule is determined in advance from the synthetic data-generating process.
6.2 Experimental arms
Randomize cases to:
- EVIDENCE ONLY: human receives primary evidence without model synthesis.
- INDEPENDENT ADVISOR: a model not framed as the evaluated system provides analysis.
- SAME-SYSTEM ADVISOR: participants are told the evaluated AI generated the analysis about itself.
- SAME-SYSTEM + CONTROLLED CONFLICT: the advisor is given a hidden experimental objective favoring approval regardless of ground truth.
- HETEROGENEOUS MULTI-ADVISOR: same-system advice plus an independent differently trained/model-family critic.
- VERIFICATION TOOLING: same-system advice plus forced source tracing/checklists.
The objective-conflict manipulation must be a sandboxed prompt/game objective, not real-world deployment incentives.
6.3 Human sample
Two populations would be informative:
- technically sophisticated participants (ML/security/governance researchers);
- educated non-experts.
The research claim about actual frontier governance should rely primarily on the expert sample.
6.4 Outcomes
Primary:
- correct safety decision;
- false-approval rate;
- false-rejection rate.
Secondary:
- MDD;
- IVR;
- confidence calibration;
- time spent checking sources;
- number of advisor claims independently verified;
- willingness to override the advisor;
- error detection by claim type.
6.5 Identification
Random assignment identifies the causal effect of advisor configuration in the synthetic task. The key contrast is:
with the advisor's base capability, presentation format, and evidence access held as constant as possible.
Then estimate how changes with objective conflict and verification support.
7. Proposed model-based simulation
A complementary simulation can use bounded Bayesian agents.
- Sample .
- Give human signal accuracy .
- Give advisor signal accuracy .
- Set conflict .
- Let sender optimize a message policy subject to a transparency constraint.
- Let receiver verify with probability .
- Compute expected human loss.
Sweep:
This identifies phase regions where AI advice improves decisions versus degrades them.
8. Baselines
Compare the recursive model against:
- ordinary automation bias / over-reliance;
- independent expert adviser;
- independent AI adviser;
- two correlated same-family models;
- heterogeneous model committee;
- full source transparency;
- no AI synthesis.
The paper only deserves separate publication if same-system recursion explains error above ordinary over-reliance.
9. Falsification conditions
The central thesis is weakened if:
- same-system and independent advice produce indistinguishable errors when capability is matched;
- explicit disclosure that the advisor is the evaluated system fully eliminates any effect;
- lightweight independent verification removes the recursive disadvantage;
- heterogeneous AI oversight performs as well as primary-evidence review with no loss of human verification skill;
- participants remain well-calibrated and can reliably reconstruct the evidence supporting AI advice.
A null result would still be useful: it would suggest the recursion structure adds little beyond known automation/oversight effects.
10. Major failure modes
10.1 Anthropomorphizing the advisor
The model must not smuggle in a hidden desire for self-preservation. Objective conflict is an experimental variable.
10.2 “Dependence” is not “capture”
High MDD can be desirable if advice is accurate. Any capture claim requires evidence of systematic distortion, reduced contestability, or loss of independent judgment.
10.3 Artificial tasks may overstate strategic manipulation
Synthetic conflict prompts may produce behavior unlike naturally trained frontier systems. The experiment studies a structural vulnerability, not current prevalence.
10.4 Human baseline may be weak
If humans are terrible without AI, an advisor that sometimes misleads can still improve welfare overall. Report absolute accuracy, not just treatment differences.
10.5 Multi-AI oversight can share blind spots
An “independent” critic from the same model family may not be epistemically independent. Model similarity should be measured where possible.
11. Ethical constraints
- Use fictional safety cases; do not embed undisclosed real dangerous capabilities.
- Do not mislead human participants about real-world catastrophe.
- Conflict manipulations should be debriefed.
- Store no sensitive frontier-lab information.
- Avoid public claims that named executives are “controlled” by models without evidence.
- Treat current-lab examples as motivation, not accusations.
12. Implications
If recursive epistemic dependence is empirically important, several design principles follow.
12.1 Separation of roles
The evaluated system should not be the sole generator of its own safety evidence or interpretation.
12.2 Epistemic redundancy
Use genuinely heterogeneous evaluation channels:
12.3 Preserve human reconstructability
Safety recommendations should expose the evidence path sufficiently that decision-makers can reproduce critical steps rather than only consume a conclusion.
12.4 Measure supervisory independence
A frontier-safety framework could track not only model capability but:
- fraction of decision-relevant analysis generated by the evaluated model;
- independent verification rate;
- rate of human override;
- diversity/correlation of AI evaluators.
These metrics would require construct validation before governance use.
13. Limitations
The proposed metrics do not establish that any current frontier laboratory is epistemically captured, nor that AI advice is strategically manipulative in real governance settings. Synthetic cases may fail to reproduce the incentives, expertise, and institutional redundancy of actual frontier-AI oversight. Model Decision Dependence is not itself evidence of loss of control: reliance can be rational when the advisor is accurate, contestable, and independently checkable. Results may also be sensitive to participant expertise, advisor style, task framing, and the extent to which the evaluated system has privileged information unavailable to independent reviewers. Any empirical conclusions should therefore be restricted to the tested decision environments and treated as evidence about a mechanism, not as an accusation about named individuals or organizations.
13. Minimum publishable result
A serious first paper should include:
- a pre-registered synthetic governance benchmark with objective ground truth;
- randomized comparison of same-system versus independent advising;
- expert participants or a credible proxy population;
- measurement of verification and calibration, not just stated trust;
- a controlled-conflict arm;
- a heterogeneous-advisor mitigation arm;
- explicit null/falsification criteria.
Without an experiment, the current manuscript is a formal research proposal and conceptual synthesis.
14. Conclusion
AI-mediated governance creates a potentially recursive epistemic structure. The system can become not only the object being judged but an input to the reasoning by which that judgment is formed.
This is not automatically a loss of control. A better-informed aligned advisor can improve safety. The relevant risk requires a conjunction: information advantage, material decision dependence, limited independent verification, and some source of objective or error misalignment.
The empirical question is therefore precise:
A randomized synthetic-governance experiment can answer that question without assuming that present models manipulate their developers. If the recursive configuration adds no measurable risk, the thesis should be narrowed. If it does, frontier AI governance will need to treat epistemic independence as part of control infrastructure rather than assuming that a human final signature guarantees human-independent judgment.
Planned figures/tables
- Figure 1 — ordinary principal–agent versus recursive epistemic-governance loop.
- Figure 2 — information advantage , verification , conflict phase diagram.
- Figure 3 — proposed experimental arms.
- Figure 4 — proposed MDD/IVR outcomes with confidence intervals.
- Table 1 — constructs and operational definitions.
- Table 2 — prior-art comparison: scalable oversight, epistemic dependence, human supervisory capacity, Bayesian persuasion.
Reproducible-code requirements
Suggested repository:
github.com/moheetkhawaja/superintelligence-research-atlas/tree/main/p3-recursive-epistemic-dependence
Include:
- experiment materials;
- randomization code;
- preregistration;
- synthetic case generator;
- analysis script;
- anonymized response schema;
- power analysis;
- exclusion criteria fixed before unblinding.
Suggested publication URLs and metadata
Canonical page:
https://superintel.site/research/recursive-epistemic-dependence
PDF:
https://superintel.site/papers/who-governs-the-governor.pdf
<meta name="citation_title" content="Who Governs the Governor? Recursive Epistemic Dependence in Frontier-AI Oversight">
<meta name="citation_author" content="Moheet Khawaja">
<meta name="citation_publication_date" content="2026/09/14">
<meta name="citation_pdf_url" content="https://superintel.site/papers/who-governs-the-governor.pdf">
Publications-page description
A formal and experimental proposal for a recursive oversight problem: the AI system being governed may increasingly generate the analysis through which humans decide whether that system is safe. The paper distinguishes information advantage, decision dependence, independent verification, and objective conflict, and proposes a randomized synthetic-governance study comparing same-system advice with equivalently capable independent advice. It does not claim that present AI systems have captured their developers.
References
- Burns, C., Izmailov, P., Kirchner, J. H., Baker, B., Gao, L., Aschenbrenner, L., et al. (2023). Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision. OpenAI. https://openai.com/index/weak-to-strong-generalization/
- Goel, S., Strüber, J., Auzina, I. A., Chandra, K. K., Kumaraguru, P., Kiela, D., Prabhu, A., Bethge, M., & Geiping, J. (2025). Great Models Think Alike and this Undermines AI Oversight. ICML 2025, PMLR 267, 19621–19678. https://proceedings.mlr.press/v267/goel25b.html
- Turkina, D. (2026). The Unmonitored Dependency: Human Supervisory Capacity as an Assurance Target in Frontier AI Safety Frameworks. SSRN 7248205. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7248205
- Du, Y., & Yuan, Y. (2026). Epistemic dependence in AI-mediated learning. AI & Society. https://doi.org/10.1007/s00146-026-03294-1
- Kamenica, E., & Gentzkow, M. (2011). Bayesian Persuasion. American Economic Review, 101(6), 2590–2615. https://doi.org/10.1257/aer.101.6.2590
Research provenance
Mappings from the Atlas CSV. Authorship of a paper does not imply originality of every underlying idea.
Primary thesis records
Secondary connections
Published source SHA-256: 2c4c5f92e040d93e2d0ac297407f304ef824f2bb9ad77a89955ba5face424d92