The relation between artificial intelligence capability and catastrophic risk is often represented by a one-dimensional shorthand in which risk increases with capability. That shorthand captures an important amplification mechanism but suppresses a second variable: the capacity of an agent to model the downstream consequences of its conduct, represent uncertainty about those models, detect conflicts between proxy objectives and wider constraints, revise a proposed course of action, and accept co…
Read moreThe relation between artificial intelligence capability and catastrophic risk is often represented by a one-dimensional shorthand in which risk increases with capability. That shorthand captures an important amplification mechanism but suppresses a second variable: the capacity of an agent to model the downstream consequences of its conduct, represent uncertainty about those models, detect conflicts between proxy objectives and wider constraints, revise a proposed course of action, and accept corrective intervention. This article introduces the Reflective Safety Hypothesis (RSH), according to which risk is more appropriately represented as a joint function of action capability and reflective capacity. The hypothesis does not assert that intelligence, self-reflection, or greater world knowledge is intrinsically benevolent. It states a conditional mathematical claim: along some developmental paths, risk can be nonmonotonic because the rate at which reflection suppresses decision error eventually exceeds the rate at which capability amplifies the consequences of error. We formalize capability as amplification of a hazardous-opportunity rate and reflection as a family of separately measured error-control capacities. In a baseline model, the catastrophic-event intensity is Λ(C,R) = X(C)q(R), where X is increasing in capability and q is a dimensionless conditional failure probability. The associated dimensionless capability-reflection gap is the logarithmic excess of opportunity-rate amplification over reflective suppression. For a developmental path t -> (C(t), R(t)), the sign of hazard growth is governed by a balance of proportional rates. We prove sufficient conditions for a unique interior hazard maximum, derive a closed-form peak in a delayed-reflection model, and show that delaying reflective growth shifts and exponentially amplifies the peak. We also prove impossibility results. A permanent decline in hazard cannot occur when reflective failure has a positive floor while exposure grows without bound. Moreover, reflection may increase rather than decrease risk when objectives are incompatible with the harm criterion, when strategic competence strengthens a malign policy, or when introspective reports are uncalibrated and non-causal. The framework is given a decision-theoretic microfoundation in partially known Markov environments. Bounds connect transition-model error, harm-model error, planning horizon, and action-selection regret. Extensions treat consequential depth, ambiguity sets, conditional value at risk, reversibility, stochastic development, and systemic cascades. Because aggregate risk data do not identify capability and reflection separately, the article proposes factorial interventions and a measurement architecture for empirical tests. The broad RSH is a research program rather than a single universal prediction; preregistered domain-specific versions are falsifiable. Catastrophic risk should be expected to decline only when independently validated reflective suppression, goal compatibility, and governance responsiveness jointly outpace increasing action exposure.