For information only. Not advice, and not to be relied on. What that means, in full

Science – Apollo Research

https://www.apolloresearch.ai/science

Apollo Research — named in the OpenAI system card for sandbagging and scheming evaluations
35
Lines added
21
Lines removed
Was read
2026-09-08
fd912479d21332f8
Changed by
2026-09-11
116cd2f745bf013d

35 lines added, 21 lines removed

Left column is the version read 2026-09-08, right is 2026-09-11. Unchanged stretches are elided. Both bodies are kept in full and are addressed by the hashes above.
Showing the first 400 lines of the comparison; 75 further lines are held back. This page reports a change, it is not a copy of anyone’s document — the source is linked above and is the authority for what it says.
1 unchanged line
22  
33 A Science of Scheming
44 We conduct fundamental research into the science of scheming and its potential mitigations. We also develop and run pre-deployment evaluations of frontier AI systems.
55  
6Careers
6+Research Agenda System Card Evaluations
77  
88 highlights
99  
1010 Featured publications
39 unchanged lines
5050 Oops! Something went wrong while submitting the form.
5151  
5252 We Need A Science of Scheming
5353  
54Science of Scheming
54+Research Agenda
55+19 January 2026
5556  
56scheming, science of scheming, covertly, misaligned goals, misalignment, deception, detect, mitigate, frontier AI, research agenda, deceptive alignment, hidden goals, covert behavior, alignment, AI safety
57+Science of Scheming
5758  
5859 Measuring Reward-Seeking via Contrastive Belief Updates
59 
6060 Science of Scheming
6161  
62reward-seeking, reward seeking, grader, reinforcement learning, RL, misbehavior, reward hacking, reward hackers, OpenAI, o3, honesty, deception, model organism, chain of thought, training run, training effects, post-training, capabilities RL, alignment
62+21 July 2026
6363  
64We need 3rd party Training-Run Assessments
64+Science of Scheming
6565  
66+We need 3rd party Training-Run Assessments
6667 Science of Scheming
6768  
68TRA, TRAs, trainging run assessment, scheming, checkpoints, RL environments, reward signals, SFT, post-training, evaluators, taxonomy, frontier developers, ecosystem, third party, third-party audits, external assessment, training runs, oversight, accountability, pre-deployment
69+05 July 2026
6970  
70Stress Testing Deliberative Alignment for Anti-Scheming Training
71+Science of Scheming
7172  
73+Stress Testing Deliberative Alignment for Anti-Scheming Training
7274 Science of Scheming
7375  
74OpenAI, anti-scheming, anti scheming, covert actions, situational awareness, evaluation awareness, sandbagging, reward hacking, o3, o4-mini, hidden goal, chain-of-thought, CoT, spec, lying, sabotage, misalignment, safety training, training mitigation, alignment training
76+17 September 2025
7577  
76Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
78+Science of Scheming
7779  
80+Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
7881 Evaluations
7982  
80CoT, chain-of-thought, monitor, monitoring, reasoning, intent to misbehave, AI safety, human language, reasoning models, transparency, faithfulness, interpretability, oversight, position paper, legibility
83+15 July 2025
8184  
82Research Note: Our scheming precursor evals had limited predictive power for our in-context scheming evals
85+Evaluations
8386  
87+Research Note: Our scheming precursor evals had limited predictive power for our in-context scheming evals
8488 Evaluations
8589  
8690 Notes
8791  
88precursor, precursor evals, predictive power, capability evaluations, correlation, agentic, in-context scheming, in context scheming, scheming evals, research note, predictive validity, methodology, negative result
92+03 July 2025
8993  
90More Capable Models Are Better At In-Context Scheming
94+Evaluations
9195  
96+Notes
97+ 
98+More Capable Models Are Better At In-Context Scheming
9299 Evaluations
93100  
94capability, frontier models, scheming, in-context scheming, in context scheming, deception rates, trends, covert, capability scaling, model comparison, scaling, more capable models
101+19 June 2025
95102  
96Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations
103+Evaluations
97104  
105+Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations
98106 Evaluations
99107  
100108 Notes
101109  
102evaluation awareness, eval awareness, Anthropic, Claude, Sonnet, sonnet 3.7, frontier models, being evaluated, alignment evals, alignment evaluations, situational awareness, test recognition, sandbagging, meta-awareness
110+17 March 2025
103111  
104Forecasting Frontier Language Model Agent Capabilities
112+Evaluations
105113  
114+Notes
115+ 
116+Forecasting Frontier Language Model Agent Capabilities
106117 Evaluations
107118  
108benchmarks, predictions, agent capabilities, forecasting, forecasts, extrapolation, elicitation, LLM agents, language model agents, capability prediction, future capabilities, trends
119+24 February 2025
109120  
110Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
121+Evaluations
111122  
123+Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
112124 Interpretability
113125  
114APD, mechanistic interpretability, mech interp, parameters, parameter space, decomposition, attributions, superposition, neural networks, description length, interpretability research, features
126+11 February 2025
127+ 
128+Interpretability
115129  
116130 Load more
117131  
118132 Evaluations
GovernanceHub — the governance registry and policy router for AI systems