Science – Apollo Research
https://www.apolloresearch.ai/science
Apollo Research — named in the OpenAI system card for sandbagging and scheming evaluations
35
Lines added
21
Lines removed
Was read
2026-09-08
fd912479d21332f8…
Changed by
2026-09-11
116cd2f745bf013d…
35 lines added, 21 lines removed
Left column is the version read 2026-09-08, right is 2026-09-11. Unchanged stretches are elided. Both bodies are kept in full and are addressed by the hashes above.
Showing the first 400 lines of the comparison; 75 further lines are held back. This page reports a change, it is not a copy of anyone’s document — the source is linked above and is the authority for what it says.
⋯ 1 unchanged line
22
33 A Science of Scheming
44 We conduct fundamental research into the science of scheming and its potential mitigations. We also develop and run pre-deployment evaluations of frontier AI systems.
55
6−Careers
6+Research Agenda System Card Evaluations
77
88 highlights
99
1010 Featured publications
⋯ 39 unchanged lines
5050 Oops! Something went wrong while submitting the form.
5151
5252 We Need A Science of Scheming
5353
54−Science of Scheming
54+Research Agenda
55+19 January 2026
5556
56−scheming, science of scheming, covertly, misaligned goals, misalignment, deception, detect, mitigate, frontier AI, research agenda, deceptive alignment, hidden goals, covert behavior, alignment, AI safety
57+Science of Scheming
5758
5859 Measuring Reward-Seeking via Contrastive Belief Updates
59−
6060 Science of Scheming
6161
62−reward-seeking, reward seeking, grader, reinforcement learning, RL, misbehavior, reward hacking, reward hackers, OpenAI, o3, honesty, deception, model organism, chain of thought, training run, training effects, post-training, capabilities RL, alignment
62+21 July 2026
6363
64−We need 3rd party Training-Run Assessments
64+Science of Scheming
6565
66+We need 3rd party Training-Run Assessments
6667 Science of Scheming
6768
68−TRA, TRAs, trainging run assessment, scheming, checkpoints, RL environments, reward signals, SFT, post-training, evaluators, taxonomy, frontier developers, ecosystem, third party, third-party audits, external assessment, training runs, oversight, accountability, pre-deployment
69+05 July 2026
6970
70−Stress Testing Deliberative Alignment for Anti-Scheming Training
71+Science of Scheming
7172
73+Stress Testing Deliberative Alignment for Anti-Scheming Training
7274 Science of Scheming
7375
74−OpenAI, anti-scheming, anti scheming, covert actions, situational awareness, evaluation awareness, sandbagging, reward hacking, o3, o4-mini, hidden goal, chain-of-thought, CoT, spec, lying, sabotage, misalignment, safety training, training mitigation, alignment training
76+17 September 2025
7577
76−Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
78+Science of Scheming
7779
80+Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
7881 Evaluations
7982
80−CoT, chain-of-thought, monitor, monitoring, reasoning, intent to misbehave, AI safety, human language, reasoning models, transparency, faithfulness, interpretability, oversight, position paper, legibility
83+15 July 2025
8184
82−Research Note: Our scheming precursor evals had limited predictive power for our in-context scheming evals
85+Evaluations
8386
87+Research Note: Our scheming precursor evals had limited predictive power for our in-context scheming evals
8488 Evaluations
8589
8690 Notes
8791
88−precursor, precursor evals, predictive power, capability evaluations, correlation, agentic, in-context scheming, in context scheming, scheming evals, research note, predictive validity, methodology, negative result
92+03 July 2025
8993
90−More Capable Models Are Better At In-Context Scheming
94+Evaluations
9195
96+Notes
97+
98+More Capable Models Are Better At In-Context Scheming
9299 Evaluations
93100
94−capability, frontier models, scheming, in-context scheming, in context scheming, deception rates, trends, covert, capability scaling, model comparison, scaling, more capable models
101+19 June 2025
95102
96−Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations
103+Evaluations
97104
105+Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations
98106 Evaluations
99107
100108 Notes
101109
102−evaluation awareness, eval awareness, Anthropic, Claude, Sonnet, sonnet 3.7, frontier models, being evaluated, alignment evals, alignment evaluations, situational awareness, test recognition, sandbagging, meta-awareness
110+17 March 2025
103111
104−Forecasting Frontier Language Model Agent Capabilities
112+Evaluations
105113
114+Notes
115+
116+Forecasting Frontier Language Model Agent Capabilities
106117 Evaluations
107118
108−benchmarks, predictions, agent capabilities, forecasting, forecasts, extrapolation, elicitation, LLM agents, language model agents, capability prediction, future capabilities, trends
119+24 February 2025
109120
110−Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
121+Evaluations
111122
123+Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
112124 Interpretability
113125
114−APD, mechanistic interpretability, mech interp, parameters, parameter space, decomposition, attributions, superposition, neural networks, description length, interpretability research, features
126+11 February 2025
127+
128+Interpretability
115129
116130 Load more
117131
118132 Evaluations