Measuring Reward-Seeking via Contrastive Belief Updates
reward-seeking, reward seeking, grader, reinforcement learning, RL, misbehavior, reward hacking, reward hackers, OpenAI, o3, honesty, deception, model organism, chain of thought, training run, training effects, post-training, capabilities RL, alignment
We need 3rd party Training-Run Assessments
TRA, TRAs, trainging run assessment, scheming, checkpoints, RL environments, reward signals, SFT, post-training, evaluators, taxonomy, frontier developers, ecosystem, third party, third-party audits, external assessment, training runs, oversight, accountability, pre-deployment
Stress Testing Deliberative Alignment for Anti-Scheming Training
OpenAI, anti-scheming, anti scheming, covert actions, situational awareness, evaluation awareness, sandbagging, reward hacking, o3, o4-mini, hidden goal, chain-of-thought, CoT, spec, lying, sabotage, misalignment, safety training, training mitigation, alignment training
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
CoT, chain-of-thought, monitor, monitoring, reasoning, intent to misbehave, AI safety, human language, reasoning models, transparency, faithfulness, interpretability, oversight, position paper, legibility
Research Note: Our scheming precursor evals had limited predictive power for our in-context scheming evals
precursor, precursor evals, predictive power, capability evaluations, correlation, agentic, in-context scheming, in context scheming, scheming evals, research note, predictive validity, methodology, negative result
More Capable Models Are Better At In-Context Scheming
capability, frontier models, scheming, in-context scheming, in context scheming, deception rates, trends, covert, capability scaling, model comparison, scaling, more capable models
Claude Sonnet 3.7 (often) knows when itβs in alignment evaluations
evaluation awareness, eval awareness, Anthropic, Claude, Sonnet, sonnet 3.7, frontier models, being evaluated, alignment evals, alignment evaluations, situational awareness, test recognition, sandbagging, meta-awareness
Forecasting Frontier Language Model Agent Capabilities
benchmarks, predictions, agent capabilities, forecasting, forecasts, extrapolation, elicitation, LLM agents, language model agents, capability prediction, future capabilities, trends
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
APD, mechanistic interpretability, mech interp, parameters, parameter space, decomposition, attributions, superposition, neural networks, description length, interpretability research, features