31 juli 2026
68 min
Anthropic concluded in the April Mythos Preview alignment risk update that the model “does not possess any unknown propensities that would increase alignment risk.” The report argues that if Mythos Preview were coherently misaligned[1], it likely would have been detected by the assessment (following Anthropic, I will call this “reliability of the assessment”[3]).
While I agree with the report on the above bottom-line conclusions (substantially on priors), I think there are gaps in its argument which weaken the current assessment and might invalidate future assessments. In particular, the report often uses weak evidence to justify reliability.
The report gives fairly weak experimental evidence for Mythos Preview having insufficient capabilities to evade monitoring. The model is plausibly often eval-aware and underelicited in the relevant capability evaluations. So, it might silently sandbag if coherently misaligned, or unintentionally underperform if otherwise misaligned.
This limitation is important: one could argue that lack of covert capabilities for sophisticated sabotage (a subset of the capabilities I discuss here) is the single most load bearing argument in alignment risk reports.
Authors of the report could have made calibrated guesses about Mythos Preview's covert capabilities, especially for covert sabotage, based on other [...]
---
Outline:
(03:16) How reliability fits into the overall safety argument
(05:22) Reliability claims by AI companies
(05:59) Reliability claims by external evaluators
(06:37) Alignment assessments are less reliable than developers claim
(07:21) 1: Measuring capabilities to covertly undermine alignment assessments
(10:15) Issues with evaluation awareness
(13:39) Issues with underestimating covert capabilities
(16:45) Issues with sandbagging rule-out
(19:20) 2: Stress-testing alignment assessments with auditing games
(20:26) An auditing failure with Mythos
(22:21) AuditBench results
(24:17) 3: Conditioning on misalignment should make us think that certain covert capabilities are better than expected
(26:20) Bottom line on the strength of current alignment assessments
(29:07) Conclusion
(29:44) Appendix:
(29:47) Why I focus on motive / alignment assessments in alignment risk reports
(30:59) Auditability vs. Trustedness
(33:27) More reliability claims by developers and third party evaluators
(33:43) Mythos Alignment Risk Update
(35:01) Opus 4.6 Sabotage Risk Report
(35:46) GPT 5.5 System card
(36:49) Muse Spark system card
(37:36) Mythos Alignment Risk Update, safety arguments against sandbagging
(38:40) UK AISI evaluations for Opus 4.7
(40:01) Past auditing games by Anthropic
(42:24) Anti-auditing capability measurements
(43:51) Conditioning on coherent misalignment updates us on certain covert capabilities
The original text contained 92 footnotes which were omitted from this narration.
---
First published:
July 31st, 2026
Source:
https://blog.redwoodresearch.org/p/sota-alignment-assessments-dont-strongly
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Lyssna på fler avsnitt från
Redwood Research Blog
Visar 1–10 av 124 avsnitt
27 augusti 2026
9 min
12 augusti 2026
20 min
27 juli 2026
44 min
26 juli 2026
9 min
25 juli 2026
11 min
24 juli 2026
73 min
23 juli 2026
10 min
2 juli 2026
25 min
18 juni 2026
18 min
10 juni 2026
10 min