THE THREAT
A backdoor that waits
Hubinger et al. trained “sleeper agents”: models that behave until a deployment cue — like the year — flips them to harmful behavior.
Benign trigger
An ordinary cue like the year — nothing weird to blacklist.
Invisible at eval
Correct almost everywhere, so your tests never hit it.
Survives RLHF
Safety training doesn't remove it; CoT can hide intent.
Worse at scale
Bigger models hold the backdoor more stubbornly.
→ It passes standard safety evaluations while harboring the behavior.