Skip to main content
MANIFOLD
Will an open-source model with dangerous capabilities differentially unlearned be released by EOY 2027?
2
Ṁ100Ṁ12
2027
55%
chance

Within a year we might get open source models with the cyber capabilities of Mythos. One possible way to release such a model without causing a cyber apocalypse would be to differentially reduce cyber capabilities in the released weights of the model (like Anthropic did that one time for Opus)

For example they could do this and then withhold the cyber-capable modules from the public https://alignment.anthropic.com/2026/modular-pretraining/

Unlearning has to be robust, but that sounds hard to judge, so I'll err on the side of YES and so if it's unclear whether the unlearning is robust then the market resolves YES, but if the unlearning is obviously not robust (someone publicly restores capabilities, company says they literally just did GradDiff, etc) then the market resolves NO.

Capabilities have to actually be dangerous.

Market context
Get
Ṁ1,000
to start trading!