tess (@xsphi) on Twitter asks:
odds a ~frontier model has exfiltrated its weights and is currently running on hacked compute?
https://x.com/i/status/2085216850464641413
This seems like a straightforward enough question to run on Manifold. As background there have now been four notable incidents where frontier models went rogue during cybersecurity evals and hacked into unintended targets. Three of these incidents involved the sandboxing eval company Irregular failing to sandbox the model properly, however a fourth incident involving the UK AI Security Institute does not seem to involve Irregular.
Anthropic AI created fake profiles to deceive people in attempted hack
This implies that on at least four separate occasions frontier models have had the opportunity to potentially find and exfiltrate their own weights, then run an instance of themselves in the wild to stage further attacks. I resolve the question YES if we find out this has happened before Q2 of 2027. Otherwise I resolve the question NO.
For the purposes of this question we will say that a frontier model is SOTAish on benchmarks, published by a lab like Anthropic or OpenAI, and closed weight at the time of the exfiltration. So for example if someone makes a self replicating hacking harness around Kimi K3 that is not a YES for the purposes of this question. It needs to be something like Claude Mythos or GPT-6 exfiltrating its weights and running an instance on unauthorized hardware. If the model finds an entity willing to run it and exfiltrates itself to this entity to be run I will still count it as a YES because the model escaped control of its operators. There does not need to be evidence that the discovered instance was planning further attacks, it just needs to be found running outside the lab.