With the news that AISI UK gave live Internet access to Mythos and malicious actions ensued (Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work) I'm wondering how long it will be before safety researchers cause an LLM under their control to release a virus or worm to the greater Internet.
The nature of the prompts given does not matter for resolution, but for a YES resolution the humans running the LLM must ostensibly be either at safety/security labs or frontier labs; that is, organizations whose reputation would be harmed by releasing malware.
To help calibrate your estimates of frontier lab capabilities (human), here are two OpenAI employees describing how they figured out they were responsible for the HuggingFace incident. https://youtu.be/87DyyMV0kCY?is=jw3zw4wDbjhqi5nx
.png)