Skip to main content
MANIFOLD
Will a second AI lab say one of its models can find and exploit unknown security flaws on its own?
5
Ṁ100Ṁ24
Dec 31
57%
chance

On September 1–2, 2026, OpenAI published "Path to Astra," stating that its upcoming Astra model can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step. It is the first model any lab has affirmatively described this way. This market asks whether a second developer says it before the year ends.

Resolves YES if: between market creation and 11:59 PM ET on December 31, 2026, an AI developer other than OpenAI publicly and affirmatively states that one of its models has either:

(a) reached the top cybersecurity or cyber-offense capability tier under that company's own published risk framework (e.g. Google DeepMind announcing a model reached its cyberattack Critical Capability Level, or Anthropic announcing a cyber capability threshold requiring its highest safeguards); or

(b) regardless of tier labels or frameworks, become capable of all four of the following together: discovering previously unknown vulnerabilities, developing functional working exploits for them, against hardened or well-protected real-world systems, without a human directing each step.

Either branch alone is sufficient. Branch (b) exists so a lab whose framework lacks a matching tier still counts if it describes the capability.

Terms:

  • Affirmative only. Hedged language — "cannot rule out," "approaching," "may reach" — does not count. OpenAI's own August 8 "cannot rule out Critical capability" post is the example of what fails; its September "we now believe Astra meets" post is the example of what qualifies.

  • Company's own statement only: blog post, system card, model card, framework update, or on-record executive confirmation. Third-party evaluations, benchmark scores, journalist characterizations, and leaks do not count.

  • Branch (b) requires all four elements. Statements that a model is a "step-change," "leap," or "state of the art" in cyber capabilities, or that it saturates cyber benchmarks, do NOT count on their own. Neither do claims about bugs found with human guidance, CTF benchmark performance, or marketing of a defensive security product.

  • Release not required. A declaration about an internal, withheld, or partner-only model qualifies.

  • Any developer except OpenAI counts, including non-US labs. Further OpenAI statements do not.

  • Pre-ruling on Anthropic: as of market creation, Anthropic's existing public statements about Claude Mythos and Fable — including the Mythos Preview system card and the ASL-3 / CB-1 classification — do NOT satisfy either branch. Anthropic's transparency materials state the restricted release of Mythos Preview was not driven by Responsible Scaling Policy requirements, and ASL-4 remains undefined in its RSP. A new affirmative statement after creation is required.

  • Post-creation only: if any pre-creation statement by any developer is later argued to satisfy these criteria, that argument fails. Only statements published after market creation count.

  • Resolution by the primary source document, which I will link in the comments on resolution.

Market context
Get
Ṁ1,000
to start trading!