Skip to main content
MANIFOLD
AGI When? [High Quality Turing Test]
1.1k
Ṁ30kṀ1.1m
2050
2,036
expected

This market resolves to the year in which an AI system exists which is capable of passing a high quality, adversarial Turing test. It is used for the Big Clock on the manifold.markets/ai page.

The Turing test, originally called the imitation game by Alan Turing in 1950, is a test of a machine's ability to exhibit intelligent behaviour equivalent to, or indistinguishable from, that of a human.

For proposed testing criteria, refer to this Metaculus Question by Matthew Barnett, or the Longbets wager between Ray Kurzweil and Mitch Kapor.

As of market creation, Metaculus predicts there is an ~88% chance that an AI will pass the Longbets Turing test before 2030, with a median community prediction of July 2028.

Manifold's current prediction of the specific Longbets Turing test can be found here:

/dreev/will-ai-pass-the-turing-test-by-202

This question is intended to determine the Manifold community's median prediction, not just of the Longbets wager specifically but of any similiarly high-quality test.


Additional Context From Longbets:

One or more human judges interview computers and human foils using terminals (so that the judges won't be prejudiced against the computers for lacking a human appearance). The nature of the dialogue between the human judges and the candidates (i.e., the computers and the human foils) is similar to an online chat using instant messaging.

The computers as well as the human foils try to convince the human judges of their humanness. If the human judges are unable to reliably unmask the computers (as imposter humans) then the computer is considered to have demonstrated human-level intelligence.

Additional Context From Metaculus:

This question refers to a high quality subset of possible Turing tests that will, in theory, be extremely difficult for any AI to pass if the AI does not possess extensive knowledge of the world, mastery of natural language, common sense, a high level of skill at deception, and the ability to reason at least as well as humans do.

A Turing test is said to be "adversarial" if the human judges make a good-faith attempt, in the best of their abilities, to successfully unmask the AI as an impostor among the participants, and the human confederates make a good-faith attempt, in the best of their abilities, to demonstrate that they are humans. In other words, all of the human participants should be trying to ensure that the AI does not pass the test.

Note: These criteria are still in draft form, and may be updated to better match the spirit of the question. Your feedback is welcome in the comments.

Market context
Get
Ṁ1,000
to start trading!
Sort by:

Can anyone explain why the distribution here is so out of sync with the related market linked in the description? /dreev/will-ai-pass-the-turing-test-by-202 Does it make sense that this market should imply such lower chances (currently ~32%) of passing by the end of 2029? (If anything, the corresponding percentage here should be higher because the resolution criteria are broader, no?)

@TimWho Good question. I'm thinking the Longbets wager might resolve to "YES AGI" on a technicality. This market I'm trusting to, as it says, hew to the spirit.

Consider: If Kurzweil and Kapor set up their long-form adversarial Turing test today, it's not obvious to me who would officially win. But it's obvious to everyone (or at least it's the general consensus -- all the markets on AGI in 2026 seem to be trading near zero) that we do not have AGI today.

@dreev Your basis seems to be the short title, but literally nothing else in the description suggests any definition of AGI other than "capable of passing a high quality, adversarial Turing test."

@TimWho Yeah, the description could really use some work. It was written 2.5 years ago now. Things that to me point to the possibility of Longbets resolving YES without triggering resolution of this market:

  1. The title, as you said

  2. Using it as the Big Clock on the AI landing page

  3. The phrase "not just of the Longbets wager specifically"

  4. Quotes that suggest that back in 2024 we still presumed that passing a sufficiently hard Turing test would require AGI

  5. The final note that the criteria may be updated to better match the spirit of the question.

@dreev It is obvious to me that if you held the test today and if the judges were selected carefully and were actually trying to distinguish the human and AI, rather than trying to answer the question, "does this seem like real intelligence to me," Kurzweil would lose. And the same thing is likely to happen even if they hold it in 2029.

But it is very plausible that Kapor will concede based on the spirit of the bet or similar, before that happens.

@dreev Re your #3, the full context is "not just of the Longbets wager specifically but of any similiarly high-quality test." Re #2, the description begins with "This market resolves to the year in which an AI system exists which is capable of passing a high quality, adversarial Turing test. It is used for the Big Clock..." Re #5, the missing context is "For proposed testing criteria, refer to this Metaculus Question by Matthew Barnett, or the Longbets wager between Ray Kurzweil and Mitch Kapor."

@TimWho I agree and am worried we're going to end up having agonizing, invidious choices to make to try to balance spirit and letter here. I was dumb to bet in this market so heavily without waiting for clarifications. Of course 2.5 years ago I didn't think it mattered and still thought the Turing test was fine for operationalizing AGI. (I'm still not entirely sure it's not fine, if it's long enough and adversarial enough and the judges can assign arbitrary tasks to the players and...)

So we had @Gen on stage tonight representing Manifold saying he believes AGI is here.

@Joshua Where do we sit on this market?

--a stage at Manifest.

@Quroe I don't work at Manifold anymore but I don't think the criteria have been fulfilled

@Joshua Yeah when I asked you offhanded at Manifest you said to defer to Metaculus, so we will do that 😃

Goal : improve the resolution criteria
Conflict of interest : I am "no" on this market (I think it will happen after 2050)

Why the Turing test is good in theory : If you can't think of any intellectual task the AI can't do (and human can), it seems hard to see why it isn't an AGI, and if you know one task it can't do, just ask the AI to do it or explain how it would to it.
(And it takes care of the part where it is an AGI, but you can tell it is an AI because you see it, and it is a computer or a robot.)
But it fails in practice in two ways :

False negative : Some ways to detect it is an AI have nothing to do with intellectual tasks, maybe the AI have some ethical constraints, maybe it has a style of writing, and both can be used to detect it.
I think it would be unconvincing for people thinking it is an AGI, if it can't pass the Turing test because it can't say some word or doesn't do any grammar mistake.

False positive : In theory you can ask for the category of task the limited AI can't do, but random people will probably not understand the limit of the AI, and will not think about these tasks.
It would be unconvincing for people thinking it is not an AGI, if it passes the Turing test, but it is still unable to do a good score in ARC-AGI or whatever tests like that

Proposition : Instead of using the Turing test, we can wait some time for people to find any intellectual task the AI can't do, and we can.
If they find it, it isn't an AGI, if after some time they don't, it is an AGI, and it resolves yes at the date of when this version of the AI was released.

It far from perfect, but I think it is more in the spirit of this market, what do you think ?

@dionisos "AGI When? [High Quality Turing Test]"

We all bet on that understanding.

@PhilosophyBear Ok, that’s fair

@dionisos I think this is right and @PhilosophyBear is also right and the way to have the best of both worlds is to require such a high quality Turing test that it doesn't have those false positives/negatives. (Nice articulation of those, btw, thank you.)

I predict we'll get AGI and nobody will ever bother running the expensive and cumbersome test.

I'm worried that having "Turing test" in the title will, more and more, lead traders astray. This market was created two years ago, back when it was easy to make an AI fall on its face answering a single question. As Dwarkesh Patel eloquently put it a few months ago, sometimes goalpost moving is fair. Because you learn that the goalposts were wrong. I tend to think that a high enough quality Turing test will continue to work, but it might get to the point where grilling the AI has to include having it go out and perform real work on the real internet, coherently over hours or days, to prove that it's a truly general intelligence.

In short, I suspect plenty of traders are betting on EARLIER in this market because they predict the Longbets version of the Turing test, if set up faithfully as originally spec'd, will soon fall. But my interpretation of the market description is that we mean something more stringent than that.

I think this market might be less misleading if we simply removed "[High Quality Turing Test]" from the title and let the existing market description convey the nuance of what we actually mean by AGI. As the market description concludes, we're aiming at the spirit of the question for what people actually mean by AGI. E.g., definitely not the models we have as of early 2026 whether or not those models can be unmasked in a couple hours of text-only interaction.

@dreev Existing market description references longbets version of Turing test multiple times and does not have a single reference for "working for days". I don't know how it can be interpreted any differently

@MikhailDoroshenko I'm looking especially at the final note, "These criteria are still in draft form, and may be updated to better match the spirit of the question. Your feedback is welcome in the comments."

It also says this market isn't about the Lonbets Turing test specifically.

I guess it's high time we get this more pinned down.

Maybe a simpler way to put this: Suppose, hypothetically, that we went all out on running the highest possible quality Turing test and the AI passed, today. Which would be our reaction for this market?

  1. Wow, apparently we hit AGI without all hell breaking loose.

  2. Oops, the Turing test apparently doesn't work to distinguish AGI.

@dreev imo doesn't matter, this market resolves yes

@MikhailDoroshenko I think I mostly agree, but, again, consider the final note in the market description about updating the resolution criteria to better match the spirit of the question. Still, the market description spends enough time on Turing tests specifically that it would feel unfair to throw that out completely.

Maybe what the market description implies is that we can go beyond the Longbets wager to a Turing test that's as stringent as necessary to match the spirit of the question, perhaps taking us closer to Aschenbrenner's drop-in remote worker. I think testing for that kind of capability is still in the scope of an expanded Turing test (as long as all the communication involved is text-based?).

How long can a market be in draft form before it needs to be locked in? With 1.1k unique traders, surely we all had something in mind when we participated.

For what it's worth, this market is also being used as the source of truth for another market: /elongatedmuskrat/will-we-burn-all-this-firewood-befo

What I had in mind is an AI capable of doing almost any intellectual task a human is able to do, at an ok level of competence (with enough time for both).
Even if the current AI are very competent, I think it still can’t because it lakes stability, and the ability to quickly learn.

Putting stuff in the context windows make the AI more knowledgeable about the current context, but also generally worse (and the training need way too much data to improve the AI compared to us).
And I find it quite plausible both are linked, if we manage to be stable over long periods, maybe it is because we are learning from the context we put ourselves in when accomplishing long tasks (by opposition to just having all the context in our mind, in fact our ability to have all the stuffs in our mind is probably completely shitty compared to what current AI can do)

Comment hidden