Third time's the charm?
Scott Alexander has a market resolving in 2028 about whether Gary Marcus will be able to find "three extremely obvious questions, that an average human teenager could certainly answer, which a leading chatbot still fails at at least half the time when asked." (Weird adversarial hacks don't count.)
A year and a half ago, I thought maybe ChatGPT o3 was smart enough that it was there already. It was not. A few months later, I thought GPT-5-Thinking might've gotten there. Still not. So now it's been more than another year and maybe Claude Opus 5.5 is there?
To resolve YES, that it still makes egregious errors, we need, by market close, just one example I can replicate in a Temporary Chat (no access to my chat history) with effort set to max. We won't worry about "average teenagers". If the answer is clear to us and Claude gets it wrong at least half the time, this market resolves YES.
As always, ask clarifying questions before betting too hard. We'll discuss in the comments and aim for the spirit of the question, namely, thinking of this as a capabilities milestone and seeing how correct Gary Marcus was in his LLM skepticism. We'll generally but not necessarily follow precedents from the previous two incarnations of this market for o3 and GPT-5.