
Resolves positively if there is an AI which can succeed at a wide variety of computer games (eg shooters, strategy games, flight simulators). Its programmers can have a short amount of time (days, not months) to connect it to the game. It doesn't get a chance to practice, and has to play at least as well as an amateur human who also hasn't gotten a chance to practice (this might be very badly) and improve at a rate not too far off from the rate at which the amateur human improves (one OOM is fine, just not millions of times slower).
As long as it can do this over 50% of the time, it's okay if there are a few games it can't learn.
People are also trading
Plays doom in a video about half way down
https://typesafe.ai/blog/introducing-system-one-models-and-jev
This market looks like it's resolving YES, so I've made a more aggressive version, which tries to operationalize ~superhuman game-playing computer use /AdamK/in-2028-will-a-public-ai-be-able-to
@AdamK I hold some yes, but you do hold a very large yes position. It is in your interest to convince the market yes is likely. And while I think yes is more likely than no here models tend to be spikey and I don’t know if labs are going to put enough effort and training in to make this possible by eoy 2027.
@HumanClanker I was literally arguing that this market's resolution criteria are too generous to YES in the previous thread...
@AdamK I think it’s likely but not obviously yes. I didn’t read all the previous threads you wrote in though. I was only referring to the fact that I didn’t think it was obvious that the market would resolve yes. One OOM is pretty generous though so I agree that it seems plausible.
@Sketchy Ooh, that’s a good point. If AI become capable of looking at a task/type of activity, making their own effective custom harness and then using that harness to do the thing nearly as well/better than humans, then we no longer need to parcel out “base model” vs “human-assisted” in one of the currently most important ways for this (kind of) market. And tool use/training/harnesses/software integration is obviously trending this direction, plus the raw capabilities are probably almost here today.
2027, year of the self-harnessed, broad computer use agent? Maybe even end of this year if things keep speeding up?
@DavidHiggs yea I don’t see any reason why humans would need to be involved in harness creation soon, especially once best practices are widely available online or even pretrained on.
@Sketchy Yup, gonna progress through humans required generally to overly niche tasks/types of harnesses and then broad AI sufficiency just like general programming is doing.
@MartinRandall Yeah, building scaffolds based on detailed knowledge of the game might go against the spirit of the market. If computer use is good enough we should care most about zero-shot performance. (But if the model is adjusting its harness while playing, that seems like fair game.)
@AdamK I mean, the market says humans can have days per game to build harness. If the AI is capable of doing that task, it seems like completely fair game.
@AdamK It doesn't go against the spirit of the market, humans are just salty that they're not superintelligent. Human gamers could also reverse compile the game and build a custom harness in a few days, it's just that humans aren't that smart.
@MartinRandall Sure. I've made a more aggressive market to get at the more aggressive zero-shot computer use possibility: /AdamK/in-2028-will-a-public-ai-be-able-to
@robm well it's playing I guess, maybe? these shots the editor didnt cut are pretty shit tier though. and the discontinuities makes me think these are the exceptionally good ones
@Stralor the author is claiming this as one continuous playthrough. I don't know if the video has edits. But it played to the closing credits in just under 24 hours real-time (custom harness pauses the game while the model thinks). The harness also lets it rotate the camera arbitrary fast, so it may look more edited than it is.

