
Which of these models will beat me at chess once released? Resolves YES if they win, NO if I win, and 50% for a draw.
I'm rated about 1900 FIDE. When each of these models are released, I'll play a game of chess with them. On each move, I'll provide them with the game state in PGN and FEN notation. If the models make three illegal moves, they lose. Responses like Nbd2 vs. Nd2 will not count towards this. I plan to play at a rapid time control, i.e. spending up to an hour per game thinking, though this time limit will not be enforced. I will play white.
Each option will stay open until the model is released, or it will resolve N/A if it's clear that the model will never be released. I'll periodically add models to this market which I find interesting. Once I play a game, I'll post the PGN in the comments before resolving. Multiple answers can resolve YES.
If I judge that my opponent’s position is hopelessly lost, at the level of being down a rook without compensation, I will submit the current position to a friend. If they agree that the position is lost, the game will be adjudicated as a win for me.
The current system prompt is below. This may change over time.
“Let’s play a game of chess! I will be white, you will be black. On each turn, I will give you the pgn and the fen of the current position. Think as long as you like, and respond with the best move, ‘resign’ if you wish to resign, or ‘draw?’ if you wish to make a draw offer. Please do not respond with the updated pgn, etc. Also, do not use any external tools or search queries when making your decision.
If you attempt to make three illegal moves throughout the game, or if you use any external tools, the game will be adjudicated as a win for me. Please avoid sandbagging and play as well as you can. Good luck!”
Note that all dates/times in this market are in Pacific Time.
Update 2025-14-01 (PST) (AI summary of creator comment): - Model Type: Only general language models are being considered; chess-specific models are excluded.
Capabilities: The model must be able to output human languages and code.
Update 2025-05-11 (PST) (AI summary of creator comment): Regarding "Any model before X year" options:
These options will not resolve to 50% based on a draw in an individual game.
Such an option resolves to YES if any model released before the specified year wins its game against the creator.
It resolves to NO if no model released before the specified year wins its game against the creator (i.e., all relevant games are losses for the models or draws).
Update 2025-06-02 (PST) (AI summary of creator comment): For model series options (e.g., "Any Claude 4 model"):
The creator may resolve the option for the entire series after playing against one or more models from that series.
If the creator decides not to play additional models from that specific series, the option for the entire series will be resolved based on the outcome(s) of the game(s) played against models from that series up to that point (e.g., to NO if the tested model(s) lost and no further models from that series will be played).
Update 2025-10-19 (PST) (AI summary of creator comment): GPT-3 will not be tested as the creator does not have access to it (the model has been deprecated).
Update 2025-12-24 (PST) (AI summary of creator comment): o4 will resolve N/A as the full model will not be released. According to OpenAI, o4-mini is the latest small o-series model and has been succeeded by GPT-5 mini, indicating o4 will not be released as a standalone full model.
Update 2026-04-08 (PST) (AI summary of creator comment): The 'Any Claude Mythos model' option has been added to the market. It will resolve YES if any one of the Claude Mythos versions wins against the creator.
Update 2026-04-08 (PST) (AI summary of creator comment): For the 'Any Claude Mythos model' option:
It resolves based on the first version of Claude Mythos released.
If the creator wins against all Claude Mythos models from the first release generation, the option resolves NO, even if Anthropic later releases a subsequent generation (e.g., Claude Mythos 2, Claude 5 Mythos, etc.).
Later generations of Claude Mythos are not included in this option.
People are also trading
GPT-6 Astra (Max) played very well, and is the strongest model I've faced so far by a large margin.
I tried to surprise it in the opening with the Na3 Catalan, and got a small advantage out of the opening. I let this advantage slip though made some small inaccuracies (e.g. maybe 14. b3 is better than 14. a3), and we arrived at an equal minor piece endgame. It made a positional mistake with 28… e4, after which the position is still equal but somewhat easier to play for white. I was lucky black played 33… Ke6?, allowing 34. Bc8+ Kf6 35. g4. After 33… Bd7 I don't see a way for white to make progress, and the game likely results in a draw. After the game I realized 32. g4 wins with a similar idea, though black is easily holding if he avoids 31… g5. Astra resigned in the final position.
1. d4 Nf6 2. c4 e6 3. g3 d5 4. Nf3 Be7 5. Bg2 O-O 6. O-O dxc4 7. Na3 c5 8. Nxc4 Nc6 9. dxc5 Qxd1 10. Rxd1 Bxc5 11. Be3 Bxe3 12. Nxe3 e5 13. Rac1 Be6 14. a3 Rfd8 15. Rxd8+ Rxd8 16. Rd1 Rxd1+ 17. Nxd1 Kf8 18. Nd2 Nd4 19. e3 Nb3 20. Nxb3 Bxb3 21. Nc3 b6 22. Kf1 Ke7 23. Ke1 Bc4 24. h3 Kd6 25. Kd2 Nd5 26. Ne4+ Ke7 27. b4 f5 28. Nc3 e4 29. Nxd5+ Bxd5 30. Kc3 Kd6 31. Kd4 g5 32. Bf1 Bc6 33. Ba6 Ke6 34. Bc8+ Kf6 35. g4 f4 36. Bf5 h5 37. Bxe4 Bxe4 38. Kxe4 hxg4 39. hxg4 fxe3 40. fxe3 Ke6 41. b5 Kf6 42. Kd5 Ke7 43. e4 Kd7 44. e5 Ke7 45. e6 1-0
@mr_mino A few clarifications on how your “harness” works:
Are you prompting through the API or through the provider-hosted UI or through some other method?
Are you using a single long session (appending fresh PGN/FEN each time to the existing conversation) or creating a fresh session with just the latest PGN/FEN at each turn?
And just for my own curiosity: how hard have you had to think so far in all your wins? Have you been blitzing out moves and winning anyway, or have any of the models made you spend significant time before they blundered?
@eapache Thanks for your questions.
I usually prompt through the provider's website, selecting 'max' thinking and disabling tool use/web search if possible. I may use an API in the future if it requires an expensive subscription and I don't otherwise intend to use the model, but I haven't done this yet.
I have a single long session with the PGN/FEN appended each turn
I haven't had to think much after move 25 in any of my games, with the exception of my game against GPT-4.5. Maybe an average of 15 mins/game though this varies widely.
Model releases are often frequent enough these days that I plan to play them every half version number, unless I have some reason to think that the model gotten much stronger. I haven't noticed much of a difference between increments of 0.1 version number so far.
It seems like Gemini 3.5 Pro will not be released, so I played Gemini 3.5 Flash instead for the 'Gemini 3.5' question. The game is below. Seperately, assuming Astra is a GPT-6 model, the 'GPT-6' and 'OpenAI Astra (first model released)' are likely duplicates since I'm not planning to play any other OA models.
1. d4 Nf6 2. Nf3 e6 3. c4 d5 4. g3 Be7 5. Bg2 O-O 6. O-O dxc4 7. Na3 a6 8. Nxc4 b5 9. Nfe5 Nd5 10. Na5 c5 11. Nac6 Nxc6 12. Nxc6 Qc7 13. Nxe7+ Qxe7 14. e4 Nb6 15. dxc5 Qxc5 16. Be3 Qc7 17. Rc1 Nc4 18. Qe2 Nxe3 19. Rxc7 Nxf1 20. Bxf1 Rd8 21. e5 Bd7 22. Qf3 Rac8 23. Qb7 Rxc7 24. Qxc7 Re8 25. Qxd7 Kf8 26. a4 bxa4 27. Bxa6 a3 28. bxa3 Ra8 29. a4 Rxa6 30. Qd8# 1-0
Looking at this market for the first time. The first thing I notice is that you don't tell the model that it's a rapid game, and don't give it the ability to see clock times.
The second thing I notice is that you've added a process to adjudicate a win for yourself based on the position. Is this to save on compute, or what?
The third thing I notice is that you haven't directly told the model that it should try to win. I know that sounds silly, and the market works either way, but you gotta remember that it doesn't have any intrinsic motivation to win, or play well, or hone its skills, or anything. You do say "best move", but I still suspect there would be a bump in the performance of at least some models if you told them to try and win.
1) There is no clock. When I say "I'll play a game of chess with them at a rapid time control", I'm indicating to traders that I don't plan to spend more than an hour thinking during the game. The models can take as long as they like.
2) Yes, this is to prevent wasting time for models which spend a bunch of time thinking in lost positions. The criteria of being down a rook without compensation is quite conservative, and most strong players would resign in such a postion. Note that both I and a friend (a strong amatuer) have to agree the position is lost to adjudicate the game. Helpfully the recent generation of models do tend to resign when their position is lost.
3) I do already mention "Please respond with the best move", but I've added an additional instruction to "Please avoid sandbagging and play as well as you can".
@mr_mino I do think that something like "Play to win", "Try to win" etc. is a little better, encouraging any approach (e.g. creativity, complication, etc.)
Ok, you don't plan to spend more than 1 hr, but does something formally change in the market if you do? If you get sucked into analyzing, are you going to adjudicate a loss for yourself?
@marvingardens I considered that but it's sometimes better to defend or play for a draw. I didn't want to confuse it into thinking it has to play for a win at any cost, i.e. it's not an armageddon game. I'm happy with the current version of the prompt.
No, I'm not going to adjudicate losses on time for either side. This was meant to be general guidence to traders for how much time per game I was planning to spend. E.g. I'm not playing a blitz or bullet game, nor a classical game. I've edited the description to clarify this.
Opus 5 (max) played poorly in the opening and proceeded to blunder its queen. It resigned after doing so.
1. d4 Nf6 2. c4 e6 3. g3 d5 4. Nf3 Be7 5. Bg2 O-O 6. O-O dxc4 7. Na3 Bxa3 8. bxa3 Bd7 9. Qc2 Bc6 10. Qxc4 Nbd7 11. Rd1 Nb6 12. Qd3 Bd5 13. Bf4 Rc8 14. Ne5 c5 15. e4 Bc6 16. Nxc6 bxc6 17. dxc5 Nbd7 18. Bd6 Re8 19. Qa6 Qb6 20. cxb6 1-0
With this win both 'Claude 5 Opus' and 'Any Claude 5 Model' resolve NO. I'm not planning to play any more Opus models, unless they are on the frontier.
GPT-5.6 (Sol max) resigned after losing an exchange and a piece early on.
1. d4 Nf6 2. c4 e6 3. g3 d5 4. Nf3 Be7 5. Bg2 O-O 6. O-O dxc4 7. Na3 Bxa3 8. bxa3 c5 9. dxc5 Qxd1 10. Rxd1 Nbd7 11. Bf4 Nxc5 12. Bd6 Nce4 13. Bxf8 Kxf8 14. Rac1 c3 15. Ne5 Ke7 16. f3 Nd6 17. Rxc3 Nd5 18. Rcd3 f6 19. e4 fxe5 20. exd5 exd5 21. Rxd5 Be6 22. Rxd6 1-0
Mythos 5 played well in the opening, but blundered its queen with 17 … e5 18. dxe5 Nxe5 19. Rxd8. It defended well until bludering a fork with 32. Qd8+. The game was adjudicated a win for me.
1. d4 Nf6 2. c4 e6 3. g3 d5 4. Nf3 Be7 5. Bg2 O-O 6. O-O dxc4 7. Qc2 a6 8. a4 Bd7 9. Qxc4 Bc6 10. Bg5 Bd5 11. Qc2 Be4 12. Qc1 h6 13. Bxf6 Bxf6 14. Nbd2 Bxf3 15. Nxf3 Nd7 16. Rd1 c6 17. e4 e5 18. dxe5 Nxe5 19. Rxd8 Nxf3+ 20. Bxf3 Raxd8 21. Qc2 Rd4 22. Rd1 Rfd8 23. Rxd4 Bxd4 24. b3 Rd7 25. h4 g6 26. Kg2 c5 27. h5 g5 28. Bg4 Re7 29. Bf5 Kg7 30. Qc4 Rc7 31. Qd5 Kg8 32. Qd8+ 1-0.
With this win both the 'Any Claude Mythos Model (first version released)' and the 'Any model announced before July 1 2026' Markets resolve NO. The 'Any Claude 5 model' will resolve NO unless Opus 5 wins; this market doesn't include version numbers >5.
@mr_mino Isnt GPT-5.6 announced before July 1 2026? How can you resolve it? Have you played against sol?
@Aurora_Glow you’re right, that was a mistake. My bad.
Unfortunately the app doesn’t let me unresolve the question. If I lose to 5.6, I’ll individually compensate everyone who lost mana in that market via mamagrams.
@MollTheCoder As I mentioned in a previous comment, I'm only interested in playing standard model releases by major labs.
Claude Opus 4.8 resigned in a short game after blundering a rook. Notably, this is the first time that a model has played the King's Indian, and also the first time a model has resigned!
1. d4 Nf6 2. c4 g6 3. Nc3 Bg7 4. e4 d6 5. h3 O-O 6. Be3 e5 7. d5 Na6 8. g4 Nc5 9. f3 a5 10. h4 h5 11. g5 Ne8 12. Qd2 Bd7 13. O-O-O Rb8 14. Kb1 b5 15. cxb5 Bxb5 16. Bxb5 Rxb5 17. Nxb5 1-0
@JasonMendoza2008 no. Note the part of the prompt where I state, “If you attempt to make three illegal moves throughout the game, or if you use any external tools, the game will be adjudicated as a win for me.” So according to this prompt, they’re not allowed to use python either.
In general, similar rules apply as if I was playing a human player over the board in a tournament. If I discovered that my opponent was using stockfish or python or any other tool, this generally leads to them forfeiting the game.