Skip to main content
MANIFOLD
Fable 5.1 VS Astra: Which will I find better?
120
Ṁ1kṀ22k
Sep 17
Astra71%

Resolves according to my judgement

The creator has blocked themselves from betting in this market.
Get
Ṁ1,000
to start trading!
Sort by:
bought Ṁ10 YES

@Bayesian what is your resolution criteria?

what do you use LLMs for?

You have access to it now?

i shall make my determination in the next days, perhaps tomorrow perhaps a couple more daya

bought Ṁ25 YES

how did you like Fable 5 vs 5.6 Sol?

@JoshYou i liked fable more

What if Astra is not released before market close?

bought Ṁ5 YES

claude is my pookie

bought Ṁ50 NO

It seems to me Anthropic rushed out a mediocre Fable 5.1 because releasing anything right after Astra's gonna make them look bad. Also based on what we do know, Fable achieved 83% cap rate on ExploitBench (88% for ACE) vs Astra which OpenAI says got 100%.

@JaundicedBaboon I feel like given what we know about Sol's tendencies to hack and the impossible problems in ExploitGym, I really wouldn't be surprised if Astra's 100% score on ExploitBench wasn't actually legitimate

bought Ṁ50 YES

@JaundicedBaboon idk my initial view of Fable 5.1 was bad, but after using it, my cross-reviewers only seem to have very low level nitpick complaints. It's doing a very good job on midsize coding tasks from concept -> final in one flow

@JaundicedBaboon OpenAI notes that 100% is likely contaminated. They get 40% on a newer set.

Regardless, I don't think the score are directly comparable

@Usaar33 I think it's questionable to assume it's any more contaminated for GPT than it is for Claude. Plus, the newer less contaminated set shows Astra scoring well over 3x Sol's score with much fewer tokens.

Between that and Meta and Google both releasing models today I think its likely the other labs know Astra will be absolutely cracked and want to get out ahead of it.