Skip to main content
MANIFOLD
In early 2028, will an AI be able to generate a full high-quality movie to a prompt?
4.6k
Ṁ21kṀ14m
2028
28%
chance

EG "make me a 120 minute Star Trek / Star Wars crossover". It should be more or less comparable to a big-budget studio film, although it doesn't have to pass a full Turing Test as long as it's pretty good. The AI doesn't have to be available to the public, as long as it's confirmed to exist.

Market context
Get
Ṁ1,000
to start trading!
Sort by:

I think this depends on the ability of the OP to identify a "pretty good" film and having good taste. AI almost certainly can create videos looks like a film in early 2028, but "high quality film"? depends on your quality bar.

@633b I don't think it does.

if you did it today (my most recent attempt here), you would end up with obvious AI tells (bad writing, repeated dialogue, characters doing things that make no sense, continuity errors, ghosting, horrifying anatomy) appearing virtually everywhere. Merely creating a film that an average human could physically tolerate watching for 90 minutes would be a massive leap over the current state of the art and well above-trendline. The trendline for "how long an ai clip [without cherry-picking] before you stop an obvious error 50% of the time" has gone from 5 seconds to maybe 30s in the last year. The trendline for "how long a written work before it outputs an obvious LLMism" is basically flat at "less then the length of an average tweet".

Unfortunately "throw more compute at it" doesn't appear to be a solution with current models. Here is a music video generated by Astra and the same video after giving Astra a "make it better" goal and letting it grind for 48 hours. There is noticeable improvement but it's still nowhere near the "normal human wouldn't know this was AI generated" level.

Ai models seem to suffer from LLM-blindness where they not only cannot identify LLM tells, but actively prefer LLM generated outputs to human ones. Most likely this is a side-effect of LLMs being trained using reinforcement-learning and LLM-as-judge, which optimizes for LLMs producing outputs that LLMs reward.

Could this change quickly if OpenAI threw a few billion dollars at a model focused on creative outputs paying actual human beings (not third world sweatshop workers who speak a non-standard dialect of English) to judge the output? Probably. Does that seem likely to happen in the next 15 months? No. All of the major ai-labs (google excepted) are laser-focused on coding because: 1) that's where the money is 2) coding feeds directly into RSI allowing them to improve faster. It's possible that RSI produces takeoff/foom and we get AI that can do basically anything in 15 months and that possibility probably contributes more probability-mass than the direct: OpenAI trains a model really good at movie generation pathway.

bought Ṁ100 NO

@LoganZoellner

"(not third world sweatshop workers who speak a non-standard dialect of English)"

That's a very weird thing to say. That is not the problem, & it's not a fair description of the workers. It's just racist.

Can anyone explain why we're not able to do this yet? How close is SotA currently?

From my understanding:

- we can generate high res images and short clips (e.g. Sora)

- current models can work on large codebases/long chats and maintain consistency

- current models can create screenplays/books (albeit v low quality)


so what we lack is maintaining coherence over a long period of time?


if we reduce the complexity of the video (e.g. stick figures or some other simple animation) are we able to maintain temporal consistency for longer, and if so how long?


is there any relationship between parameters and (consistent) length of video produced?

intuitively, I would've thought consistency can be maintained via some high level features (extracted from frames + story/screenplay) and continually cross-referencing new frames with existing frames and re-generating when needed

Or via the tools themselves (e.g. Pixar uses 3D models of characters/objects etc. and puppeteers them in the environment vs generating each frame from scratch.
Could Astra or future versions use e.g. Blender to do the same?)


apologies if these are dumb questions

@elf the answer is: for every single one of those points, models are not "good enough". with some amount of future scaling they might be, but on the current trendline that scaling will not happen before "early 2028"

short clips from Sora: even with the best current model (seeddream2/google-omni) it is rare to generate even 30s of ai video without some obvious glaring glitch that is obvious to a human being

large codebases: LLMs are UNIQUELY BAD at writing and UNIQUELY GOOD at code. even so, any programmer will tell you that LLM generated code is hideous to look at full of weird constructions unnecessary comments and basically nothing like what a human would write. Reinforcement-learning can be used to train "pass this test" or even (as we learned yesterday "solve this math problem") but it cannot be used to teach taste

screenplay writing: it is probably true that if every other aspect were good enough a mediocre human written screenplay would be sufficient. But LLMs cannot produce even a mediocre human-written-screenplay. An llm-written screenplay would be filled with repeated use of the same phrases, LLM-sims ("it's not just X, it's Y") and heavy-handed reuse of the same three tropes throughout the film. Unfortunately, this part is getting WORSE not better. Modern LLMs talk less like a human than the original gpt-3.

There is undoubtedly some correlation between number of parameters and length of usable clip, but this parameters range from low billions = ever single frame has problems to 10's of billions = glitches are visible every few seconds. Currently there is no ongoing effort to train a true-multimodal (video-in, video-out) transformer at the scale of headline models like Bel (trillions of parameters) which solved Navier-Stokes.

the "best hope" for a positive resolution is that the Bel class of models has sufficient transfer-learning that it is able to build some kind of hybrid-solution: block out scenes in blender or another 3d modeling program, use a video model to turn this into realistic video, inspect resulting videos for problems in sound/visual glitches. Unfortunately (trust me I have been trying) LLMs are completely blind to LLM failure modes and there is 0 indication that scaling larger LLMs fixes this.

bought Ṁ1,000 YES

@elf it’s hard to say how close we are because it’s unclear how much compute is being thrown at the components of this specific problem. I’ve flipped my bet to Yes today because I suspect that OpenAI or Google or Anthropic will be able to just throw $50M worth of tokens at the problem in early 2028 and get a movie that fulfills the questions criteria.

The question didn’t specify that this sort of output has to be cheap or accessible to outside consumers, so even one movie should resolve this to Yes.

bought Ṁ500 NO

@nsokolsky it’d still have to be made in response to a single prompt

@MachiNi hm... but what if its a single prompt, but then the AI automatically spawns a bazillion sub-agents in a bazillion different attempts? That should fulfill the criteria, even though it would be defacto impractical for anyone without a spare $50M to replicate.

@nsokolsky Yeah, I was about to say maybe by the target date of the market it is possible to do but it’s just cost prohibitive because there isnt a financial or reputational/PR incentive to do it unlike with, say, the Millennium Prize. Especially if it’s likely not going to be high quality.

Film/Entertainment also has a pretty small TAM so it’s rational for the labs to focus more on other stuff like coding, enterprise, drug discovery, insurance etc

bought Ṁ50 NO

AI psychosis

@AlanTennant nah, just moderately bullish at this point on progress rates + priorities of frontier model companies. I’m starting to think the price is (nearly) reasonable at the moment.

Main problem is this isn’t going to happen borderline accidentally like the recent maths capabilities, and it will likely be too costly to deliberately go for it within the next 16 months.

Yes I’m saying successful auto-prompted movies won’t present a good enough market value to justify their opportunity cost.

Q: Since you mention Star Trek crossovers...would something on the order of "Star Wreck: In the Pirkinning" qualify?

WaveSpeedAI (@wavespeed_ai) on X

we are approximately 1 gpt away from this process being good enough to make a full film. This market is more or less a: will gpt-7 release before early 2028 market now.

bought Ṁ500 NO

@LoganZoellner and what’s the probability of that?

@Frogswap seems like @LoganZoellner is leaving a lot of money on the table then

@LoganZoellner according to the market @Frogswap linked to there’s a 76% chance GPT-7 will be released before 2028. Unless it takes several months to actually make the movie once it’s available, one of these two markets is badly mispriced according to your hypothesis.

@MachiNi Two things I try never do: bet against AI progress being faster than expected and bet that I'm smarter than a prediction market. What happens when an unstoppable force hits an immovable object? idk.

@LoganZoellner Isn't any bet on a prediction market a bet that you are smarter than a prediction market?

@Balasar if I have specific non public information, no

still pretty glitchy, but seems like people are starting to use harnesses to fully automate video gen

so early 2028, with gpt 7 + seedance 4, how far can that be pushed?

haven't seen a fully-automated "harness" pull off a coherent multi-shot scene like this though:

https://www.reddit.com/r/aivideo/comments/1w8xqk4/one_small_bite_seedance_wip_trailer/

Even if it's just 30s long

@0xseraphim 16 months left

@bens real movies are jumbles of barely coherent shots strung together and are full of continuity errors. There are very few steps left to go from current tech to a movie that resolves this question YES

@0xseraphim AI can't even currently generate a "full high-quality" novel, let alone a movie.

@bens are you familiar with the plot and dialogue of Transformers: Revenge of the Fallen (2009)?

@bens Megan fox runs around in it. And then there are explosions. And mythology.

@bens

Almost made a billion dollars at the box office

The plot and dialogue of Transformers: Revenge of the Fallen (2009) isn't even the floor. This question's criteria just says "It should be more or less comparable to a big-budget studio film"

@bens my confidence is high that automated plot/dialogue generation passed the threshold set by Transformers: Revenge of the Fallen (2009) some time ago

@0xseraphim I can tell you that I am extremely locked into SOTA hour-length vertical microdramas right now. They (1) involve some degree of human-in-the-loop still, and (2) are not remotely at the quality of "high-quality big budget studio films". That being said, progress is quite fast! That's why this market is at 28%. I'd price it closer to 10-15%, but not zero!

@bens we need the metr graph going from the microdrama to the feature film 😄

IDK how anyone looks at AI progress and thinks "wow 16 months is a long time"

@jim long time or short time?

@0xseraphim I can't figure it out, but the one that means we're going to get AI movies this year

@jim oh this year, really? short timelines indeed

bought Ṁ200 NO

@jim I’m confused. Aren’t you the one who always believes AGI and all that entails is just around the corner?

@MachiNi yes but I used the word "long" when I ought to have used the word "short". My best guess is that a qualifying system for this market to resolve as 'yes' will be constructed on 29 December, 2026.