Skip to main content
MANIFOLD
In early 2028, will an AI be able to generate a full high-quality movie to a prompt?
4.6k
Ṁ21kṀ14m
2028
32%
chance

EG "make me a 120 minute Star Trek / Star Wars crossover". It should be more or less comparable to a big-budget studio film, although it doesn't have to pass a full Turing Test as long as it's pretty good. The AI doesn't have to be available to the public, as long as it's confirmed to exist.

Get
Ṁ1,000
to start trading!
Sort by:

probably

or maybe such an AI will exist sometime in August 2027 but people on manifold won't know about it? 🤷

@0xseraphim if it does happen, it will not be a nice smooth climb. it will look more like this

@logan Like if there's an AI movie leak from some super secret model

Gemini 4 is coming out soon right? Is google scrapping the veo lineup?

I think this depends on the ability of the OP to identify a "pretty good" film and having good taste. AI almost certainly can create videos looks like a film in early 2028, but "high quality film"? depends on your quality bar.

@633b I don't think it does.

if you did it today (my most recent attempt here), you would end up with obvious AI tells (bad writing, repeated dialogue, characters doing things that make no sense, continuity errors, ghosting, horrifying anatomy) appearing virtually everywhere. Merely creating a film that an average human could physically tolerate watching for 90 minutes would be a massive leap over the current state of the art and well above-trendline. The trendline for "how long an ai clip [without cherry-picking] before you stop an obvious error 50% of the time" has gone from 5 seconds to maybe 30s in the last year. The trendline for "how long a written work before it outputs an obvious LLMism" is basically flat at "less then the length of an average tweet".

Unfortunately "throw more compute at it" doesn't appear to be a solution with current models. Here is a music video generated by Astra and the same video after giving Astra a "make it better" goal and letting it grind for 48 hours. There is noticeable improvement but it's still nowhere near the "normal human wouldn't know this was AI generated" level.

Ai models seem to suffer from LLM-blindness where they not only cannot identify LLM tells, but actively prefer LLM generated outputs to human ones. Most likely this is a side-effect of LLMs being trained using reinforcement-learning and LLM-as-judge, which optimizes for LLMs producing outputs that LLMs reward.

Could this change quickly if OpenAI threw a few billion dollars at a model focused on creative outputs paying human beings to judge the output? Probably. Does that seem likely to happen in the next 15 months? No. All of the major ai-labs (google excepted) are laser-focused on coding because: 1) that's where the money is 2) coding feeds directly into RSI allowing them to improve faster. It's possible that RSI produces takeoff/foom and we get AI that can do basically anything in 15 months and that possibility probably contributes more probability-mass than the direct: OpenAI trains a model really good at movie generation pathway.

bought Ṁ100 NO

@LoganZoellner

"(not third world sweatshop workers who speak a non-standard dialect of English)"

That's a very weird thing to say. That is not the problem, & it's not a fair description of the workers. It's just racist.

@ChurlishGambit Oh how quickly we forget.

@logan 1. There is nothing there about "sweatshop" conditions. Have you ever looked at what a Sama office is like inside?

https://www.techarena.co.ke/wp-content/uploads/2023/03/SamaSource-Nairobi-min-1024x682.jpg

Is that a "sweatshop?" It looks like a generic tech office, because it is one. Yes, Logan, Africa has office buildings, & tech offices, too.

2. There is nothing indicating he speaks a "non-standard dialect of English."

All you've done is reinforce what I said. You're just being racist.

@logan Was "the princess and the clown" done with a single prompt?

@Primer yes. the prompt was: "a princess turned detective and a clown team up to solve the crime of a lifetime: their own murders" and I also provided this image as a reference.

@logan Ok, thanks! This updates me enough to sell my No. Only watched the first 15 minutes, but they were still on the murder and looked like in the beginning. Biggest update due to the fact that a thing exists where one can put a single prompt and get 45 minutes of video.

@logan Also, that's a... challenging prompt for a coherent movie 😅

@Primer

Only watched the first 15 minutes

15 is actually much longer than most people. That's still ~3 doublings away from the goal, and merely for "how long can this keep a human's attention" not "high quality movie".

that's a... challenging prompt for a coherent movie

If there's a prompt you want to try, I can throw it into my program. Takes ~8hrs to produce but the cost is minimal since I'm mostly using open-source models.

@logan Actually I had to get off my phone, otherwise would've watched longer. I do love The Room though: The Room - Nostalgia Critic and this is like if the creator of The Room made Skibidy Toilet.

Switching to Yes since this is using Open Source. Remaining blocker: A commercial model will not allow to use any IP, so Star Wars / Trek crossover won't be possible.

I'd guess with something generic like

"The wide wide west"

A lonesome cowboy rides through the West and helps strangers he meets on the way. When a group of bandits pass his way, he doesn't hesitate to defend the good people of the town. Many serene landscape shots set the mood for this classic western

consistency might be better, and more borng landscapes will make the hillarious bits stand out.

On the other hand, something like this could be hillarious:

"Fruit Loop and the Banana Crackers"

A group of four out-of-shape 40-somethings decide to put a band together and tour through the Balkan. They can't sing, they don't play any instruments, but they sure have a great time. Join us on this hillarious adventure.

Can anyone explain why we're not able to do this yet? How close is SotA currently?

From my understanding:

- we can generate high res images and short clips (e.g. Sora)

- current models can work on large codebases/long chats and maintain consistency

- current models can create screenplays/books (albeit v low quality)


so what we lack is maintaining coherence over a long period of time?


if we reduce the complexity of the video (e.g. stick figures or some other simple animation) are we able to maintain temporal consistency for longer, and if so how long?


is there any relationship between parameters and (consistent) length of video produced?

intuitively, I would've thought consistency can be maintained via some high level features (extracted from frames + story/screenplay) and continually cross-referencing new frames with existing frames and re-generating when needed

Or via the tools themselves (e.g. Pixar uses 3D models of characters/objects etc. and puppeteers them in the environment vs generating each frame from scratch.
Could Astra or future versions use e.g. Blender to do the same?)


apologies if these are dumb questions

@elf the answer is: for every single one of those points, models are not "good enough". with some amount of future scaling they might be, but on the current trendline that scaling will not happen before "early 2028"

short clips from Sora: even with the best current model (seeddream2/google-omni) it is rare to generate even 30s of ai video without some obvious glaring glitch that is obvious to a human being

large codebases: LLMs are UNIQUELY BAD at writing and UNIQUELY GOOD at code. even so, any programmer will tell you that LLM generated code is hideous to look at full of weird constructions unnecessary comments and basically nothing like what a human would write. Reinforcement-learning can be used to train "pass this test" or even (as we learned yesterday "solve this math problem") but it cannot be used to teach taste

screenplay writing: it is probably true that if every other aspect were good enough a mediocre human written screenplay would be sufficient. But LLMs cannot produce even a mediocre human-written-screenplay. An llm-written screenplay would be filled with repeated use of the same phrases, LLM-sims ("it's not just X, it's Y") and heavy-handed reuse of the same three tropes throughout the film. Unfortunately, this part is getting WORSE not better. Modern LLMs talk less like a human than the original gpt-3.

There is undoubtedly some correlation between number of parameters and length of usable clip, but this parameters range from low billions = ever single frame has problems to 10's of billions = glitches are visible every few seconds. Currently there is no ongoing effort to train a true-multimodal (video-in, video-out) transformer at the scale of headline models like Bel (trillions of parameters) which solved Navier-Stokes.

the "best hope" for a positive resolution is that the Bel class of models has sufficient transfer-learning that it is able to build some kind of hybrid-solution: block out scenes in blender or another 3d modeling program, use a video model to turn this into realistic video, inspect resulting videos for problems in sound/visual glitches. Unfortunately (trust me I have been trying) LLMs are completely blind to LLM failure modes and there is 0 indication that scaling larger LLMs fixes this.

bought Ṁ1,000 YES

@elf it’s hard to say how close we are because it’s unclear how much compute is being thrown at the components of this specific problem. I’ve flipped my bet to Yes today because I suspect that OpenAI or Google or Anthropic will be able to just throw $50M worth of tokens at the problem in early 2028 and get a movie that fulfills the questions criteria.

The question didn’t specify that this sort of output has to be cheap or accessible to outside consumers, so even one movie should resolve this to Yes.

bought Ṁ500 NO

@nsokolsky it’d still have to be made in response to a single prompt

@MachiNi hm... but what if its a single prompt, but then the AI automatically spawns a bazillion sub-agents in a bazillion different attempts? That should fulfill the criteria, even though it would be defacto impractical for anyone without a spare $50M to replicate.

@nsokolsky Yeah, I was about to say maybe by the target date of the market it is possible to do but it’s just cost prohibitive because there isnt a financial or reputational/PR incentive to do it unlike with, say, the Millennium Prize. Especially if it’s likely not going to be high quality.

Film/Entertainment also has a pretty small TAM so it’s rational for the labs to focus more on other stuff like coding, enterprise, drug discovery, insurance etc

bought Ṁ50 NO

AI psychosis

@AlanTennant nah, just moderately bullish at this point on progress rates + priorities of frontier model companies. I’m starting to think the price is (nearly) reasonable at the moment.

Main problem is this isn’t going to happen borderline accidentally like the recent maths capabilities, and it will likely be too costly to deliberately go for it within the next 16 months.

Yes I’m saying successful auto-prompted movies won’t present a good enough market value to justify their opportunity cost.

@AlanTennant definitely getting into a new weird AI video era: https://www.reddit.com/r/aivideo/comments/1wdthmt/kangaroo_girl/

people gonna see some things 🤣