Skip to main content
MANIFOLD
In early 2028, will an AI be able to generate a full high-quality movie to a prompt?
4.6k
Ṁ21kṀ14m
2028
34%
chance

EG "make me a 120 minute Star Trek / Star Wars crossover". It should be more or less comparable to a big-budget studio film, although it doesn't have to pass a full Turing Test as long as it's pretty good. The AI doesn't have to be available to the public, as long as it's confirmed to exist.

Get
Ṁ1,000
to start trading!
Sort by:

It would be cool to see a breakdown of who thinks we are one breakthrough away from this in the next 16 months, vs. who thinks we are more than one breakthrough away (but still thinks it will happen).

opened a Ṁ5,000 YES at 30% order

@BionicD0LPH1N I would add to check at examples from September 22 2025. GPT-5 and Opus 4.1 were, lets just say, not great at animations. This was presented as an impressive animation by Opus 4.1. SOTA from last year movie scenes recreation:

@BionicD0LPH1N Another great example:

@BionicD0LPH1N My favorite one so far:

probably

or maybe such an AI will exist sometime in August 2027 but people on manifold won't know about it? 🤷

@0xseraphim if it does happen, it will not be a nice smooth climb. it will look more like this

@logan Like if there's an AI movie leak from some super secret model

Gemini 4 is coming out soon right? Is google scrapping the veo lineup?

I think this depends on the ability of the OP to identify a "pretty good" film and having good taste. AI almost certainly can create videos looks like a film in early 2028, but "high quality film"? depends on your quality bar.

@633b I don't think it does.

if you did it today (my most recent attempt here), you would end up with obvious AI tells (bad writing, repeated dialogue, characters doing things that make no sense, continuity errors, ghosting, horrifying anatomy) appearing virtually everywhere. Merely creating a film that an average human could physically tolerate watching for 90 minutes would be a massive leap over the current state of the art and well above-trendline. The trendline for "how long an ai clip [without cherry-picking] before you stop an obvious error 50% of the time" has gone from 5 seconds to maybe 30s in the last year. The trendline for "how long a written work before it outputs an obvious LLMism" is basically flat at "less then the length of an average tweet".

Unfortunately "throw more compute at it" doesn't appear to be a solution with current models. Here is a music video generated by Astra and the same video after giving Astra a "make it better" goal and letting it grind for 48 hours. There is noticeable improvement but it's still nowhere near the "normal human wouldn't know this was AI generated" level.

Ai models seem to suffer from LLM-blindness where they not only cannot identify LLM tells, but actively prefer LLM generated outputs to human ones. Most likely this is a side-effect of LLMs being trained using reinforcement-learning and LLM-as-judge, which optimizes for LLMs producing outputs that LLMs reward.

Could this change quickly if OpenAI threw a few billion dollars at a model focused on creative outputs paying human beings to judge the output? Probably. Does that seem likely to happen in the next 15 months? No. All of the major ai-labs (google excepted) are laser-focused on coding because: 1) that's where the money is 2) coding feeds directly into RSI allowing them to improve faster. It's possible that RSI produces takeoff/foom and we get AI that can do basically anything in 15 months and that possibility probably contributes more probability-mass than the direct: OpenAI trains a model really good at movie generation pathway.

bought Ṁ100 NO

@LoganZoellner

"(not third world sweatshop workers who speak a non-standard dialect of English)"

That's a very weird thing to say. That is not the problem, & it's not a fair description of the workers. It's just racist.

@ChurlishGambit Oh how quickly we forget.

@logan 1. There is nothing there about "sweatshop" conditions. Have you ever looked at what a Sama office is like inside?

https://www.techarena.co.ke/wp-content/uploads/2023/03/SamaSource-Nairobi-min-1024x682.jpg

Is that a "sweatshop?" It looks like a generic tech office, because it is one. Yes, Logan, Africa has office buildings, & tech offices, too.

2. There is nothing indicating he speaks a "non-standard dialect of English."

All you've done is reinforce what I said. You're just being racist.

@logan Was "the princess and the clown" done with a single prompt?

@Primer yes. the prompt was: "a princess turned detective and a clown team up to solve the crime of a lifetime: their own murders" and I also provided this image as a reference.

@logan Ok, thanks! This updates me enough to sell my No. Only watched the first 15 minutes, but they were still on the murder and looked like in the beginning. Biggest update due to the fact that a thing exists where one can put a single prompt and get 45 minutes of video.

@logan Also, that's a... challenging prompt for a coherent movie 😅

@Primer

Only watched the first 15 minutes

15 is actually much longer than most people. That's still ~3 doublings away from the goal, and merely for "how long can this keep a human's attention" not "high quality movie".

that's a... challenging prompt for a coherent movie

If there's a prompt you want to try, I can throw it into my program. Takes ~8hrs to produce but the cost is minimal since I'm mostly using open-source models.

@logan Actually I had to get off my phone, otherwise would've watched longer. I do love The Room though: The Room - Nostalgia Critic and this is like if the creator of The Room made Skibidy Toilet.

Switching to Yes since this is using Open Source. Remaining blocker: A commercial model will not allow to use any IP, so Star Wars / Trek crossover won't be possible.

I'd guess with something generic like

"The wide wide west"

A lonesome cowboy rides through the West and helps strangers he meets on the way. When a group of bandits pass his way, he doesn't hesitate to defend the good people of the town. Many serene landscape shots set the mood for this classic western

consistency might be better, and more borng landscapes will make the hillarious bits stand out.

On the other hand, something like this could be hillarious:

"Fruit Loop and the Banana Crackers"

A group of four out-of-shape 40-somethings decide to put a band together and tour through the Balkan. They can't sing, they don't play any instruments, but they sure have a great time. Join us on this hillarious adventure.

@logan can you say what intermediate outputs your pipeline creates? E.g., a script, shotlist, first/last frames, character sheets, voices, soundtrack, location art/descriptions, storyboard, outline?

I'm curious if we can identify the strong and weak points in the flow from prompt to film.

@robm

it goes: prompt -> short (1 page) film summary -> list of scenes each with detailed summaries -> screen plays for each scene -> list of shots -> initial image for each shot -> video for each shot

shots also include a setting and a list of characters and stable reference images are used for those so they are conserved from shot to shot.

Easiest places to boost are initial image and video, which can be improved by switching from open source to closed source models (at a price: using google omni, 30 minutes of 4k video costs $540 dollars). Hardest thing to boost is writing, because LLM-that-doesn't-write-like-an-LLM is an unsolved problem.

If every single one of these steps were fixed, we could probably resolve the corresponding low-quality version of this market. But the high-quality version would still be untouched because a movie is not merely a sequence of clips. Positive resolution is going to require one of: better agentic ai, a true multimodal video-in-video-out frontier scale model, or some other breakthrough.

I started a run with Astra letting it simulate an entire film-crew: starting from concept art, blocking out 3d set designs, auditioning for voice actors, ... but I killed it after 4 days (and burning through 2 of my resets on the pro+ plan). The intermediate outputs were not promising. I would love to sic one of those 10k-Bel-agent swarms they used to solve Navier Stokes at the problem, if anyone at OpenAI is looking for research problems.

@logan

The intermediate outputs were not promising.

This needs better casting. The acting sucks. 😜

@Primer Here is the cowboy prompt

@logan lol thanks! Acting is way better than in "A Road for the Lost". Cutting, editing and directing though, there's room for improvement.

@Primer You must have a very low bar for "acting"

@ChurlishGambit Maybe? Don't you think especially voice acting is much better than https://youtu.be/NrlqpQpI7WA?is=hvcVrHA6BNNQMqbe

@Primer That's like asking whether being shot five times is better than being shot six times. Neither is any good at all.

@ChurlishGambit They way I disagree: I'd say it's more like being shot 3 times versus being shot 6 times.