Three dimensions and vision · 02

What a finished second of generated film costs

A 43 second animated short about Quad Day, made in a week of evenings on one desktop machine, picture and sound generated together in eight passes. The render time is the easy number. The one that decides what a project like this costs is the ratio between compute that shipped and compute that did not.

August 2026 · MiniMax H3, text to video with audio, 1344×768, 124 frames, 20 steps, INT8 · one GB10 desktop, 122 GB unified memory

42.83 s
finished film
8,779 s
generation, first pass
210 s
compute per finished second
8.2×
per-step cost of two knobs
Eight frames from the final segment of the film, in two rows: a red-haired student signing a clipboard under a green tent while a squirrel returns a flyer, then a crane shot rising over rows of tents along the quad.
A contact print of the last segment, eight frames of the 124 the model produced in one pass. Every frame and every sound in the film came out of one text-to-video model; the grade, the cuts and the audio work are afterwards. Watch the film on Vimeo.

What it is

42.83 seconds at 1344×768 and 24 frames per second: eight segments of about five seconds each and a frozen tail for the end title. The model generates picture and sound in the same pass, so the ambient noise, the voices and the effects are generated too. The story is invented. The history around it is not, and it is sourced under the video.

This is the creative side of a question the note on running a video model on 2017 GPUs asks from the hardware side. That one asks whether the model runs at all on cards nobody wants. This one assumes it runs, and asks what it takes to finish something with it.

The arithmetic of one clip

The first complete pass of the film, eight segments at the same settings on the same desktop, every one of them usable on the first attempt:

segment wall clock segment wall clock
01 dawn 1,085 s 05 corn 1,104 s
02 unfurl 1,088 s 06 gonzo 1,107 s
03 squirrel 1,090 s 07 today 1,102 s
04 1971 1,100 s 08 signup 1,103 s
total 8,779 s

Two hours and twenty-six minutes of generation for 41 seconds of picture, with 2% spread across the eight runs. That is about 210 seconds of compute for every second of screen time, on the pass where nothing went wrong.

Which makes 8,779 seconds a floor and not a total, because the film that shipped is not that pass. One of those eight segments was cut from the edit, one was split into two shots that had to be generated separately, and others were regenerated after review. Everything below is what that difference is made of.

Two knobs, multiplied

The settings above were chosen from a sweep, and the sweep has one result worth carrying to any diffusion video model. Going from 960×544 at 124 frames to 1344×768 at 243 frames doubles the pixels and doubles the frames, roughly 3.9× the tokens. Time per step went from 17.1 seconds to 140 seconds, a factor of 8.2. Attention is quadratic in sequence length, so the two knobs do not add, they multiply and then get squared.

resolution and length clip wall clock
960×544, 124 frames 5.2 s 410 s
1344×768, 124 frames (the film) 5.2 s 1,085 to 1,107 s
1344×768, 243 frames 10.1 s about 50 min
1344×768, 362 frames 15.1 s 4,877 s

So ten seconds is the knee. One ten second take costs fifty minutes and one fifteen second take costs eighty-one, against about eighteen minutes for five seconds at the same resolution. Cut more short segments rather than one long one, and as a bonus you can re-run one of them without paying for the rest.

One thing that is nearly free on this machine: memory contention. Running the same job with a 27B language model holding 73 of the 122 GB of unified memory moved time per step from 17.1 to 18.3 seconds, about 6%. Whether to stop the other model is a question about 6%, not about whether the job fits. The bottleneck here is arithmetic, not memory.

What the retakes cost

The largest single item is one shot. A five-second segment came back with a performance problem: the character laughs too hard, and the frames sit in the uncanny valley. The seed is fixed, so re-running the same prompt returns an identical clip. The only lever is the text.

Five rewrites were generated and compared side by side. The four that were timed took 1,310 to 1,316 seconds each, so about an hour and fifty minutes of generation to change one shot, roughly three quarters of what the entire first pass of the whole film had cost.

Two rows of five frames of the same animated character at the same moment. In the top row her eyes are squeezed shut and her inner brows are raised; in the bottom row her eyes are open and looking around and her brows are level.
Two of the five rewrites at the same five frames. Top: the version that moved the laugh into the eyes. Bottom: the version that occludes the peak of each laugh with a prop. Both fixed the complaint about the mouth, and the whole decision came down to the eyebrows: the top row reads as suppressed pain, the bottom row is present and looking at something.

The comparison paid for itself with a finding. A rewrite that replaced the laughter with other physical actions, holding it in, turning away, pressing the lips together, came back looking like distress: brows knitted, eyes shut, mouth turned down. Closing the mouth without pinning the brow makes the model render suffering. The missing constraint was never about the mouth at all.

A second branch produced a cleaner negative. Two segments were merged into a single ten second take to buy out the cut between them, three drafts deep. The rule that came out of it: in this prompt format, a timecode at the start of a sentence is itself an instruction to cut. Writing "At about 00:04.900, WITHOUT ANY CUT, the camera gives ground" produced a cut; deleting the timecode from that sentence removed it. The capitalised negation did not beat the sentence pattern. The merged version was abandoned anyway, in favour of a two frame transition.

Where the model has to be helped

Text does not lock architecture. The first pass described the real quad in words, a domed columned stone auditorium, long rows of tall elms, and got a generic east coast campus back: white New England spires, a little white rotunda, red brick walkways, autumn foliage. Some of that was the prompt's own fault, since the real quad has pale concrete paths and is deep green in late August, but correcting the words fixed only half of it. The place had to arrive as an image instead: real photographs as geometry reference, frames of the film itself as style reference, an image model in between, and the result locked in as the video model's first frame. Photographs alone produce a photoreal matte painting that does not cut with animation.

Locking the first frame is not enough. One segment given only a first frame walked the camera into the tents and lost the only thing the shot existed to show, which was the scale of the lawn. First and last are both locked now.

Lettering does not survive. Short signs in the subject position of the frame come out clean, and a hand-painted "CORN ON THE QUAD" board did so on the first try. Small lettering in the background is always mush. The end title is burned in afterwards with ffmpeg, for repeatability rather than for looks.

Period look is a grade, not a prompt. The prompt for the 1971 sequence asked for warm grainy film stock, halation, faded colour and gate weave. The model returned exactly the right content, a folk trio, a bow-tied dean, hand-painted signs, and clean modern rendering. The look came from a colour grade in post, which also let the switch land on the precise frame she passes through the curtain of light.

What none of it bought

The film was posted to the campus subreddit with a short disclosure of how it was made. Most of the replies read it as a one-shot: type a prompt, get a film. The post was deleted.

The diagnosis is worth more than the film. A one-line disclosure reads exactly like no disclosure. Everything that constitutes the human work here, the storyboard, a locked first and last frame per segment, five rewrites of one performance, the grade, the cuts, the audio surgery, lived outside the artifact, in a repository nobody has any reason to open. Anything published as a finished object gets judged as a finished object. If the work is to count, the evidence of it has to be inside the thing, not next to it.

That is the same instinct the rest of this site runs on, which is why every note here carries its measurements rather than pointing at them.

Made at home. Not an official University of Illinois production. Music: "Under the Sun" by Michael Ramir C., from Mixkit. The stock music licence cannot be passed on and the film shows a university trademark, so the film itself is all rights reserved and is not offered under an open licence. The numbers and the findings on this page are free to use.

The companion note on this cluster is rebuilding the same quad from public data, where the medium is measurement rather than invention.