FilmGPT

Autoregressive Modeling of Film
with Applications in Video Montage

Marcelo Sandoval-Castañeda¹, Fabian Caba Heilbron², Shiry Ginosar¹, Bryan Russell²,
Josef Sivic²³, Alexei A. Efros²⁴, Greg Shakhnarovich¹
¹TTI-Chicago ²Adobe ³CIIRC, CTU ⁴UC Berkeley

TTI-Chicago
Adobe
CIIRC, CTU
Berkeley AI Research

FilmGPT is an autoregressive model for video that assembles edited videos from massive collections of real raw footage.

Paper – Code (coming soon)

Abstract

This work introduces FilmGPT, an autoregressive transformer designed to address the challenge of video montage—turning a collection of raw, “unwatchable” footage into coherent cinematic sequences. Inspired by language learning in modern LLMs, we train a long-context autoregressive transformer on a large corpus of movies. The aim is to implicitly capture the “grammar” of film directly from data rather than from hand-coded rules. Unlike other generative models, FilmGPT does not generate any new video frames. Instead, at inference time, we introduce a footage-constrained decoding algorithm to select the best next shot from the input raw footage according to the statistical patterns learned from films. We first evaluate these learned statistics directly by using the FilmGPT autoregressive model for next shot prediction on a standard benchmark of shot sequence ordering, outperforming the previous state of the art. We then evaluate our footage-constrained decoding algorithm on the full film editing task via a user study, and find that our FilmGPT-based editing significantly outperforms previous approaches. Finally, we demonstrate the applicability of FilmGPT to a wide range of applications in video montage, from automatic video segment trimming to human-in-the-loop film editing.

Results
Segment Trimming

Aligned Multi-Camera Editing

Automatic B-Roll

Human-in-the-Loop B-Roll


Movie Idioms

Cutting on Action

Point-of-View Shot

Establishing Sequence

BibTeX
@inproceedings{sandoval2026filmgpt,
author = {Sandoval-Casta\~{n}eda, Marcelo and Caba Heilbron, Fabian and Ginosar, Shiry and Russell, Bryan and Sivic, Josef and Efros, Alexei A. and Shakhnarovich, Gregory},
title = {Autoregressive Modeling of Film with Applications in Video Montage},
booktitle = {Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference (SIGGRAPH)},
year = {2026},
}