Skip to content

The Text-First Trick Behind Every LLM Video Editor

#video-generation #llm-agents #ffmpeg #automation #multimodal

The Text-First Trick Behind Every LLM Video Editor ​

There's a moment in every LLM video editing project where you realize the model can't actually watch the video. Not really. It can process frames, but dump 30,000 of them into a context window and you're looking at 45 million tokens of noise. The output quality doesn't justify the cost, and the model still can't tell you which clip the good take is in.

The tools that actually work this year all converged on the same answer: turn the video into text first, let the LLM reason over the text, and hand the resulting edit list to ffmpeg. Three projects shipped this pattern in the last few months. ReelCraft, a Python CLI for turning phone photos and clips into 9:16 shorts. video-use, from the browser-use team, which edits raw footage through Claude Code. And MoneyPrinterTurbo, the most-starred automated short-video pipeline on GitHub. The differences between them tell you everything about where this space is heading.

The problem: video is a terrible input format for an LLM ​

Start with what the model vendors actually support. Gemini Omni Flash, the new video editing model, has a limitation section that reads like a dare: "Referencing or reasoning across multiple videos is not supported. Attempting multi-video prompting may result in degraded model performance or unexpected outputs." Video references up to 3 seconds are accepted by the API schema but not correctly processed.

So the obvious approach, throwing a folder of footage at a model and asking it to edit, is dead on arrival. The standard Gemini models can handle up to 10 videos per request with a 1M token context, which covers about an hour of footage at default resolution. But ask for highlights across ten videos in one prompt and the model confuses the timelines. It loses track of which second belongs to which clip.

The workaround is to stop thinking of the model as a video editor and start thinking of it as a film critic who writes detailed notes. You give it one video at a time, it produces a structured analysis. Then you feed all the analyses, as text, to a second call that makes the actual editing decisions.

Key numbers: 1M token context handles roughly an hour of footage. 10 videos max per request, and only if you accept timeline confusion. 3 seconds is the longest video reference Omni Flash actually processes, regardless of what the schema accepts. 12KB of packed transcript replaces 45M tokens of raw frames in the video-use pipeline.

The shared architecture: transcribe, reason, render ​

Every working pipeline follows the same five stages.

The analysis layer is where the real engineering happens. ReelCraft calls Gemini once per asset to get precise internal timestamps and descriptions. That per-file approach sidesteps the multi-video limitation entirely, and a failed analysis doesn't take down the whole batch. Three retries, then the failure gets logged and the pipeline moves on.

video-use takes a different route into the same architecture. It runs ElevenLabs Scribe over every take to get word-level timestamps, speaker diarization, and audio events like laughter or applause. The whole session packs into a single 12KB markdown file that the LLM reads as its primary view. A visual composite, filmstrip plus waveform plus word labels, is generated on demand only at decision points: ambiguous pauses, retake comparisons, cut-point sanity checks. The model never watches the video, but it always knows exactly what was said and when.

MoneyPrinterTurbo goes even further upstream. You give it a topic, it generates the script, extracts material search keywords, pulls stock footage from Pexels or Pixabay, and composites the final video. The LLM never sees any footage at all. It writes the words and picks the search terms; the matching happens outside the model.

Quick Take: the winning pattern is to compress video into a text representation the LLM can reason over, then let deterministic tooling do the actual cutting.

Three pipelines, one pattern ​

ReelCraftvideo-useMoneyPrinterTurbo
LLMGemini 3.7 FlashClaude CodeKimi, Gemini, DeepSeek, others
Inputphotos + clipsraw footagetopic or keyword
Text representationper-file analysis JSONword-level transcript + diarizationscript + search keywords
Human checkpointedl.yamlstrategy approvalWebUI review
Rendererffmpeg xfadeffmpeg + parallel animation agentsffmpeg composite
MusicLyria 3 generatednonelocal BGM files
Extrasauto-burned subtitlesself-eval loop, color gradeauto-publish to TikTok, IG, YouTube

The human checkpoint is the piece most people skip and regret. I decided from the start not to make my pipeline one-click fully automatic. LLM-provided edit points will inevitably have irrationalities, and re-running the whole pipeline costs API money again. The intermediate artifact is a YAML file where each clip has a source path, an in/out range, and a note explaining why the model chose it. Change a number to move an edit point. Move a line to reorder a clip. Save, render.

video-use builds the same checkpoint into its agent loop. The LLM inventories the sources, proposes a strategy, and waits for approval before touching the cut. Then it renders, evaluates its own output at every cut boundary, and fixes issues up to three times before showing you anything.

The pattern is spreading beyond these three. Community Spaces like reel-lab and Omni-videos-custom are wrapping the same idea in Gradio UIs, and Lightricks keeps shipping generation models like LTX-2.5 that feed the other end of the pipeline. The generation side gets the headlines. The editing side is where the actual workflow value sits.

Model choice matters more than you'd think ​

When I switched my analysis stage from gemini-2.5-flash to gemini-3.7-flash, the descriptions changed in ways that directly affected edit quality. For the same lecture video, 2.5 wrote: "a woman on stage uses a microphone to introduce herself... The large screen behind her shows her name and her job description." 3.7 wrote: "a female speaker (Zona Wang, LINE Technology Evangelist) is giving a self-introduction... followed by a camera pan across the audience."

The difference is that 3.7 actually read the small text on the slide. The aggregation stage showed an even bigger gap. 2.5 described a conference as "vitality and diversity." 3.7 named the full event, identified booth names, and described specific photos down to the semiconductor-shaped snacks being handed out. None of those details were in my prompt. They came from OCR-level reading of the actual assets.

For an application where asset understanding quality directly determines editing quality, that's the difference between a generic slideshow and a video that feels like it was made by someone who was there.

The rendering layer is where things break ​

Every project in this space hits the same wall: ffmpeg returns exit code 0 and produces a video that's subtly wrong. I found two specific failure modes, and both are the kind of thing you only catch by actually playing the output.

First, a transition longer than the clip it connects. Two 1-second clips with a 2-second crossfade produce a negative offset. ffmpeg accepts the negative number, finishes normally, and the second clip silently disappears from the output. Since the transition field is free text, typing "3s" instead of "0.3s" in the EDL gives you no warning at all.

Second, an out-point that exceeds the actual asset length. A 10-second video with an EDL range of 8.0 to 15.0 only yields 2 seconds of usable footage. Everything scheduled after it gets truncated or dropped. This one matters more because the EDL is LLM-generated, and hallucinating an out-of-bounds end time is the most natural failure mode in the world.

The fix is explicit validation before rendering: throw an error on negative offsets, and ffprobe every asset to check requested ranges against actual durations. A crash you can fix. A "successful" render with a missing segment means you don't discover the problem until you watch the whole thing and think "wait, where did that clip go?"

Subtitles have their own silent failure. I burn SRT subtitles using libass rather than drawtext, because drawtext requires manual font path handling and escaping hell. But the first version set each subtitle's display interval to its clip's start and end times. With a 0.3s crossfade between adjacent clips, that overlap produced two lines of white text stacked on screen. The fix: end each subtitle when the next clip starts, so at most one line is visible at any moment.

video-use adds a 30ms audio fade at every cut so you never hear a pop. Small detail, but it's the difference between output that feels edited and output that feels spliced.

Common pitfalls ​

The pattern is consistent across all three projects. Get these wrong and you'll ship broken videos.

The most dangerous failure mode is trusting ffmpeg's exit code. It returns 0 on outputs that are missing clips or truncated segments. Validate offsets and asset durations before rendering, and watch the output at every cut boundary.

Avoid feeding multiple videos to a single analysis prompt. The models degrade badly at cross-video reasoning, confusing timelines and attributing moments to the wrong clip. Per-file analysis with a text aggregation step costs a few extra API calls and eliminates an entire class of errors.

Don't assume the music API works like the text API. Lyria 3 uses a completely different calling convention with no structured parameters. Length, BPM, and mood all go into a natural language prompt, generation is single-turn with no iteration, and audio comes back with a SynthID watermark. Plan for it to fail and fall back to silence rather than crashing the render.

Subtitle intervals need their own care when transitions are involved. If each subtitle spans its clip's full range, the crossfade overlap stacks two lines of text on screen. End each subtitle when the next clip starts.

The human checkpoint is not optional. LLM edit points have irrationalities baked in, and re-running an entire pipeline to fix one bad cut wastes API credits. A YAML file or strategy approval step costs minutes and saves reruns.

One thing to remember ​

Keep the LLM on reasoning, keep ffmpeg on rendering. Every tool that works treats the model as the layer that produces an edit decision list, and treats ffmpeg as the deterministic executor. The moment you blur those roles, you inherit every failure mode at once.

The Bottom Line ​

If you're building a video editing pipeline, adopt the text-first architecture: per-file analysis or transcription, a packed text representation, an EDL with a human review point, and ffmpeg for rendering. It's the only pattern that scales past a handful of clips.

If you're constrained by API costs, the 12KB transcript approach from video-use is the cheapest path. It replaces millions of tokens of frame data with a few thousand tokens of text, and generates visuals only at decision points.

One thing to watch: the model vendors are closing the gap on native video understanding. Omni Flash still can't reason across multiple videos, but that limitation is explicitly temporary. When multi-video reasoning lands, the per-file analysis workaround becomes optional. The text-first architecture survives either way, because the human checkpoint and the ffmpeg validation layer solve problems that have nothing to do with model capabilities.