An AI video generator creates footage that did not exist before, an AI-assisted editing tool works on footage you recorded. Two tools for two starting points, which a lot of comparisons online throw into one pot. The difference shows up fastest in the unit each one works in: a generator delivers a clip of a few seconds. Google's Veo 3.1 produces videos of 4, 6 or 8 seconds according to the developer documentation, 720p by default, 1080p and 4K only at the 8-second length, in 16:9 or 9:16, and it generates the audio along with the picture. In the Gemini app, according to Google's product page, the model behind it is now Gemini Omni 1.1 Flash, with 10 seconds per video. An editing program, by contrast, works in minutes: on what is already there. In practice that means the generator is strong where no particular person and no particular place has to be shown. If a face is supposed to be recognised as yours, though, the avatar route starts with a photo or a short video recording. All model names, lengths and prices in this article were checked against the provider pages on 20 September 2026.
I do not generate my reels, I record them and have them edited. Generated material only shows up here as supply, a few seconds of B-roll in a video that otherwise consists of real recordings. Everything I collect about AI at work is on the AI page; how I decide what gets handed off at all is in Which Tasks You Can Hand Off to AI.
Two tools, two starting points
The row that explains the most is the third one. A generator delivers one shot, a video consists of many of them: in an order, with transitions, with audio that carries across the cut. Producing that order is editing work, no matter where the individual shots came from. So anyone searching for "create AI video from text" and expecting a finished five-minute video is really looking for two tools, not one.
What the generator is good for
The generator plays to its strength wherever you want to show something you cannot shoot or do not want to shoot: a camera move across a landscape that does not exist, an abstract scene as an underlay, an object for which there is no set and no shooting day.
Prices differ less in amount than in model. Runway bills in credits: the Standard plan costs 12 US dollars a month billed annually according to the pricing page and includes 625 credits, which the provider puts at roughly 52 seconds of Gen-4.5 or 104 seconds of Gen-4 Turbo, with watermarks gone from that tier up. Free access consists of 125 one-time credits, which is a few attempts, not a working mode. At Google, video generation in the Gemini app depends on a paid subscription, features differ by plan and region, and the minimum age is 18. Pricing pages retrieved on 20 September 2026; tariffs and credit allowances change quickly in this field.
Without a provider account, what remains is open model weights on your own hardware, and then you pay in compute time instead of money. Alibaba's Wan 2.2 is under Apache 2.0 and creates five-second clips in 480P or 720P; the large A14B variant states at least 80 GB of VRAM for a single graphics card in its model card, while the smaller TI2V-5B runs on a consumer card such as the 4090 according to the same card and produces a 720P video in under nine minutes. LTX-Video from Lightricks delivers up to 1216 by 704 pixels at 30 frames per second and up to 257 individual frames per generation according to its model card, again in the range of seconds. With open weights it pays to look past the licence line: the model terms, the provenance of the training data and the hardware requirement all decide whether the route holds up. Anyone who would rather not work through the setup alone brings the question into the community.
Three points belong in the planning. First, length: at eight or ten seconds per clip you need six to eight generations for a minute on paper. On paper is not finished: not every attempt is usable, rejects and retries belong in the calculation, and a minute that holds together only comes out of the edit. Second, recognisability across several clips, the same problem as with image series; what the providers offer against it is in AI Image Generation Compared. Third, storage: Google writes in the Veo documentation that generated videos sit on the server for two days and are removed after that. Whatever you want to keep, you download.
What the editing program is good for
Everything you recorded yourself. That sounds banal and is still the dividing line: as soon as footage exists, the task is no longer "create", it is "select, trim, label, deliver". The AI mostly does supporting work here, transcript and subtitles and framing, and you or someone acting for you makes the editing decisions. Which program takes over which part, and where the limit runs, I went through in AI Video Editing Software Compared, so there is no second table here.
One point belongs here anyway, because it applies to generators just as much: with browser services, your raw footage travels to someone else's servers. If clients, staff or conversations appear in it, that is a deliberate decision. How I handle it is in AI and Privacy: What the AI Gets to See.
The mixed form that comes up most often in practice
The normal case is neither one nor the other, it is your own footage with generated inserts. You stand in front of the camera yourself and say what this is about, and in three places something sits on top that you did not shoot: an illustration, a scene for which no footage exists, a transition.
Four things keep that mixture usable day to day:
- Set the format first. If the finished video is 9:16, generate the inserts in 9:16 right away. Veo supports both aspect ratios according to the documentation; cropping afterwards costs picture area.
- Keep them short. Generated shots stand out less the shorter they stay on screen.
- Match the look. A generated shot that is brighter, smoother or more saturated than your recording reads as a foreign body. That is editing work, not prompt work.
- Label it where it could mislead. More on that below.
In order, that means: record first, then review, then name the gaps, then generate. Anyone starting from the other end builds the video around the clips the generator happened to get right.
A face that speaks for you
This is the line the whole comparison rests on, and it runs finer than it looks. A talking face is something the avatar providers deliver without a shoot: Synthesia names more than 240 ready-made avatars on its avatar page and says a first video with them needs no filming and no consent process, HeyGen more than 1,100, cleared for commercial use according to the provider page. That is then someone else's face speaking your text. If your own face is meant to speak, the one clients recognise, the avatar route starts with a photo or a short video recording, depending on the provider. Synthesia also requires a live on-camera consent verification that, in the provider's words, "cannot be uploaded or bypassed", and the person giving consent has to be the same person as in the avatar video, at least 18 years old. HeyGen names a short clip as the entry point, says 30 seconds is enough, offers a photo as an alternative, and requires demonstrable consent from the person being reproduced for custom avatars.
In other words: the route that looks like "video without a shoot" needs one clean piece of source material from you. After that, the avatar can speak in many videos. And where someone else's face is involved, German law applies anyway: section 22 of the Kunsturhebergesetz says that portraits may only be distributed or publicly displayed with the consent of the person depicted. This article is not legal advice, but the direction is clear enough to build into your choice of tool.
Read the other way around, that is the good news for the generator: where nobody in particular has to be on screen, exactly this problem disappears. Landscape, abstraction, object, mood. That is what it was built for.
Labelling generated video
Labelling here is not one duty but two, and they fall on different parties. Technically: Google writes in the Veo documentation that videos created with Veo are watermarked using SynthID, and names the same invisible watermark for the Gemini app, where you can check a file by uploading it. Legally: Article 50 of the AI Act has applied since 2 August 2026. Paragraph 2 obliges the providers of the AI systems to mark their outputs in a machine-readable format as artificially generated; that mark is not visible, and where a tool only performs an assistive function for standard editing, the European Commission says the duty does not apply at all. For systems placed on the market before 2 August 2026 the Commission names, in its FAQ with a page status of 24 July 2026, a grace period until 2 December 2026, and only for this marking; content created before 2 August 2026 does not have to be labelled retroactively. Under paragraph 4 you have a disclosure duty of your own for deepfakes, at the latest on first exposure and in a clearly recognisable manner. You cannot rely on the provider's marking to discharge it, the Commission states: recognisable means perceivable without technical tools. For special effects in film productions and for artistic, creative and satirical works the Commission names exemptions or a limited duty.
What applies to images, which markers the individual providers set and how easily they get lost on upload, I documented in Creating AI Images: Free and Free to Use; for video the situation is the same, so this is just the pointer. This is not legal advice either.
How it runs here
In my business, one employee makes this distinction visible. Eddi is my AI employee for video: he gets raw footage and a script and returns a finished, technically checked video, pauses out, subtitles on, overlays placed, format 9:16. He works on the command line with ffmpeg, a locally running Whisper model and Remotion, not in an editing program with a mouse. He does not shoot anything himself, and he never overwrites raw footage. AI generations cost money, so they only run with my approval, and publishing is a gate as well.
That puts the answer to the generator question in one sentence: a generator creates moving images from a prompt, Eddi edits real footage. Both can appear in the same video, and neither of them replaces the recording you are in. How a role like that comes about is in Hiring an AI Employee: The Process; how the tasks are split up here is in My AI Workforce.
Frequently asked questions
What is the difference between AI video editing and an AI video generator?
The generator creates footage that did not exist before, from a text and often from a reference image as well. AI video editing works on footage you recorded: trimming, subtitling, setting the format. The generator delivers individual shots of a few seconds, the edit turns them into a sequence. You can combine the two, you cannot swap one for the other.
Can I create a finished video from a text?
A clip yes, a finished longer video not in one go. Veo 3.1 creates videos of 4, 6 or 8 seconds with natively generated audio according to Google, and in the Gemini app it is 10 seconds per video. For a minute you therefore need several generations that fit together, plus the edit that connects them.
Can I create a video from an image and a text?
Yes, that is the common route at the big providers. Google names up to five photo references for video generation in the Gemini app. A reference image helps above all where the look is meant to stay the same across several clips. It is not a guarantee of recognisability, and for faces the consent rules apply on top.
Are there free AI video generators, or ones without sign-up?
Free tiers yes, permanently and without an account rather not. Runway states 125 one-time credits for free access, which covers attempts, not ongoing work. Video generation in the Gemini app requires a paid Google AI subscription and a minimum age of 18. Without a provider account, the local route via open model weights remains.
Are there open source video generators?
Yes, with hardware as the price. Wan 2.2 from Alibaba is under Apache 2.0 and creates five-second clips in 480P or 720P; the large variant states at least 80 GB of VRAM, the smaller one runs on a 4090 according to the model card. LTX-Video from Lightricks delivers up to 1216 by 704 pixels at 30 frames per second. Setup and compute time are on you.
Do I have to label an AI video?
Under Article 50(2) of the AI Act, the duty to mark content in a machine-readable format falls on the providers of the systems, not on you. Under paragraph 4 you have a disclosure duty of your own for deepfakes, at the latest on first exposure and clearly recognisable; the provider's invisible mark does not discharge it. None of this replaces legal advice.
Where to go from here
Run the test on a video that is coming up anyway. Write two lists: the shots you record yourself, and the ones you cannot shoot. Generate only for the second list. How long those two lists turn out answers the original question before you open a single tool.
If it turns out that what is missing is not a tool but someone who turns raw footage into a finished video every week: that exact process, from raw footage through the edit to approval, is what I show in my community, together with the ready-made templates for the employees who handle it here.