
How to Make Videos With AI for Free
I made a whole short video using only free AI tools and no camera, no studio, nothing, here's the honest breakdown of what worked, what was frustrating, and what the final product actually looked like.
Real talk: when I first heard people were making videos entirely with AI tools, I assumed the results looked terrible and the process was a mess. I was half right. Some of it is still pretty rough. But I made a two-minute explainer video using only free tools and no camera, and I'm genuinely surprised by how it turned out.
This is a real walkthrough. I'll tell you exactly what I used, what the free tiers actually give you, and where the friction is, because there's friction and I'm not going to pretend there isn't.
First: what kind of video are we talking about
Before you start downloading apps and signing up for accounts, you need to decide what type of video you're making because the toolchain is completely different depending on your answer.
If you want a video with a realistic AI presenter or avatar talking to camera, that's one path. If you want a video with AI-generated images and a voiceover, that's another. If you want animated explainer-style content, that's a third. I'm going to focus on the approach I actually used: AI voiceover over AI-generated images, stitched together with transitions. This is the easiest to do for free and produces a watchable result.
Step one: write the script with AI
I wrote a script for a two-minute video explaining how to make cold brew coffee at home. Nothing fancy. I just wanted to test the whole pipeline end-to-end with a topic I knew well enough that I could catch errors.
I used Claude to help me tighten the script. Not to write it from scratch. I wrote a rough version first and then asked for feedback on pacing and clarity. The main thing it caught was that I was front-loading too much information before getting to the actual instructions, which would have lost people in the first 30 seconds. Good catch.
Two minutes of video is roughly 260-300 words of script. That's shorter than you think. I kept trying to make it longer and had to cut myself back.
Step two: turn the script into voiceover
ElevenLabs has a free tier that gives you 10,000 characters per month. For a short video that's plenty. The voice quality on the free tier is genuinely good. I used a voice called "Rachel" which sounds like a real person and not a robot reading at you.
The thing I'd suggest: read your script out loud before generating the audio. The AI reads exactly what you write, so weird punctuation produces weird pauses and run-on sentences produce breathless audio. I had to go back and add commas and period breaks to places where I'd written continuous sentences, just to get the pacing right.
Once you've got the audio file, download it as an MP3. That's your backbone, everything else gets timed to this.
Step three: generate the images
This is where it gets a little more involved. You need one image for roughly every 5-10 seconds of video, so for a two-minute video that's 12-24 images. Free image generation tools have daily limits so you might need to spread this across a couple of days or use a few different tools.
I used a combination of tools here. Leonardo AI has a free tier with a daily token allowance. Microsoft Designer (which runs on DALL-E) is free with a Microsoft account. Between the two I had enough to generate everything I needed.
For consistency across images, which matters a lot for a video because you want it to look like one coherent thing. I kept my prompts really consistent. I picked a style early (I went with clean, bright, slightly stylized food photography style) and described it the same way in every prompt. "Warm natural lighting, shallow depth of field, minimal background" on every single image. That consistency is what makes the final video not look like a Frankenstein of different aesthetics.
Honestly this was the most time-consuming part. Image generation with free tools is slow and the results are unpredictable. I generated probably 60 images to end up with 20 I was happy with. Plan for that.
Step four: put it together
CapCut has a free desktop and mobile version that's more than capable for this. I'm going to be honest that the free version has a CapCut watermark if you export at high quality. There are ways around this, you can export at lower quality watermark-free, or use DaVinci Resolve which is entirely free and watermark-free but has a much steeper learning curve.
For my test video I used CapCut because I just wanted to see the pipeline work end-to-end without spending hours learning new software. The watermark bothered me aesthetically but not enough to derail the whole test.
The actual assembly process: import your voiceover audio first and lay it on the timeline. Then start adding your images on top, roughly timed to when the voiceover is talking about that thing. Ken Burns effect (slow pan/zoom on static images) makes it feel way more dynamic than just static slides. CapCut has a one-click auto-apply for this and it's fine.
Add captions if you want them. CapCut can auto-generate captions from your audio and they're pretty accurate. I had to fix maybe four words across the whole two-minute video. Captions matter because a lot of people watch short videos without sound.
Where the free tier walls actually hit you
ElevenLabs free tier resets monthly and 10,000 characters goes fast if you're making longer videos. A five-minute video would burn through your monthly allowance in one go. For anything beyond short content you'd need to pay.
Image generation daily limits are the real bottleneck. If you need 30+ images for a longer video you're either waiting multiple days to generate them all on free tiers or juggling multiple accounts across multiple platforms. I did both of these things. It's annoying but it works.
The watermark situation on CapCut is genuinely frustrating. DaVinci Resolve is the real answer if you want to do this properly for free, but you have to be willing to invest some time learning it. There are good beginner tutorials on YouTube and for basic cut-and-paste video editing you can get functional pretty fast.
What the final video actually looked like
Honestly? Pretty good for something made entirely with free tools and no camera. The voiceover sounded natural. The images were consistent and looked intentional. The pacing worked. If I hadn't told you it was made with AI you probably would have assumed it was a well-made but modest explainer video from a food creator.
What would have made it better: more time on image selection and prompting, a better editor than CapCut free tier, and probably one or two more rounds of script tightening. All of that is doable with free tools, it just takes longer.
I could be wrong but I think a lot of people are surprised by how accessible this pipeline is. You don't need a camera. You don't need a studio. You don't need to spend money. You need a few free accounts, a couple of afternoons, and the patience to generate way more images than you'll end up using.
If you're trying to test a content idea before investing in real production, or you want to start building an audience before you have any equipment, this pipeline is a completely reasonable place to start. It won't replace great production quality. But it's genuinely workable for what it is.
Emily in AI
Emily in AI is a plain-English guide to AI tools, tips, and beginner guides. Every tool gets tested and written up without the hype or the jargon, so you can figure out what actually helps. New posts every week.
About Emily in AI →

