
ElevenLabs Tutorial: Your First 10 Minutes
A friendly, click-by-click ElevenLabs walkthrough so you can make a clean voiceover today, plus the settings I wish I had known about sooner.
The first time I opened ElevenLabs, I pasted a paragraph, hit generate, and got a robot reading a ransom note. Too fast, weirdly emotional, kind of unsettling. So this ElevenLabs tutorial is the version I wish someone had handed me. We're going to go from a blank account to a clean, natural voiceover in about ten minutes, and I'll point out the two sliders that fix 90 percent of beginner problems.
I'm assuming you have zero experience and just want something that sounds good without reading a manual. That's the whole vibe here. No jargon, no twenty-tab research project, just the handful of choices that actually change how your audio sounds.
Make your account and find the playground
Sign up at elevenlabs.io. The Free plan gives you 10,000 credits a month, which is roughly 10 minutes of speech, and lets you create up to three custom voices. That's plenty for testing. One honest catch: the Free plan does not include commercial rights. If you're publishing this for a business or a monetized channel, you need at least the Starter plan, which is 6 dollars a month. For just learning the tool, Free is genuinely fine, so don't feel pushed to pay on day one.
Once you're in, look for Text to Speech in the left menu. That's the playground. You'll see a big text box in the middle and a row of controls around it. Ignore most of it for now. The screen looks busier than it is, and you really only need three things to make your first clip: a voice, some text, and the generate button.
Your first generation, step by step
Here's the fastest path to a real result:
- Pick a voice. Bottom left of the screen, you'll see your voices. Start with one of the built-in ones like Rachel or Adam. Don't overthink it. You can swap later in two seconds.
- Paste a short bit of text. Two or three sentences, not a whole essay. You want to hear a sample fast, and short tests cost almost nothing in credits.
- Choose a model. Multilingual v2 is the high-quality default. The Turbo and Flash models are faster and cost about half the credits, but I'd start with v2 so you hear the best version first, then decide if speed matters more.
- Hit Generate. The audio plays right there, and you can download it as an MP3 with one click.
That's it. You've made a voiceover. Genuinely, that's the whole core loop. Now let's make it not sound weird, because the default settings are not always the flattering ones.
The two sliders that fix everything
This is the part I needed someone to explain in plain English. Under the voice, you'll find Stability and Similarity. Almost every beginner problem traces back to one of these two being off.
Stability controls how consistent the voice is between runs. Low stability gives you more emotion and variation, but push it too low and the voice rushes and gets dramatic in a way nobody asked for. Too high and it goes flat and monotone, like a hold-music recording. A good starting point is around 50, which keeps some life in the delivery without letting it run wild.
Similarity controls how closely the AI sticks to the original voice. Around 75 is a solid default. If you ever clone your own voice from a noisy recording, keep this lower, because cranking it high makes the AI copy the background hiss and room echo too. High similarity on clean source audio is great. High similarity on bad source audio just faithfully reproduces the badness.
One thing nobody tells you: these sliders are not exact. ElevenLabs is not deterministic, so the same settings can give slightly different takes each time. Think of them as a range, not a dial with one correct number. If a take sounds off, just regenerate. It's often a one-click fix, and that surprised me at first because I assumed I had set something wrong.
Speed and small tweaks
There's a Speed setting too. Default is 1.0. You can slow it down to 0.7 or speed it up to 1.2, and honestly that range is narrow on purpose, because pushing speech too far in either direction sounds artificial. For narration I usually nudge it to about 0.95 so it feels less hurried. Small change, big difference, especially for longer listening.
For punctuation, commas and periods actually matter. They create natural pauses. If a sentence reads too fast, breaking it into two sentences usually sounds better than fiddling with sliders. The model reads your formatting, so write the way you'd want it spoken. Ellipses can add a beat of hesitation, and a paragraph break gives a slightly longer rest. Once you notice this, you start writing scripts a little differently, in a good way.
A quick workflow that just works
Here's the loop I use every time so I'm not stuck regenerating forever:
- Write the script in short sentences, the way a person actually talks, not the way you write an email.
- Set stability to 50, similarity to 75, speed to about 0.95.
- Generate a short test chunk first, not the whole thing. Confirm the voice and tone feel right before committing.
- If it sounds rushed or flat, regenerate before touching sliders. Half the time the next take is the one you wanted.
- Once a chunk sounds right, do the rest in similar-length pieces. Long blocks are harder to fix and tend to drift in energy near the end.
Working in chunks also protects your credits. You burn 1 credit per character on v2, so a few short tests cost almost nothing, while regenerating a giant block over and over adds up fast. I learned that the expensive way, watching my monthly allowance disappear on one stubborn paragraph.
Cloning your own voice, briefly
People always ask about this, so here's the honest version. Instant voice cloning, where you give it a short sample, is available starting on the Starter plan. It's fun and fast, but the quality leans on how clean your recording is. Record somewhere quiet, close to the mic, no fan or traffic in the background. The professional-grade cloning that sounds eerily like you lives on the Creator plan at 22 dollars a month and uses more training audio. If a casual clone is all you need, Starter handles it. If you want it to truly pass as you for serious work, that's the Creator tier.
Who should skip ElevenLabs for now
I like it a lot, but it's not for everyone. If you only need a voiceover once or twice a year, the Free plan covers you, but the commercial-rights catch means casual business use pushes you into a paid plan. If you need hours of audio every month, watch your credit math, because the cheaper tiers run out faster than you'd expect. And if you want pro voice cloning trained on your real voice, that lives on the Creator plan, not the cheaper tiers, so budget accordingly. None of this is a dealbreaker, it's just the stuff I'd want a friend to know before they got attached.
The Bottom Line
You can get a genuinely good voiceover out of ElevenLabs in your first ten minutes. Pick a built-in voice, set stability to 50 and similarity to 75, test in short chunks, and regenerate freely. Skip the deep settings until you actually need them. Start on Free to learn the feel, then upgrade only when you hit the commercial-rights wall or run low on credits. That's the whole ElevenLabs tutorial, and the rest is just practice and trusting your ears.
Emily in AI
Emily in AI is a plain-English guide to AI tools, tips, and beginner guides. Every tool gets tested and written up without the hype or the jargon, so you can figure out what actually helps. New posts every week.
About Emily in AI →

