Why doesn't the teddy move its lips?
Common lip-sync models are designed for human faces. Kling AI explicitly excludes animal characters. A talking teddy needs an avatar model that supports animal characters, and a paid plan.
Promo video · Storyboard · Guide
The Teddynews teddy introduces the company. The video runs 64 seconds. It has a German AI voice, burned-in German subtitles and no music. The footage comes from existing material, plus 5 cards in the Teddynews style. The voice cost 0.03 US dollars. Everything else ran on our own computer.
Below you will find the video, all 10 stills of the sequence, the full script in English translation and a guide to building your own.
No sign-up, no account. The video only starts when you click.
The video
The video runs 64 seconds and contains speech but no music. Voice and subtitles are in German; the English translation of every line is in the script below. The subtitles are burned in because many people watch without sound. The voice is an AI voice from Google, the same technology behind NotebookLM's Audio Overviews.
What you see: 5 teddy clips from existing material, 5 explainer cards in the Teddynews style, subtitles and a closing frame with the address. What's missing: the teddy doesn't move its lips. That would need an avatar model that supports animal characters.
Portrait format
On social media, the 9:16 portrait format is what counts. The version below has the same text, voice and timing as the video above — just laid out anew: clip at the top, message below, subtitles above the bottom edge. It comes from the same source; a second render is enough. Extra cost: none, because the voice tracks already exist.
Why the edges stay clear: Instagram and TikTok place their own controls over the picture — the header at the top, caption and buttons at the bottom, the icon column on the right. That is why every text in this version sits between 210 and 1500 pixels in height, inside the safe zone. The subtitles sit above that so the post caption doesn't cover them.
Storyboard
Each section has a picture, a sentence and an on-screen caption. Teddy clips and cards alternate, which gives the eye something to hold on to. Each card shows a header, a message and progress, as in good explainer videos. The stills below come from the finished video; the lines are translated from German.
00:00 · Teddy at the desk
Your mascot can do more.
Voice (translated): Your mascot can do more than sit on a shelf.
00:05 · Card 01: The situation
Many mascots stay invisible.
Voice (translated): Many mascots stay invisible: no video, no page of their own, hardly any enquiries.
00:13 · Production in our own factory
From Austria, since 1998.
Voice (translated): Teddynews has been making mascots since 1998. Over 2 million pieces, for more than 3,000 companies.
00:18 · Card 02: Figures
1998 · 2 million+ mascots · 3,000+ business customers
No voice, caption only
00:22 · Finished teddies at the factory
Sewn, not stamped.
Voice (translated): Made in our own factory, tested to CE and EN 71.
00:29 · Teddy at the laptop
From the shelf into the feed.
Voice (translated): Your mascot gets videos, its own page and appearances on social media.
00:35 · Card 03: Overview
Video · Page · App
No voice, caption only
00:38 · Card 04: without an agency
Get seen — without an agency.
Voice (translated): Get seen without an agency: a page in 60 seconds, videos from your material and, if you like, your own app.
00:48 · Card 05: Process
Sample · Production · Content
Voice (translated): Here's how it works: free sample, production, finished content for your channels.
00:55 · Teddy says thank you
See first, then decide.
Voice (translated): The sample is free and handmade. You decide only after that.Script
The text follows a clear order: first the loss, then the proof, then the offer. It ends with risk reversal. Every figure comes from teddynews.at or lindner-stofftiere.at. No sentence is longer than 25 words, which keeps the pitch short and verifiable. The video is spoken in German; the table shows English translations.
| Time | Picture | Voice | Caption |
|---|---|---|---|
| 00:00 | Teddy at the desk | Your mascot can do more than sit on a shelf. | Your mascot can do more. |
| 00:05 | Card 01: The situation | Many mascots stay invisible: no video, no page of their own, hardly any enquiries. | Many mascots stay invisible. |
| 00:13 | Production in our own factory | Teddynews has been making mascots since 1998. Over 2 million pieces, for more than 3,000 companies. | From Austria, since 1998. |
| 00:18 | Card 02: Figures | — | 1998 · 2 million+ mascots · 3,000+ business customers |
| 00:22 | Finished teddies at the factory | Made in our own factory, tested to CE and EN 71. | Sewn, not stamped. |
| 00:29 | Teddy at the laptop | Your mascot gets videos, its own page and appearances on social media. | From the shelf into the feed. |
| 00:35 | Card 03: Overview | — | Video · Page · App |
| 00:38 | Card 04: without an agency | Get seen without an agency: a page in 60 seconds, videos from your material and, if you like, your own app. | Get seen — without an agency. |
| 00:48 | Card 05: Process | Here's how it works: free sample, production, finished content for your channels. | Sample · Production · Content |
| 00:55 | Teddy says thank you | The sample is free and handmade. You decide only after that. | See first, then decide. |
Guide
The video was made without an agency and without a camera. All you need is existing footage and a script, plus a text-to-speech service and 2 free tools. Every step runs on a laptop; only the voice comes from the internet. The whole process takes a few hours.
Measure the existing clips and pick the sharpest. Crop out watermarks, remove the sound, scale to the final frame size.
Result: 5 clips, silent, cropped to fit.
8 short sections, each statement with proof. Keep sentences short, numbers as digits, no advertising clichés.
Result: 8 sentences, about 120 words in total.
Each section is voiced separately, so any part can be swapped later without regenerating everything.
Model: gemini-3.1-flash-tts-preview
Voice: Puck · Language: German
Cost: 25 tokens per second, $20 per 1M tokens
Result: 8 voice files, 60 s in total. Cost: $0.03.
Trim silence at the edges, even out the loudness, raise the tempo slightly. That keeps the pitch under 65 seconds.
ffmpeg -i vo.wav -af "silenceremove=…,loudnorm=I=-16:TP=-1.5,atempo=1.07" final.wav
Result: 60.1 s of speech instead of 69.8 s.
Pictures, cards, subtitles and audio tracks sit on a timeline in HTML. Each element gets a start time and a duration.
npx hyperframes check # checks timing, contrast, media
npx hyperframes render -q high -f 30
Result: 64 s in 1920 × 1080, rendered in 37 s. Cost: €0.
Scale down to 1280 × 720, add the audio, make the file start quickly. Plus a poster image for the page.
ffmpeg -i raw.mp4 -vf scale=1280:720 -c:v libx264 -crf 23 -c:a aac -movflags +faststart teddy-pitch.mp4
Result: 5.8 MB instead of 34 MB.
Costs
Only the voice was paid for. It cost 0.03 US dollars for 60 seconds of speech. Editing, cards, subtitles and export run on free tools, all on our own computer. The footage already existed, so a second video again costs only cents.
| Item | Cost | Status |
|---|---|---|
| Text-to-speech, 60 s | $0.03 | calculated |
| Free tier of the text-to-speech service | $0 | checked |
| Editing, cards, subtitles | €0 | checked |
| Export and compression | €0 | checked |
| Footage | existing | checked |
Text-to-speech price: $20 per 1 million tokens, 25 tokens per second of audio, according to Google's price list. German is supported according to Google's documentation.
Questions
Common lip-sync models are designed for human faces. Kling AI explicitly excludes animal characters. A talking teddy needs an avatar model that supports animal characters, and a paid plan.
Music needs a cleared licence and, in a 64-second pitch, distracts from the text. Voice and subtitles carry the message on their own.
Yes. Each section has its own voice file. A new sentence costs a few cents, then the video is exported again. Pictures and cards stay untouched.
Yes. The 9:16 portrait version is further up. It uses the same script and audio track; only the layout is reset for Reels, Shorts and TikTok.
No. Voice and subtitles are in German. With the same method, an English voice track would again cost only cents; the subtitles and cards would need translating and rendering again.
Sample, production and finished content come from one source. The sample is handmade for you and costs nothing.
Free demos, no credit card.