Our challenge: Amber Williamson, Digital Willow’s CEO, like many business leaders, has a jam-packed schedule and can find that getting in front of the camera tends to be both a time-vacuum and sometimes challenging.
Our goal was to replicate Amber’s voice and personality, not just her look, but her gestures, tone, and charisma, using AI.
Naively, we thought creating a believable AI ‘Amber-style’ avatar would be as simple as uploading a few photos and hitting “generate.”
It wasn’t.
So here are the lessons from the trenches; bloopers, wins and all the fun in-between.
Lesson One: Image-Based Avatars & the Multi-Angle Myth
Our first idea was simple enough: feed AI video software multiple photos of Amber from every angle and let it work its magic, blending them into one dynamic avatar. Spoiler alert, it did not.
Next, we tried animating a single smiling photo. The result? Think mannequin possessed by caffeine, exaggerated grins, twitchy movements, and a general air of uncanny horror. This method looked the least realistic of all.
Amber’s AI Clone, Subject 1:
Amber’s AI Clone, Subject 2 (Caution, video may cause nightmares!):
Across a variety of platforms we tried, they only allowed one asset (a single image or video) to be used as the source for an avatar. So they don’t blend multiple inputs (yet). This was just the first of many limitations we would uncover.
Lesson Two: From Script to Screen, AI’s Curious Interpretation
Next, we used a video of Amber sitting at a table, talking and gesturing naturally. We paired it with a written script, hoping the AI would catch her spark. It did not.
While the result captured basic movements, it lacked finesse:
- Gestures were robotic
- Voiceovers sounded monotone and overly synthetic
- Facial expressions were often misaligned or missing altogether
We also noticed that when there were pauses in the audio, AI filled them with strange micro-expressions and exaggerated gestures. This made the output feel unnatural. A quick lesson:
Be careful allowing too many pauses between sentences as the AI software likes to get creative with your avatars expressions, often in the weirdest ways.
Amber’s AI Clone, Subject 3:
Lesson Three: The Partial Win, Sound Syncing Without the Smooth Moves
Recording a separate voiceover and using the lip-sync function was definitely a step up, finally, it sounded like a real human and not a satnav.
But the win was short-lived. The faces were still doing their own interpretive dance, either frozen like a wax figure or flailing like they’d had one too many energy drinks. And those mouth movements? Let’s just say they didn’t quite get the memo on timing.
Amber’s AI Clone, Subject 4:
Lesson Four: When AI Gestures Go Off Script
During the editing process we explored functions that allowed us to edit gestures like pointing or smiling, with the ambition of thinking we could finesse the output. Unfortunately, we found that inserting gestures at certain time stamps during the video didn’t look natural and really disrupted the flow. We dropped use of that feature quickly.
Lesson Five: Close-Up Chaos – When AI Mouths Go Rogue
Returning to video inputs, we tested a version where Amber was seated directly in front of the camera, no angled shots. It was better, but her mouth movements still looked unnatural, and it broke the illusion.
When it comes to realism in AI video, it all hinges on the mouth, and that’s exactly where AI still bites the dust.
Pro tip for your source video: Think of it like directing a scene, your avatar is only as good as your source performance. The tone, energy, and mood need to line up perfectly. If your actor looks drained or distracted in the source footage, that’s exactly what the AI will replicate. But give it a spark, even the smallest smile or glimmer of intent, and the whole clone comes to life.
Amber’s AI Clone, Subject 5:
Lesson Six: Finally, A Breakthrough: Up Close and Realistic, Subtlety is Key
Our best results came from a clear, close-up, front-facing video filmed with Amber positioned closely to the camera. The frame included only her head, neck, and upper body, no hand movements, just subtle facial expressions.
This input produced the most realistic avatar to date. There were still minor glitches, but we simply edited them out during post-production for a smoother final product.
Amber’s AI Clone, Subject 6:
Key Takeaways from Our AI Video Journey
- Image-based avatars lack naturalism, use video instead. Don’t upload multiple images expecting a “blended” avatar.
- Use a neutral-to-positive mood in the source video, remember your avatar will mirror your subject.
- Avoid pauses and paragraph breaks in your audio.
- Experiment with recording audio as a separate file and syncing it to your AI video. This sometimes has better results than uploading a complete audio and visual asset all together.
- Pick a source video with minimal movement.
- Editing in gestures during post production rarely improved final output.
Your Next Step Towards Effortless Video Creation
If you’ve been putting off video thought leadership because it feels uncomfortable or takes up too much time then talk to us about this solution. Fill out the form above or head over to our Say Hello page to get in touch. Now’s the time to get ahead, before your competitors do.