How to build a consistent AI creator cast, and why one-off avatars fail
Consistency is the whole game for AI UGC at scale. How to pick portraits that hold up, keep identity locked across takes, cast three to five characters by energy and audience, diagnose drift, and run a cast like a brand asset instead of a new face every time.
Workflow6 min read
Updated
Key takeaways
- A creative test only measures the hook if the creator is identical across variations. Drift ruins the test.
- Start from a clean, well-lit, front-facing portrait with the face uncovered. Reference-driven models reward it.
- Cast three to five characters by energy and audience, not by looks, and reuse them.
- Treat every character like a brand asset: named, owned, permission-cleared, and stored where the next render can find it.
Why one-off avatars quietly wreck your tests
The first time a team uses AI creators, they generate a new face for every concept. It feels like variety. It is actually noise. When variation A has a warm creator and variation B has a cool one, the click-through difference is about the faces, and the team walks away believing something about hooks that the data never said.
Consistency is what makes AI creators useful for testing at all. Same face, same energy, same delivery, six different opening lines: now the only thing that changed is the thing you wanted to measure. It also compounds outside the test. A presenter that shows up in every ad becomes familiar, and familiarity is most of what a brand spokesperson is for.
There is a second, quieter cost to one-off avatars: nobody owns them. A face generated on a Tuesday for one concept has no name, no file, no brief, and no future. Six months later the team cannot reproduce the ad that worked, because the face that made it work was never saved.
Start with a portrait that will hold
Reference-driven video models take one image to lock identity and one clip to drive motion. The identity image is doing a lot of work, so choose it deliberately. Everything the model knows about the face comes from that single frame, so anything ambiguous in the frame becomes ambiguous in every render.
- Front-facing or close to it, with both eyes clearly visible.
- Even lighting. Hard shadows on one side become a different face when the head turns.
- Face uncovered: no sunglasses, no hands, no hair across the eyes, no heavy filters.
- Neutral to mild expression. The reference clip supplies the performance; a portrait mid-laugh fights it.
- Plain background and a clean crop from the chest up. Busy backgrounds leak into the scene match.
- High resolution and sharp. Upscaling a soft portrait produces a soft character.
| Portrait problem | What it does to renders | Fix |
|---|---|---|
| Side lighting | Face changes shape when the head turns | Reshoot facing a window or a soft light |
| Big smile | Mouth shapes fight the reference clip's speech | Neutral or slight smile |
| Sunglasses or hat | Model invents eyes or hairline, differently each time | Uncovered face |
| Wide shot | Too few face pixels; identity drifts | Crop to chest-up |
| Heavy filter | Skin texture reads as plastic | Unfiltered photo |
In the VibesUGC workspace a character is a single named JPEG, PNG, or WebP portrait saved to your account. Saving it does not start a render; it is a free creative step that makes the next render start with the right face preselected.
Cast by energy, not by looks
A useful cast is small. Three to five characters cover most of what a brand needs, and the difference between them should be about delivery and audience fit rather than appearance. One calm explainer. One high-energy hype presenter. One skeptical, dry reviewer type. Maybe one that mirrors the exact customer you are targeting.
| Role | Energy | Use it for |
|---|---|---|
| The explainer | Calm, direct, trustworthy | Feature walkthroughs, problem and solution bodies |
| The hype | Fast, expressive, big gestures | Pattern-interrupt hooks, launches, offers |
| The skeptic | Dry, deadpan, slightly tired | Objection handling, comparison ads, contrarian hooks |
| The mirror | Matches the target customer | Retargeting, direct-address hooks |
| The expert | Measured, specific, unhurried | Ingredient, spec, or how-it-works angles |
Then match the reference clip to the role. A high-energy motion reference on the calm explainer produces a performance that looks wrong even when identity holds perfectly, because the body language and the face were never meant to go together. Build a small reference library per role: two or three motion clips whose pacing suits each character.
Diversity in a cast is about audience fit, not a checkbox. If your customers are mostly one demographic, the mirror character should look like them. If your customers span several, so should the cast, and the test data will tell you which character each segment trusts.
Keep identity locked across takes
- Use the same saved portrait for every render of that character. Never re-generate it, and never swap in a "better" photo halfway through a test.
- Keep reference clips short and framed similarly. Wild framing changes push the model harder than it needs to be pushed.
- Review the first render against the portrait before cutting variations. If identity drifts, fix the portrait, not the prompt.
- Give characters names and use them in file names, briefs, and reporting. A cast people refer to by name gets reused.
- Keep wardrobe and setting consistent within a campaign. The face is only part of recognition.
| Symptom | Likely cause | Fix |
|---|---|---|
| Face looks different when turned | Side-lit or wide portrait | Reshoot the portrait |
| Mouth looks wrong on speech | Portrait expression fights the reference | Neutral portrait, or a reference with less extreme mouth shapes |
| Body looks stiff | Reference clip is static or over-cropped | Use a reference with natural, moderate movement |
| Skin looks plastic | Filtered or upscaled portrait | Unfiltered, native-resolution photo |
| Character ages or changes between renders | Different portraits used | One saved character, always |
Running the cast as a brand asset
A consistent cast compounds. Audiences start to recognize the presenter, the team stops relitigating what the creator should look like, and every winning body becomes a template the next hook can run on. It also makes disclosure straightforward, because a named synthetic spokesperson is easy to label honestly, and the rules for that reward exactly that clarity.
- Keep a one-page cast sheet: name, role, energy, the portrait, the two or three reference clips that suit it, and the ads it has appeared in.
- Record the rights basis for each portrait: who the person is, what they agreed to, and for how long.
- Retire characters deliberately. When a character's ads fatigue, rest it for a quarter rather than replacing it.
- Add a character only when a role is missing, not when a concept feels like it wants a new face.
The right face stays ready for the next take
Save a reusable character or pick from the creator catalog, then carry that choice straight into Copy Video.
Frequently asked questions
- How many AI characters does a brand need?
- Three to five, chosen by role and energy: an explainer, a hype presenter, a skeptic, a mirror of the target customer, and optionally an expert. More than that dilutes recognition and makes tests harder to read.
- What makes a good portrait for an AI character?
- Front-facing, evenly lit, uncovered face, neutral expression, plain background, chest-up crop, sharp and unfiltered. Most identity drift traces back to a portrait that breaks one of those rules.
- Can I use a stock photo or a photo of a real person as a character?
- Only with rights that cover advertising use of that person's likeness. A stock license may not include it, and a photo found online never does. VibesUGC asks you to attest to rights before a render.
- Why does my character look different in every render?
- Usually because different portraits were used, or the portrait is side-lit, wide, or filtered. Save one portrait as a named character and use it every time.