AI Video Avatars in Training: When They Work and When They Fall Flat in 2026
- Odin Training
- Jul 3
- 4 min read
Training video used to mean a camera, a presenter, a quiet room, and a long edit. In 2026, a growing number of training developers are producing finished video by typing a script and choosing a digital presenter. AI video avatars have moved from novelty to a standard option in the L&D toolkit, and the platforms behind them are now valued in the billions.
Speed is the obvious draw. The harder question for experienced training developers is where avatar video actually improves learning, where it simply saves time, and where it quietly works against the outcome you care about. This post looks at the current adoption data, what the research says about learning from avatars, and a practical way to decide when to use them.
For context, here is a sample AI avatar video that I created in Synthesia to test a new feature of customizing the clothes and location of an avatar. View here.
Where AI Avatar Video Adoption Stands in 2026
Adoption is no longer marginal. In Synthesia's 2026 AI in Learning and Development report, 57 percent of L&D teams said they are actively using AI in their programs, with another 30 percent running early pilots. Video creation specifically is now used by 52 percent of teams, behind voice generation at 63 percent and content or quiz drafting at 60 percent.
The scale of the platforms reflects that demand. Synthesia alone reports more than 60,000 business customers, including over 90 percent of the Fortune 100, with more than 240 avatars across 160 plus languages. The single most cited value area in its report was localization, named by 54 percent of respondents, which points to where avatar video is strongest: producing the same content in many languages without re-shooting anything.
The same teams are clear about the blockers. Security was the most common concern at 58 percent, followed by accuracy at 52 percent, integration at 46 percent, and legal restrictions at 41 percent. Those concerns carry more weight in regulated and high-risk environments than in general corporate onboarding.
What the Research Says About Learning From Avatars
The marketing claims are strong, so it is worth separating them from the evidence. A UCL study of 500 adult learners found no significant difference in engagement or retention between an AI avatar video and a human instructor video. Learners completed the avatar version roughly 20 percent faster. For straightforward informational content, a well-made avatar video can match a human presenter and take less of the learner's time.
The picture changes with authenticity. Analysis of on-screen presenters has found a retention lift of up to 38 percent from a visible human face, but that lift depends on the face being a real, recognized person rather than a generic stock avatar. Distraction also rises when an avatar fills the screen and learners start noticing robotic facial movement. Avatar video helps most when the presenter is a supporting element, not the entire experience.
Researchers reviewing AI-generated instructional video also flag structural limits: passive information processing, a weaker sense of the overall structure of a lesson, and fixed pacing. The technology produces fluent delivery, but it does not supply the pedagogical scaffolding, retrieval practice, or decision points that drive durable learning. That work still belongs to the designer.
Where AI Avatars Earn Their Place
Several use cases are a good fit. Localization is the clearest one: a single script can be delivered in dozens of languages with consistent quality, which is otherwise expensive and slow. Microlearning is another, since short explainer clips can be produced and updated quickly when a policy or procedure changes. Avatar video also works well for content that dates fast, because re-recording a section means editing text rather than rebooking a studio.
There is a specific high-value case for this audience. Some medical and professional programs now use AI-generated patient or witness avatars so learners can practice difficult conversations, such as a diagnostic interview or an investigative interview, without scheduling live role-players. Used this way, the avatar is a practice partner inside a scenario, which is far more valuable than a talking head reading a policy.
Where They Fall Flat
Avatar video is a poor choice when authenticity is part of the message. Leadership communications, values content, and anything where learners need to trust the person speaking are weakened by a synthetic presenter. The same applies to high-stakes safety and use-of-force content, where credibility and the visible judgment of a real expert carry weight that a generated face cannot replicate.
There is also a growing compliance dimension. Under the EU AI Act, transparency obligations for AI-generated content take effect on August 2, 2026, which means clearly disclosing synthetic media to viewers. Content provenance standards such as C2PA are emerging to label AI-generated media. For training developers in law enforcement and other regulated fields, undisclosed synthetic presenters create avoidable risk, and disclosure is becoming the baseline expectation rather than a courtesy.
Finally, avatar video does not fix weak instructional design. A generated presenter narrating dense bullet points is still passive click-through learning, just produced faster. If the underlying module relies on the learner doing something, an avatar reading at them is the wrong format regardless of how polished it looks.
A Practical Way to Decide
Before generating anything, ask what job the video is doing. If the goal is to deliver consistent information at scale, across languages, or to update content frequently, avatar video is a strong, cost-effective option. If the goal is to build trust, model expert judgment, or drive practice and decision-making, either keep a real human on camera or move to an interactive format instead of a passive video.
A reasonable default is to use avatar video for the informational layer of a program and reserve real presenters and scenario-based practice for the parts that change behavior. Disclose synthetic media every time, keep a designer in control of structure and assessment, and treat the avatar as one production tool among several rather than a replacement for instructional design.
Sources
A Note on AI Use
This post was researched and drafted with AI assistance, then reviewed and edited for accuracy and voice. All practical recommendations reflect my own instructional design experience.



Comments