图学学报 ›› 2026, Vol. 47 ›› Issue (4): 757-765.DOI: 10.11996/JG.j.2095-302X.2026040757
收稿日期:2026-01-12
接受日期:2026-04-13
出版日期:2026-08-31
发布日期:2026-08-31
通讯作者:刘永进,E-mail:liuyongjin@tsinghua.edu.cn
ZHOU Yu, LV Tian, LI Ming, LIU Yongjin(
)
Received:2026-01-12
Accepted:2026-04-13
Published:2026-08-31
Online:2026-08-31
Contact:
LIU Yongjin,E-mail:liuyongjin@tsinghua.edu.cn摘要:
语音驱动的三维人脸动画生成是计算机视觉与计算机图形学中的重要研究方向,在虚拟数字人、实时交互系统以及沉浸式媒体等应用场景中具有广泛需求。该任务旨在根据输入语音信号生成与语音内容、节奏和发音特征高度一致的三维人脸运动,但现有方法往往依赖复杂的生成过程或多步采样机制,导致推理延迟较高,难以满足实时流式应用的性能要求。为解决上述问题,提出了一种面向实时生成的语音驱动三维人脸动画方法,将人脸运动建模为基于离散表示的自回归序列生成过程。首先利用VQ-VAE将连续的人脸运动参数映射到紧凑的离散token空间,从而降低序列建模难度,并在此基础上采用类语言模型的自回归Transformer,在语音条件约束下对离散人脸运动序列进行逐帧预测,实现低延迟、可流式的人脸动画生成。在人脸表示方面,统一采用FLAME2023参数化模型,并基于VHAP tracker对音视频数据进行重新拟合,以获得高质量且时间一致的人脸运动参数序列。实验结果表明,该方法在唇部同步、面部表情动态以及头部运动一致性等方面均取得了具有竞争力的表现,同时在推理效率和实时性能上显著优于现有生成方法,验证了其在实时交互场景中的实用性。相关研究为构建高效、可实时交互的三维虚拟人系统提供了一种可行且可扩展的技术路径,并为后续在离散自回归框架下进一步引入风格、情绪状态等高层语义控制因素奠定了基础。
中图分类号:
周禹, 吕天, 李明, 刘永进. 基于离散自回归序列建模的实时语音驱动三维说话人动画生成[J]. 图学学报, 2026, 47(4): 757-765.
ZHOU Yu, LV Tian, LI Ming, LIU Yongjin. Real-time speech-driven 3D talking face animation via discrete autoregressive sequence modeling[J]. Journal of Graphics, 2026, 47(4): 757-765.
| 方法 | LVE/mm↓ | FDD/×10-5m↓ | FPS↑ |
|---|---|---|---|
| FaceFormer[ | 10.07 | 16.85 | 2.07 |
| SelfTalk[ | 12.03 | 15.93 | 23.91 |
| ARTalk[ | 10.31 | 11.01 | 18.31 |
| DiffPoseTalk[ | 8.83 | 10.46 | 0.14 |
| 本文方法 | 10.12 | 10.97 | 27.63 |
表1 不同方法的生成质量评估
Table 1 Evaluation of generation quality across different methods
| 方法 | LVE/mm↓ | FDD/×10-5m↓ | FPS↑ |
|---|---|---|---|
| FaceFormer[ | 10.07 | 16.85 | 2.07 |
| SelfTalk[ | 12.03 | 15.93 | 23.91 |
| ARTalk[ | 10.31 | 11.01 | 18.31 |
| DiffPoseTalk[ | 8.83 | 10.46 | 0.14 |
| 本文方法 | 10.12 | 10.97 | 27.63 |
| 方法 | LVE/mm↓ | FDD/×10-5 m↓ |
|---|---|---|
| 本文方法(使用FLAME2020) | 11.04 | 11.41 |
| 本文方法(去除平滑损失) | 10.21 | 11.09 |
| 本文方法 | 10.12 | 10.97 |
表2 消融实验结果
Table 2 Ablation study
| 方法 | LVE/mm↓ | FDD/×10-5 m↓ |
|---|---|---|
| 本文方法(使用FLAME2020) | 11.04 | 11.41 |
| 本文方法(去除平滑损失) | 10.21 | 11.09 |
| 本文方法 | 10.12 | 10.97 |
| [1] | CHU X G, LI Y, ZENG A L, et al. GPAvatar: generalizable and precise head avatar from image(s)[EB/OL]. [2024-01-18]. http://arxiv.org/abs/2401.10215. |
| [2] | MA S J, WENG Y L, SHAO T J, et al. 3D Gaussian blendshapes for head avatar animation[C]// 2024 ACM SIGGRAPH Conference Papers. New York: ACM, 2024: 1-10. |
| [3] | LI T Y, BOLKART T, BLACK M J, et al. Learning a model of facial shape and expression from 4D scans[J]. ACM Transactions on Graphics, 2017, 36(6): 194. |
| [4] | FAN Y R, LIN Z J, SAITO J, et al. FaceFormer: speech-driven 3D facial animation with transformers[C]// 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2022: 18770-18780. |
| [5] | VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need[J]. Advances in Neural Information Processing Systems, 2017, 30. |
| [6] | PENG Z Q, LUO Y H, SHI Y, et al. SelfTalk: a self-supervised commutative training diagram to comprehend 3D talking faces[C]// The 31st ACM International Conference on Multimedia. New York: ACM, 2023: 5292-5301. |
| [7] | HO J, JAIN A, ABBEEL P. Denoising diffusion probabilistic models[J]. Advances in Neural Information Processing Systems, 2020, 33: 6840-6851. |
| [8] | SONG J M, MENG C L, ERMON S. Denoising diffusion implicit models[EB/OL]. [2020-10-06]. http://arxiv.org/abs/2010.02502. |
| [9] | SUN Z Y, LV T, YE S, et al. DiffPoseTalk: speech-driven stylistic 3D facial animation and head pose generation via diffusion models[J]. ACM Transactions on Graphics, 2024, 43(4): 1-9. |
| [10] | VAN DEN OORD A, VINYALS O, KAVUKCUOGLU K. Neural discrete representation learning[J]. Advances in Neural Information Processing Systems, 2017, 30. |
| [11] | CHU X G, GOSWAMI N, CUI Z T, et al. ARTalk: speech- driven 3D head animation via autoregressive model[C]// 2025 ACM SIGGRAPH Conference Papers. New York: ACM, 2025: 1-9. |
| [12] | MASSARO D W, COHEN M M, TABAIN M, et al. Animated speech: research progress and applications[M]//BAILLY G, PERRIER P, VATIKIOTIS-BATESON E. Audiovisual Speech Processing. Cambridge: Cambridge University Press, 2012: 309-345. |
| [13] | EDWARDS P, LANDRETH C, FIUME E, et al. JALI: an animator-centric viseme model for expressive lip synchronization[J]. ACM Transactions on Graphics, 2016, 35(4): 1-11. |
| [14] | BAEVSKI A, ZHOU Y H, MOHAMED A, et al. wav2vec 2.0: a framework for self-supervised learning of speech representations[J]. Advances in Neural Information Processing Systems, 2020, 33: 12449-12460. |
| [15] |
HSU W N, BOLTE B, TSAI Y H H, et al. HuBERT: self-supervised speech representation learning by masked prediction of hidden units[J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021, 29: 3451-3460.
DOI URL |
| [16] | PENG Z Q, WU H Y, SONG Z B, et al. EmoTalk: speech- driven emotional disentanglement for 3D face animation[C]// 2023 IEEE/CVF International Conference on Computer Vision. New York: IEEE Press, 2023: 20687-20697. |
| [17] |
ZHANG C X, NI S F, FAN Z P, et al. 3D talking face with personalized pose dynamics[J]. IEEE Transactions on Visualization and Computer Graphics, 2021, 29(2): 1438-1449.
DOI URL |
| [18] | CUDEIRO D, BOLKART T, LAIDLAW C, et al. Capture, learning, and synthesis of 3D speaking styles[C]// 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2019: 10101-10111. |
| [19] | HAQUE K I, YUMAK Z. FaceXHuBERT: text-less speech-driven EXpressive 3D facial animation synthesis using self-supervised speech representation learning[C]// The 25th International Conference on Multimodal Interaction. New York: ACM, 2023: 282-291. |
| [20] | XING J B, XIA M H, ZHANG Y C, et al. CodeTalker: speech-driven 3D facial animation with discrete motion prior[C]// 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2023: 12780-12790. |
| [21] | THAMBIRAJA B, HABIBIE I, ALIAKBARIAN S, et al. Imitator: personalized speech-driven 3D facial animation[C]// 2023 IEEE/CVF International Conference on Computer Vision. New York: IEEE Press, 2023: 20621-20631. |
| [22] | STAN S, HAQUE K I, YUMAK Z. FaceDiffuser: speech- driven 3D facial animation synthesis using diffusion[C]// The 16th ACM SIGGRAPH Conference on Motion, Interaction and Games. New York: ACM, 2023: 1-11. |
| [23] | QIAN S H, KIRSCHSTEIN T, SCHONEVELD L, et al. GaussianAvatars: photorealistic head avatars with rigged 3D Gaussians[C]// 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2024: 20299-20309. |
| [24] | DANĚČEK R, CHHATRE K, TRIPATHI S, et al. Emotional speech-driven animation with content-emotion disentanglement[C]// 2023 SIGGRAPH Asia Conference Papers. New York: ACM, 2023: 1-13. |
| [25] | ZHANG Z M, LI L C, DING Y, et al. Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset[C]// 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2021: 3661-3670. |
| [1] | 田硕, 黄炎, 齐家望, 石超君, 戚银城. 流场置信度引导与各向异性约束的无监督图像拼接方法[J]. 图学学报, 2026, 47(4): 683-694. |
| [2] | 欧阳泽洪, 沈旭昆, 任曦, 胡勇, 黄勇. 基于共视引导的大规模场景运动恢复结构[J]. 图学学报, 2026, 47(4): 695-703. |
| [3] | 王紫威, 王录涛, 李桉同, 沈艳. 单目深度模糊感知估计的少视图三维高斯重建[J]. 图学学报, 2026, 47(4): 704-713. |
| [4] | 许航, 谢雪光, 夏清, 高阳, 禹鹏, 胡珈皓. 基于语义感知和混合物质点法的高斯动态重建[J]. 图学学报, 2026, 47(4): 714-725. |
| [5] | 李煜华, 姜杉, 杨志永, 王禹泽, 周泽洋. 物理增强-深度协同的自由式三维超声重建[J]. 图学学报, 2026, 47(4): 726-735. |
| [6] | 赵啦啦, 杨亦卓, 段晨龙, 郭辰昊, 王清龙, 王宏都. 一种融合频率扰动的KL展开3D颗粒随机建模方法[J]. 图学学报, 2026, 47(4): 736-745. |
| [7] | 刘渠, 陈斌, 黄元正. QC-ORF:基于弱提示的三维高斯条件查询对象响应场构建方法[J]. 图学学报, 2026, 47(4): 746-756. |
| [8] | 唐晓腾, 姚君, 胡鹤凡, 邵将, 束云峰. 虚拟现实环境下多类型眼控选择任务的用户意图识别模型研究[J]. 图学学报, 2026, 47(4): 766-775. |
| [9] | 温瑞祺, 吕健, 宋定安, 苏乐, 梁智斌. 具身视域下康复训练虚拟现实系统设计与评估[J]. 图学学报, 2026, 47(4): 776-787. |
| [10] | 陆相江, 姜豪, 王爱增, 宁涛. 一种G2连续插值过渡曲面建模方法[J]. 图学学报, 2026, 47(4): 788-800. |
| [11] | 陈国军, 孔赟艺, 陈家乐, 宋双双. 基于计算着色器的并行约束Delaunay三角剖分算法[J]. 图学学报, 2026, 47(4): 801-811. |
| [12] | 张智博, 郑联语. 基于点云数据的管路几何特征统一自动提取方法及应用[J]. 图学学报, 2026, 47(4): 812-819. |
| [13] | 李凯, 刘绍华, 何梓豪, 刘康凡, 周思龙. 大规模线束拓扑并行归约与上下文敏感增量更新方法[J]. 图学学报, 2026, 47(4): 820-833. |
| [14] | 尹思琪, 刘利刚. 基于层次概率路线图的多目标点路径规划[J]. 图学学报, 2026, 47(4): 834-843. |
| [15] | 王雨涛, 杨超, 况立群, 杨晓文, 韩燮, 焦世超. 基于层次化对齐的三维模型零样本草图检索[J]. 图学学报, 2026, 47(4): 844-853. |
| 阅读次数 | ||||||
|
全文 |
|
|||||
|
摘要 |
|
|||||