Journal of Graphics ›› 2026, Vol. 47 ›› Issue (4): 757-765.DOI: 10.11996/JG.j.2095-302X.2026040757
• Computer Graphics and Virtual Reality • Previous Articles Next Articles
ZHOU Yu, LV Tian, LI Ming, LIU Yongjin(
)
Received:2026-01-12
Accepted:2026-04-13
Online:2026-08-31
Published:2026-08-31
Contact:
LIU Yongjin
CLC Number:
ZHOU Yu, LV Tian, LI Ming, LIU Yongjin. Real-time speech-driven 3D talking face animation via discrete autoregressive sequence modeling[J]. Journal of Graphics, 2026, 47(4): 757-765.
Add to citation manager EndNote|Ris|BibTeX
URL: http://www.txxb.com.cn/EN/10.11996/JG.j.2095-302X.2026040757
| 方法 | LVE/mm↓ | FDD/×10-5m↓ | FPS↑ |
|---|---|---|---|
| FaceFormer[ | 10.07 | 16.85 | 2.07 |
| SelfTalk[ | 12.03 | 15.93 | 23.91 |
| ARTalk[ | 10.31 | 11.01 | 18.31 |
| DiffPoseTalk[ | 8.83 | 10.46 | 0.14 |
| 本文方法 | 10.12 | 10.97 | 27.63 |
Table 1 Evaluation of generation quality across different methods
| 方法 | LVE/mm↓ | FDD/×10-5m↓ | FPS↑ |
|---|---|---|---|
| FaceFormer[ | 10.07 | 16.85 | 2.07 |
| SelfTalk[ | 12.03 | 15.93 | 23.91 |
| ARTalk[ | 10.31 | 11.01 | 18.31 |
| DiffPoseTalk[ | 8.83 | 10.46 | 0.14 |
| 本文方法 | 10.12 | 10.97 | 27.63 |
| 方法 | LVE/mm↓ | FDD/×10-5 m↓ |
|---|---|---|
| 本文方法(使用FLAME2020) | 11.04 | 11.41 |
| 本文方法(去除平滑损失) | 10.21 | 11.09 |
| 本文方法 | 10.12 | 10.97 |
Table 2 Ablation study
| 方法 | LVE/mm↓ | FDD/×10-5 m↓ |
|---|---|---|
| 本文方法(使用FLAME2020) | 11.04 | 11.41 |
| 本文方法(去除平滑损失) | 10.21 | 11.09 |
| 本文方法 | 10.12 | 10.97 |
| [1] | CHU X G, LI Y, ZENG A L, et al. GPAvatar: generalizable and precise head avatar from image(s)[EB/OL]. [2024-01-18]. http://arxiv.org/abs/2401.10215. |
| [2] | MA S J, WENG Y L, SHAO T J, et al. 3D Gaussian blendshapes for head avatar animation[C]// 2024 ACM SIGGRAPH Conference Papers. New York: ACM, 2024: 1-10. |
| [3] | LI T Y, BOLKART T, BLACK M J, et al. Learning a model of facial shape and expression from 4D scans[J]. ACM Transactions on Graphics, 2017, 36(6): 194. |
| [4] | FAN Y R, LIN Z J, SAITO J, et al. FaceFormer: speech-driven 3D facial animation with transformers[C]// 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2022: 18770-18780. |
| [5] | VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need[J]. Advances in Neural Information Processing Systems, 2017, 30. |
| [6] | PENG Z Q, LUO Y H, SHI Y, et al. SelfTalk: a self-supervised commutative training diagram to comprehend 3D talking faces[C]// The 31st ACM International Conference on Multimedia. New York: ACM, 2023: 5292-5301. |
| [7] | HO J, JAIN A, ABBEEL P. Denoising diffusion probabilistic models[J]. Advances in Neural Information Processing Systems, 2020, 33: 6840-6851. |
| [8] | SONG J M, MENG C L, ERMON S. Denoising diffusion implicit models[EB/OL]. [2020-10-06]. http://arxiv.org/abs/2010.02502. |
| [9] | SUN Z Y, LV T, YE S, et al. DiffPoseTalk: speech-driven stylistic 3D facial animation and head pose generation via diffusion models[J]. ACM Transactions on Graphics, 2024, 43(4): 1-9. |
| [10] | VAN DEN OORD A, VINYALS O, KAVUKCUOGLU K. Neural discrete representation learning[J]. Advances in Neural Information Processing Systems, 2017, 30. |
| [11] | CHU X G, GOSWAMI N, CUI Z T, et al. ARTalk: speech- driven 3D head animation via autoregressive model[C]// 2025 ACM SIGGRAPH Conference Papers. New York: ACM, 2025: 1-9. |
| [12] | MASSARO D W, COHEN M M, TABAIN M, et al. Animated speech: research progress and applications[M]//BAILLY G, PERRIER P, VATIKIOTIS-BATESON E. Audiovisual Speech Processing. Cambridge: Cambridge University Press, 2012: 309-345. |
| [13] | EDWARDS P, LANDRETH C, FIUME E, et al. JALI: an animator-centric viseme model for expressive lip synchronization[J]. ACM Transactions on Graphics, 2016, 35(4): 1-11. |
| [14] | BAEVSKI A, ZHOU Y H, MOHAMED A, et al. wav2vec 2.0: a framework for self-supervised learning of speech representations[J]. Advances in Neural Information Processing Systems, 2020, 33: 12449-12460. |
| [15] |
HSU W N, BOLTE B, TSAI Y H H, et al. HuBERT: self-supervised speech representation learning by masked prediction of hidden units[J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021, 29: 3451-3460.
DOI URL |
| [16] | PENG Z Q, WU H Y, SONG Z B, et al. EmoTalk: speech- driven emotional disentanglement for 3D face animation[C]// 2023 IEEE/CVF International Conference on Computer Vision. New York: IEEE Press, 2023: 20687-20697. |
| [17] |
ZHANG C X, NI S F, FAN Z P, et al. 3D talking face with personalized pose dynamics[J]. IEEE Transactions on Visualization and Computer Graphics, 2021, 29(2): 1438-1449.
DOI URL |
| [18] | CUDEIRO D, BOLKART T, LAIDLAW C, et al. Capture, learning, and synthesis of 3D speaking styles[C]// 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2019: 10101-10111. |
| [19] | HAQUE K I, YUMAK Z. FaceXHuBERT: text-less speech-driven EXpressive 3D facial animation synthesis using self-supervised speech representation learning[C]// The 25th International Conference on Multimodal Interaction. New York: ACM, 2023: 282-291. |
| [20] | XING J B, XIA M H, ZHANG Y C, et al. CodeTalker: speech-driven 3D facial animation with discrete motion prior[C]// 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2023: 12780-12790. |
| [21] | THAMBIRAJA B, HABIBIE I, ALIAKBARIAN S, et al. Imitator: personalized speech-driven 3D facial animation[C]// 2023 IEEE/CVF International Conference on Computer Vision. New York: IEEE Press, 2023: 20621-20631. |
| [22] | STAN S, HAQUE K I, YUMAK Z. FaceDiffuser: speech- driven 3D facial animation synthesis using diffusion[C]// The 16th ACM SIGGRAPH Conference on Motion, Interaction and Games. New York: ACM, 2023: 1-11. |
| [23] | QIAN S H, KIRSCHSTEIN T, SCHONEVELD L, et al. GaussianAvatars: photorealistic head avatars with rigged 3D Gaussians[C]// 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2024: 20299-20309. |
| [24] | DANĚČEK R, CHHATRE K, TRIPATHI S, et al. Emotional speech-driven animation with content-emotion disentanglement[C]// 2023 SIGGRAPH Asia Conference Papers. New York: ACM, 2023: 1-13. |
| [25] | ZHANG Z M, LI L C, DING Y, et al. Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset[C]// 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2021: 3661-3670. |
| [1] | TIAN Shuo, HUANG Yan, QI Jiawang, SHI Chaojun, QI Yincheng. Unsupervised image stitching method guided by flow field confidence and anisotropic constraints [J]. Journal of Graphics, 2026, 47(4): 683-694. |
| [2] | OUYANG Zehong, SHEN Xukun, REN Xi, HU Yong, HUANG Yong. Covisibility-based large-scale structure from motion [J]. Journal of Graphics, 2026, 47(4): 695-703. |
| [3] | WANG Ziwei, WANG Lutao, LI Antong, SHEN Yan. Few-shot 3D Gaussian splatting based on monocular depth ambiguity-aware estimation [J]. Journal of Graphics, 2026, 47(4): 704-713. |
| [4] | XU Hang, XIE Xueguang, XIA Qing, GAO Yang, YU Peng, HU Jiahao. Gaussian dynamic reconstruction based on semantic perception and hybrid material point method [J]. Journal of Graphics, 2026, 47(4): 714-725. |
| [5] | LI Yuhua, JIANG Shan, YANG Zhiyong, WANG Yuze, ZHOU Zeyang. Physically-augmented and depth synergized freehand 3D ultrasound reconstruction [J]. Journal of Graphics, 2026, 47(4): 726-735. |
| [6] | ZHAO Lala, YANG Yizhuo, DUAN Chenlong, GUO Chenhao, WANG Qinglong, WANG Hongdu. A 3D random particle modeling method integrating KL expansion and frequency perturbation [J]. Journal of Graphics, 2026, 47(4): 736-745. |
| [7] | LIU Qu, CHEN Bin, HUANG Yuanzheng. QC-ORF: constructing query-conditioned object response fields in 3D Gaussians via weak prompts [J]. Journal of Graphics, 2026, 47(4): 746-756. |
| [8] | TANG Xiaoteng, YAO Jun, HU Hefan, SHAO Jiang, SHU Yunfeng. User intention recognition for multi-type gaze-based target selection tasks in virtual reality [J]. Journal of Graphics, 2026, 47(4): 766-775. |
| [9] | WEN Ruiqi, LV Jian, SONG Dingan, SU Le, LIANG Zhibin. Embodied virtual reality system design and evaluation for rehabilitation training [J]. Journal of Graphics, 2026, 47(4): 776-787. |
| [10] | LU Xiangjiang, JIANG Hao, WANG Aizeng, NING Tao. A modeling method for G2-continuous interpolating blending surface [J]. Journal of Graphics, 2026, 47(4): 788-800. |
| [11] | CHEN Guojun, KONG Yunyi, CHEN Jiale, SONG Shuangshuang. Parallel constrained Delaunay triangulation algorithm based on compute shaders [J]. Journal of Graphics, 2026, 47(4): 801-811. |
| [12] | ZHANG Zhibo, ZHENG Lianyu. Unified automatic extraction method for pipeline geometric features based on point cloud data and its application [J]. Journal of Graphics, 2026, 47(4): 812-819. |
| [13] | LI Kai, LIU Shaohua, HE Zihao, LIU Kangfan, ZHOU Silong. Parallel reduction and context-sensitive incremental update method for large-scale wiring harness topology [J]. Journal of Graphics, 2026, 47(4): 820-833. |
| [14] | YIN Siqi, LIU Ligang. Multi-goal path planning based on hierarchical probabilistic roadmaps [J]. Journal of Graphics, 2026, 47(4): 834-843. |
| [15] | WANG Yutao, YANG Chao, KUANG Liqun, YANG Xiaowen, HAN Xie, JIAO Shichao. Hierarchical alignment for zero-shot sketch-based 3D shape retrieval [J]. Journal of Graphics, 2026, 47(4): 844-853. |
| Viewed | ||||||
|
Full text |
|
|||||
|
Abstract |
|
|||||