Welcome to Journal of Graphics

Journal of Graphics ›› 2026, Vol. 47 ›› Issue (4): 757-765.DOI: 10.11996/JG.j.2095-302X.2026040757

• Computer Graphics and Virtual Reality • Previous Articles     Next Articles

Real-time speech-driven 3D talking face animation via discrete autoregressive sequence modeling

ZHOU Yu, LV Tian, LI Ming, LIU Yongjin()   

  1. Department of Computer Science and Technology, Tsinghua University, Beijing 100084, China
  • Received:2026-01-12 Accepted:2026-04-13 Online:2026-08-31 Published:2026-08-31
  • Contact: LIU Yongjin

Abstract:

Speech-driven 3D facial animation generation is an important research topic in computer vision and computer graphics, with broad demand in applications such as virtual digital humans, real-time interactive systems, and immersive media. The goal of this task is to generate 3D facial motions that are highly consistent with the content, rhythm, and articulation characteristics of the input speech signal. However, existing methods often rely on complex generation pipelines or multi-step sampling mechanisms, resulting in high inference latency and making them difficult to deploy in real-time streaming scenarios. To address these limitations, a real-time-oriented approach for speech-driven 3D facial animation was proposed, in which facial motion generation was formulated as an autoregressive sequence modeling problem based on discrete representations. Specifically, continuous facial motion parameters were first mapped into a compact discrete token space using VQ-VAE, thereby reducing the difficulty of sequence modeling. On this basis, a language-model-like autoregressive Transformer was employed to predict discrete facial motion tokens frame by frame under speech-conditioned constraints, enabling low-latency and streamable facial animation generation. For facial representation, the FLAME2023 parametric model was adopted in a unified manner, and audio-visual data were re-fitted using the VHAP tracker to obtain high-quality and temporally consistent facial motion parameter sequences. Experimental results demonstrated that this approach achieved competitive performance in lip synchronization, facial-expression dynamics, and head-motion consistency, while significantly outperforming existing generation methods in terms of inference efficiency and real-time performance, validating its practical applicability in real-time interactive scenarios. This research provides a feasible and scalable technical pathway for building efficient, real-time interactive 3D virtual human systems, and lays the groundwork for further incorporation of higher-level semantic controls such as speaking style and emotional states within a discrete autoregressive framework.

Key words: 3D talking face video generation, real-time streaming inference, discrete autoregressive generation, parametric 3D face model, virtual digital humans

CLC Number: