Welcome to Journal of Graphics

Journal of Graphics ›› 2026, Vol. 47 ›› Issue (4): 766-775.DOI: 10.11996/JG.j.2095-302X.2026040766

• Computer Graphics and Virtual Reality • Previous Articles     Next Articles

User intention recognition for multi-type gaze-based target selection tasks in virtual reality

TANG Xiaoteng1,2, YAO Jun1(), HU Hefan3, SHAO Jiang1, SHU Yunfeng2   

  1. 1 School of Architecture and Design, China University of Mining and Technology, Xuzhou Jiangsu 221000, China
    2 College of Computer Science and Technology, Zhejiang University, Hangzhou Zhejiang 310000, China
    3 Undergraduate Innovation and Entrepreneurship Training Center, China University of Mining and Technology, Xuzhou Jiangsu 221000, China
  • Received:2025-12-23 Accepted:2026-04-28 Online:2026-08-31 Published:2026-08-31
  • Contact: YAO Jun
  • Supported by:
    Jiangsu Provincial Philosophy and Social Science Planning Project(25YSB015);Fundamental Research Funds for the Central Universities(2026QNSK47)

Abstract:

Gaze-based interaction, as a natural interaction modality driven by eye movements, has demonstrated great potential for improving interaction efficiency and immersive experience in target-selection tasks within virtual reality (VR) environments, due to its fast, intuitive, and hands-free characteristics. Existing studies mainly rely on input techniques such as fixation, smooth pursuit, and gaze gestures to trigger targets through predefined input parameters and interface layouts. However, fixed dwell-time thresholds and activation areas were commonly adopted during target selection, which fail to account for individual differences and dynamic user states. This limitation often leads to increased false activations and reduced triggering efficiency. To address this issue, user interaction intention in VR was investigated from an intention-driven perspective. Five representative gaze-based selection tasks were designed in a VR environment, and task-performance data, eye-movement behaviors, and subjective intention labels were systematically collected. A multi-source eye-tracking feature space was constructed based on the acquired data. On this basis, five non-sequential models based on static features (SVM, LR, RF, LightGBM, and MLP) and three sequential deep-learning models (Bi-LSTM, 1D CNN, and CNN-LSTM) were developed and evaluated. Model performance was compared using subject-independent five-fold cross-validation and an independent test set. In addition, deployment-related factors such as online inference latency and model size were jointly considered to select the optimal model. The results indicated that the LightGBM-based intention-recognition model achieved a superior balance among recognition accuracy, inference efficiency, and deployment feasibility. This study provided an effective methodological foundation for real-time intention recognition, adaptive triggering mechanisms, and personalized interaction design in gaze-controlled virtual reality interfaces.

Key words: virtual reality, eye-controlled interaction, target selection tasks, gaze input, user intention recognition, eye movement features

CLC Number: