欢迎访问《图学学报》

图学学报 ›› 2026, Vol. 47 ›› Issue (4): 766-775.DOI: 10.11996/JG.j.2095-302X.2026040766

• 计算机图形学与虚拟现实 • 上一篇    下一篇

虚拟现实环境下多类型眼控选择任务的用户意图识别模型研究

唐晓腾1,2, 姚君1(), 胡鹤凡3, 邵将1, 束云峰2   

  1. 1 中国矿业大学建筑与设计学院江苏 徐州 221000
    2 浙江大学计算机科学与技术学院浙江 杭州 310000
    3 中国矿业大学大学生创新训练中心江苏 徐州 221000
  • 收稿日期:2025-12-23 接受日期:2026-04-28 出版日期:2026-08-31 发布日期:2026-08-31
  • 通讯作者:姚君,E-mail:yaojun@cumt.edu.cn
  • 基金资助:
    江苏省哲学社会科学规划项目(25YSB015);中央高校基本科研业务费专项资金项目(2026QNSK47)

User intention recognition for multi-type gaze-based target selection tasks in virtual reality

TANG Xiaoteng1,2, YAO Jun1(), HU Hefan3, SHAO Jiang1, SHU Yunfeng2   

  1. 1 School of Architecture and Design, China University of Mining and Technology, Xuzhou Jiangsu 221000, China
    2 College of Computer Science and Technology, Zhejiang University, Hangzhou Zhejiang 310000, China
    3 Undergraduate Innovation and Entrepreneurship Training Center, China University of Mining and Technology, Xuzhou Jiangsu 221000, China
  • Received:2025-12-23 Accepted:2026-04-28 Published:2026-08-31 Online:2026-08-31
  • Contact: YAO Jun,E-mail:yaojun@cumt.edu.cn
  • Supported by:
    Jiangsu Provincial Philosophy and Social Science Planning Project(25YSB015);Fundamental Research Funds for the Central Universities(2026QNSK47)

摘要:

眼控交互作为一种基于视线行为的自然交互方式,凭借快捷、直观和解放双手等特性,在虚拟现实(VR)场景下的目标选择任务中展现出明显的交互效率与沉浸体验提升潜力。已有研究大多采用凝视、平滑追踪和眼势等输入方式,利用预设的输入参数和界面布局来完成实现目标触发,但在目标选择过程中普遍采用固定的停留时间和激活区域设置,忽略了个体差异和状态变化对眼动模式的影响,易造成误触率增加、触发效率降低。因此,本研究从用户交互意图角度出发,在VR环境中设计了5种典型眼控任务场景,采集任务绩效、眼动行为及主观意图标注数据,并构建多源眼动特征空间。基于该数据集,分别构建并评估了五类基于静态特征的传统机器学习模型(SVM,LR,RF,LightGBM,MLP)和三类时序深度模型(Bi-LSTM,1D CNN,CNN-LSTM),通过被试独立的五折交叉验证与独立测试集进行性能比较,并结合在线推理时延与模型体积等部署指标综合筛选出基于 LightGBM 的用户选择意图识别模型。结果表明,该模型在识别精度、推理效率与实际部署可行性之间取得了较优平衡,可为VR眼控界面中的实时意图识别、自适应触发机制及个性化交互设计提供方法支撑。

关键词: 虚拟现实, 眼控交互, 目标选择任务, 凝视输入, 用户意图识别, 眼动特征

Abstract:

Gaze-based interaction, as a natural interaction modality driven by eye movements, has demonstrated great potential for improving interaction efficiency and immersive experience in target-selection tasks within virtual reality (VR) environments, due to its fast, intuitive, and hands-free characteristics. Existing studies mainly rely on input techniques such as fixation, smooth pursuit, and gaze gestures to trigger targets through predefined input parameters and interface layouts. However, fixed dwell-time thresholds and activation areas were commonly adopted during target selection, which fail to account for individual differences and dynamic user states. This limitation often leads to increased false activations and reduced triggering efficiency. To address this issue, user interaction intention in VR was investigated from an intention-driven perspective. Five representative gaze-based selection tasks were designed in a VR environment, and task-performance data, eye-movement behaviors, and subjective intention labels were systematically collected. A multi-source eye-tracking feature space was constructed based on the acquired data. On this basis, five non-sequential models based on static features (SVM, LR, RF, LightGBM, and MLP) and three sequential deep-learning models (Bi-LSTM, 1D CNN, and CNN-LSTM) were developed and evaluated. Model performance was compared using subject-independent five-fold cross-validation and an independent test set. In addition, deployment-related factors such as online inference latency and model size were jointly considered to select the optimal model. The results indicated that the LightGBM-based intention-recognition model achieved a superior balance among recognition accuracy, inference efficiency, and deployment feasibility. This study provided an effective methodological foundation for real-time intention recognition, adaptive triggering mechanisms, and personalized interaction design in gaze-controlled virtual reality interfaces.

Key words: virtual reality, eye-controlled interaction, target selection tasks, gaze input, user intention recognition, eye movement features

中图分类号: