Welcome to Journal of Graphics

Journal of Graphics ›› 2026, Vol. 47 ›› Issue (4): 844-853.DOI: 10.11996/JG.j.2095-302X.2026040844

• Digital Design and Manufacture • Previous Articles     Next Articles

Hierarchical alignment for zero-shot sketch-based 3D shape retrieval

WANG Yutao1,2,3, YANG Chao1, KUANG Liqun1,2,3, YANG Xiaowen1,2,3, HAN Xie1,2,3, JIAO Shichao1,2,3()   

  1. 1 School of Computer Science and Technology, North University of China, Taiyuan Shanxi 030051, China
    2 Shanxi Key Laboratory of Machine Vision and Virtual Reality, Taiyuan Shanxi 030051, China
    3 Shanxi Province’s Vision Information Processing and Intelligent Robot Engineering Research Center, Taiyuan Shanxi 030051, China
  • Received:2026-03-14 Accepted:2026-05-22 Online:2026-08-31 Published:2026-08-31
  • Contact: JIAO Shichao
  • Supported by:
    National Natural Science Foundation of China(62272426);Key Research and Development Program of Shanxi Province(202402020101001);Natural Science Foundation of Shanxi(202303021212189)

Abstract:

Sketch-based 3D shape retrieval has gradually become an important research direction in the field of 3D shape retrieval, mainly because it provides an intuitive mode of human-computer interaction for users. However, in real application scenarios, the number of 3D shape categories is very large and continues to grow, which makes it hard for training data to cover all potential categories, and thus traditional supervised retrieval methods based on category labels often struggle to meet the demand for new-category retrieval in open environments. This limitation motivates the study of zero-shot sketch-based 3D shape retrieval, which aims to improve the adaptability of retrieval systems to open scenarios. In this task, there are two main difficulties: addressing the modality gap between sketches and 3D shapes, and achieving effective knowledge transfer from seen categories to unseen categories in the absence of samples from unseen classes. To address these problems, a hierarchical alignment method was designed, which performs progressive alignment at both the input level and the feature level to jointly reduce cross-modal differences and support knowledge transfer. At the input level, an adversarial learning-based pseudo-view generation method was introduced to convert sketches into pseudo multi-view representations that are closer to 3D shape distributions, thereby reducing modality discrepancy from the source. At the feature level, a shared semantic embedding space was constructed, and a multi-granularity metric learning strategy was further introduced, in which semantic consistency constraints, semantic enhancement mechanisms, and class-level proxy constraints were jointly used to improve feature representation. Experimental results on the SHREC 2013 and SHREC 2014 datasets showed that the proposed hierarchical alignment framework achieved competitive performance across multiple evaluation metrics, while ablation studies further verified the effectiveness of each component, indicating that input-level alignment and feature-level semantic alignment can work together to enhance cross-modal representation learning and improve zero-shot retrieval performance.

Key words: sketch-based 3D shape retrieval, zero-shot learning, deep learning, adversarial generation, cross-modal alignment

CLC Number: