arXiv 论文速递

Snapshot: 20260515_0457

Unlocking Patch-Level Features for CLIP-Based Class-Incremental Learning

Authors: Hao Sun, Zi-Jun Ding, Da-Wei Zhou

First: 2026-05-13T17:56:23+00:00 · Latest: 2026-05-13T17:56:23+00:00

Abstract

Class-Incremental Learning (CIL) enables models to continuously integrate new knowledge while mitigating catastrophic forgetting. Driven by the remarkable generalization of CLIP, leveraging pre-trained vision-language models has become a dominant paradigm in CIL. However, current work primarily focuses on aligning global image embeddings (i.e., [CLS] token) with their corresponding text prompts (i.e., [EOS] token). Despite their good performance, we find that they discard the rich patch-level semantic information inherent in CLIP's encoders. For instance, when recognizing a rabbit, local patches may encode its distinctive cues, such as long ears and a fluffy tail, which can provide complementary evidence for recognition. Based on the above observation, we propose SPA (Semantic-guided Patch-level Alignment) for CLIP-based CIL, which aims to awaken long-neglected local representations within CLIP. Specifically, for each class, we first construct representative and diverse visual samples and feed them to GPT-5 as visual guidance to generate class-wise semantic descriptions. These descriptions are used to guide the selection of discriminative patch-level visual features. Building upon these selected patches, we further employ optimal transport to align selected patch tokens with semantic tokens from class-wise descriptions, yielding a structured cross-modal alignment that improves recognition. Furthermore, we introduce task-specific projectors for effective adaptation to downstream incremental tasks, and sample pseudo-features from stored class-wise Gaussian statistics to calibrate old-class representations, thereby mitigating catastrophic forgetting. Extensive experiments demonstrate that SPA achieves state-of-the-art performance.

中文标题/摘要

标题：基于CLIP的类增量学习中解锁patches级别的特征

类增量学习(CIL)使模型能够不断整合新知识并减轻灾难性遗忘。受CLIP卓越泛化能力的驱动，利用预训练的跨模态模型已成为CIL中的主导范式。然而，当前工作主要集中在对齐全局图像嵌入（即[CLS]标记）与其相应的文本提示（即[EOS]标记）。尽管它们表现出色，但我们发现它们忽略了CLIP编码器中固有的丰富patches级别的语义信息。例如，在识别兔子时，局部patches可能编码其独特的线索，如长耳朵和蓬松的尾巴，这些线索可以为识别提供补充证据。基于上述观察，我们提出了基于CLIP的CIL方法SPA（语义引导的patches级别对齐），旨在唤醒CLIP中长期被忽视的局部表示。具体而言，对于每个类别，我们首先构建代表性且多样的视觉样本，并将其输入GPT-5作为视觉指导以生成类别级别的语义描述。这些描述用于指导选择具有区分性的patches级别的视觉特征。在此基础上，我们进一步利用最优传输将所选的patches标记与类别级别的描述中的语义标记对齐，从而获得一种结构化的跨模态对齐，以提高识别效果。此外，我们引入了特定任务的投影器以有效适应下游增量任务，并从存储的类别级别的高斯统计中采样伪特征以校准旧类别的表示，从而减轻灾难性遗忘。广泛的实验表明，SPA达到了最先进的性能。

Summary / 总结

The paper proposes SPA (Semantic-guided Patch-level Alignment) to enhance CLIP-based class-incremental learning by leveraging patch-level semantic information. It constructs class-wise semantic descriptions using GPT-5 and aligns these with selected patch tokens via optimal transport, improving recognition. Task-specific projectors and pseudo-features from stored Gaussian statistics are also introduced to mitigate catastrophic forgetting. SPA achieves state-of-the-art performance in extensive experiments.

论文提出了SPA（语义引导的局部特征对齐），通过利用局部语义信息来增强基于CLIP的类增量学习。它使用GPT-5构建类别的语义描述，并通过最优传输将这些描述与选定的局部特征对齐，从而提高识别效果。此外，还引入了任务特定的投影器和从存储的高斯统计中抽取的伪特征来减轻灾难性遗忘。广泛的实验表明，SPA达到了最先进的性能。

Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context

Authors: Zhaowei Wang, Lishu Luo, Haodong Duan, Weiwei Liu, Sijin Wu, Ji Luo, Shen Yan, Shuai Peng, Sihang Yuan, Chaoyi Huang, Yi Lin, Yangqiu Song

First: 2026-05-13T17:52:53+00:00 · Latest: 2026-05-13T17:52:53+00:00

Comments: work in progress