面向高效文档表征的略读感知对比学习 (Skim-Aware Contrastive Learning for Efficient Document Representation) - 专知论文

会员服务 ·

0

对比学习 · 法律 · Transformer · 稀疏 · 资源消耗 ·

2025 年 12 月 30 日

Skim-Aware Contrastive Learning for Efficient Document Representation

翻译：面向高效文档表征的略读感知对比学习

Waheed Ahmed Abro,Zied Bouraoui

Although transformer-based models have shown strong performance in word- and sentence-level tasks, effectively representing long documents, especially in fields like law and medicine, remains difficult. Sparse attention mechanisms can handle longer inputs, but are resource-intensive and often fail to capture full-document context. Hierarchical transformer models offer better efficiency but do not clearly explain how they relate different sections of a document. In contrast, humans often skim texts, focusing on important sections to understand the overall message. Drawing from this human strategy, we introduce a new self-supervised contrastive learning framework that enhances long document representation. Our method randomly masks a section of the document and uses a natural language inference (NLI)-based contrastive objective to align it with relevant parts while distancing it from unrelated ones. This mimics how humans synthesize information, resulting in representations that are both richer and more computationally efficient. Experiments on legal and biomedical texts confirm significant gains in both accuracy and efficiency.

翻译：尽管基于Transformer的模型在词级和句级任务中表现出色，但有效表征长文档（尤其在法律和医学等领域）仍然具有挑战性。稀疏注意力机制虽能处理更长输入，但资源消耗大且常难以捕获全文语境。分层Transformer模型提供了更好的效率，但未能清晰阐释文档不同部分间的关联机制。相比之下，人类常通过略读文本、聚焦关键段落来把握整体主旨。受此人类认知策略启发，我们提出一种新型自监督对比学习框架以增强长文档表征。该方法随机掩蔽文档的某个段落，并基于自然语言推理（NLI）的对比学习目标，使该段落与相关部分对齐，同时与无关部分分离。这种机制模拟了人类整合信息的方式，最终生成既语义丰富又计算高效的表征。在法律与生物医学文本上的实验验证了该方法在准确性与效率上的显著提升。

0

相关内容

对比学习

通过潜在空间的对比损失最大限度地提高相同数据样本的不同扩充视图之间的一致性来学习表示。对比式自监督学习技术是一类很有前途的方法，它通过学习编码来构建表征，编码使两个事物相似或不同

【KDD2024】面向鲁棒推荐的决策边界感知图对比学习

【KDD2024】面向鲁棒推荐的决策边界感知图对比学习

专知会员服务

21+阅读 · 2024年8月8日

如何检测大模型“幻觉”？剑桥提出SelfCheckGPT: 针对生成型大型语言模型的零资源黑盒子幻觉检测

如何检测大模型“幻觉”？剑桥提出SelfCheckGPT: 针对生成型大型语言模型的零资源黑盒子幻觉检测

专知会员服务

43+阅读 · 2023年8月22日

【CVPR 2022】基于实例深度估计的统一深度感知全景分割 PanopticDepth: Per-Instance Depth Estimation for Unified Depth-Aware Panoptic Segmentation

【CVPR 2022】基于实例深度估计的统一深度感知全景分割 PanopticDepth: Per-Instance Depth Estimation for Unified Depth-Aware Panoptic Segmentation

专知会员服务

18+阅读 · 2022年3月19日

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

专知会员服务

87+阅读 · 2020年8月28日

元迁移学习的小样本学习，Meta-transfer Learning for Few-shot Learning

元迁移学习的小样本学习，Meta-transfer Learning for Few-shot Learning

专知会员服务

159+阅读 · 2020年2月29日

【AAAI2021】知识图谱增强的预训练模型的生成式常识推理

【AAAI2021】知识图谱增强的预训练模型的生成式常识推理

专知

29+阅读 · 2021年1月25日

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

专知

11+阅读 · 2020年8月28日

Python图像处理，366页pdf，Image Operators Image Processing in Python

Python图像处理，366页pdf，Image Operators Image Processing in Python

专知

15+阅读 · 2020年7月23日

【阿里巴巴-WWW2020】对抗性多模态表示学习的点击率预测，Adversarial Multimodal RL

【阿里巴巴-WWW2020】对抗性多模态表示学习的点击率预测，Adversarial Multimodal RL

专知

11+阅读 · 2020年3月17日

论文浅尝 | 当知识图谱遇上零样本学习——零样本学习综述

论文浅尝 | 当知识图谱遇上零样本学习——零样本学习综述

开放知识图谱

22+阅读 · 2018年9月26日

语义Web知识库补全关键技术研究

国家自然科学基金

17+阅读 · 2017年12月31日

基于多样化查询的多标记主动学习研究

国家自然科学基金

0+阅读 · 2015年12月31日

不确定知识图谱中面向结构查询的众包清洗研究

国家自然科学基金

4+阅读 · 2015年12月31日

面向甲骨学知识图谱的实体发现及语义关系挖掘研究

国家自然科学基金

3+阅读 · 2015年12月31日

面向大规模多步学习问题的学习分类元系统技术研究

国家自然科学基金

5+阅读 · 2015年12月31日

Beyond Degradation Redundancy: Contrastive Prompt Learning for All-in-One Image Restoration

Arxiv

0+阅读 · 2025年12月30日

Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance

Arxiv

0+阅读 · 2025年12月29日

Object-Centric Representation Learning for Enhanced 3D Scene Graph Prediction

Arxiv

0+阅读 · 2025年12月29日

Fast Inference of Visual Autoregressive Model with Adjacency-Adaptive Dynamical Draft Trees

Arxiv

0+阅读 · 2025年12月26日

Adaptive Focus Memory for Language Models

Arxiv

0+阅读 · 2025年12月24日

VIP会员

文章信息

相关主题

相关VIP内容

【KDD2024】面向鲁棒推荐的决策边界感知图对比学习

【KDD2024】面向鲁棒推荐的决策边界感知图对比学习

专知会员服务

21+阅读 · 2024年8月8日

如何检测大模型“幻觉”？剑桥提出SelfCheckGPT: 针对生成型大型语言模型的零资源黑盒子幻觉检测

如何检测大模型“幻觉”？剑桥提出SelfCheckGPT: 针对生成型大型语言模型的零资源黑盒子幻觉检测

专知会员服务

43+阅读 · 2023年8月22日

【CVPR 2022】基于实例深度估计的统一深度感知全景分割 PanopticDepth: Per-Instance Depth Estimation for Unified Depth-Aware Panoptic Segmentation

【CVPR 2022】基于实例深度估计的统一深度感知全景分割 PanopticDepth: Per-Instance Depth Estimation for Unified Depth-Aware Panoptic Segmentation

专知会员服务

18+阅读 · 2022年3月19日

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

专知会员服务

87+阅读 · 2020年8月28日

元迁移学习的小样本学习，Meta-transfer Learning for Few-shot Learning

元迁移学习的小样本学习，Meta-transfer Learning for Few-shot Learning

专知会员服务

159+阅读 · 2020年2月29日

热门VIP内容

开通专知VIP会员享更多权益服务

生成式人工智能导论：可靠性、负责任开发及实际应用（第二版）

《2025财年美陆军转型倡议（ATI）部队结构与组织提案》

【CMU博士论文】分布偏移下的可信机器学习

智能体 EDA 的曙光：自主数字芯片设计综述

相关资讯

【AAAI2021】知识图谱增强的预训练模型的生成式常识推理

【AAAI2021】知识图谱增强的预训练模型的生成式常识推理

专知

29+阅读 · 2021年1月25日

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

专知

11+阅读 · 2020年8月28日

Python图像处理，366页pdf，Image Operators Image Processing in Python

Python图像处理，366页pdf，Image Operators Image Processing in Python

专知

15+阅读 · 2020年7月23日

【阿里巴巴-WWW2020】对抗性多模态表示学习的点击率预测，Adversarial Multimodal RL

【阿里巴巴-WWW2020】对抗性多模态表示学习的点击率预测，Adversarial Multimodal RL

专知

11+阅读 · 2020年3月17日

论文浅尝 | 当知识图谱遇上零样本学习——零样本学习综述

论文浅尝 | 当知识图谱遇上零样本学习——零样本学习综述

开放知识图谱

22+阅读 · 2018年9月26日

相关论文

Beyond Degradation Redundancy: Contrastive Prompt Learning for All-in-One Image Restoration

Arxiv

0+阅读 · 2025年12月30日

Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance

Arxiv

0+阅读 · 2025年12月29日

Object-Centric Representation Learning for Enhanced 3D Scene Graph Prediction

Arxiv

0+阅读 · 2025年12月29日

Fast Inference of Visual Autoregressive Model with Adjacency-Adaptive Dynamical Draft Trees

Arxiv

0+阅读 · 2025年12月26日

Adaptive Focus Memory for Language Models

Arxiv

0+阅读 · 2025年12月24日

相关基金

语义Web知识库补全关键技术研究

国家自然科学基金

17+阅读 · 2017年12月31日

基于多样化查询的多标记主动学习研究

国家自然科学基金

0+阅读 · 2015年12月31日

不确定知识图谱中面向结构查询的众包清洗研究

国家自然科学基金

4+阅读 · 2015年12月31日

面向甲骨学知识图谱的实体发现及语义关系挖掘研究

国家自然科学基金

3+阅读 · 2015年12月31日

面向大规模多步学习问题的学习分类元系统技术研究

国家自然科学基金

5+阅读 · 2015年12月31日

微信扫码咨询专知VIP会员