计算机科学
关系抽取
杠杆(统计)
标记数据
人工智能
瓶颈
一致性(知识库)
机器学习
数据挖掘
概化理论
正规化(语言学)
水准点(测量)
接头(建筑物)
关系(数据库)
多任务学习
组分(热力学)
任务(项目管理)
数据建模
语言模型
估计员
自然语言处理
可比性
过度拟合
稳健性(进化)
成对比较
数据集成
结构化预测
统一模型
信息抽取
粒度
面子(社会学概念)
作者
Xiyang Liu,Hang Shen,Honglei Qi,Yuanfei Dai
摘要
Joint entity and relation extraction represents a critical task in knowledge representation, but often suffers from the bottleneck of requiring large amounts of labeled data, which is expensive and laborious to obtain. While semi-supervised learning (SSL) offers a way to leverage unlabeled data, traditional methods face limitations in generating high-quality, diverse augmentations for text. This paper introduces a novel framework that synergistically combines SSL with large language models (LLMs) to improve joint entity and relation extraction, especially in low-resource settings. Our approach utilizes LLMs to generate semantically coherent and diverse augmented data from unlabeled samples. These augmented samples, along with limited labeled data, are used within an SSL framework employing consistency regularization and pseudo-labeling to train the extraction model. Crucially, the framework incorporates an iterative refinement mechanism where the performance of the SSL component informs the parameter-efficient fine-tuning of the LLM, leading to progressively better data augmentation and model accuracy. We demonstrate through extensive experiments on four benchmark datasets that our proposed method significantly outperforms existing state-of-the-art approaches, particularly when labeled data is scarce. The framework’s design is adaptable and can be integrated with various existing joint extraction models, showcasing its generalizability and practical utility.
科研通智能强力驱动
Strongly Powered by AbleSci AI