遇见数据集

Training Hyperparameters.

收藏
Figshare2025-06-27 更新2026-04-28 收录
官方服务:

资源简介:

Large language models (LLMs) have demonstrated remarkable performance across various linguistic tasks. However, existing LLMs perform inadequately in information extraction tasks for both Chinese and English. Numerous studies attempt to enhance model performance by increasing the scale of training data. However, discrepancies in the number and type of schemas used during training and evaluation can harm model effectiveness. To tackle this challenge, we propose ChunkUIE, a unified information extraction model that supports Chinese and English. We design a chunked instruction construction strategy that randomly and reproducibly divides all schemas into chunks containing an identical number of schemas. This approach ensures that the union of schemas across all chunks encompasses all schemas. By limiting the number of schemas in each instruction, this strategy effectively addresses the performance degradation caused by inconsistencies in schema counts between training and evaluation. Additionally, we construct some challenging negative schemas using a predefined hard schema dictionary, which mitigates the model’s semantic confusion regarding similar schemas. Experimental results demonstrate that ChunkUIE enhances zero-shot performance in information extraction.

大语言模型(Large Language Models,LLMs)在各类语言任务中已展现出卓越性能。然而,现有大语言模型在中英双语的信息抽取任务中表现欠佳。诸多研究尝试通过扩大训练数据规模来提升模型性能,但训练与评估阶段所用抽取模式的数量与类型存在差异,会削弱模型的实际效能。为应对这一挑战,我们提出了支持中英双语的统一信息抽取模型ChunkUIE。我们设计了一种分块指令构建策略,该策略可随机且可复现地将所有抽取模式划分为包含相同数量模式的分块,确保所有分块的抽取模式集合的并集覆盖全部抽取模式。通过限制单条指令中的抽取模式数量,该策略有效解决了训练与评估阶段抽取模式数量不一致所导致的性能退化问题。此外,我们通过预定义的难例抽取模式词典构建了若干具有挑战性的负向抽取模式,以此缓解模型对相似抽取模式的语义混淆问题。实验结果表明,ChunkUIE可提升信息抽取任务的零样本性能。

创建时间:
2025-06-27
二维码
社区交流群
二维码
科研交流群
商业服务