遇见数据集

NobleMind/Context-to-knowledge

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个多任务自然语言处理数据集,包含四个配置:decontextualization(去上下文化)、default(默认)、ner(命名实体识别)和relation_extraction(关系抽取)。每个配置针对不同任务设计,例如decontextualization配置包含上下文段落、依赖句子和目标独立句子等字段,用于处理文本去上下文化;ner配置包含原始文本、令牌和NER标签等,用于实体识别;relation_extraction配置包含句子、标记文本和实体关系信息,用于关系抽取。数据集基于维基百科数据构建,支持多种语言,训练集示例数量为10025(除default配置为10)。

This is a multi-task natural language processing dataset that encompasses four configurations: decontextualization, default, ner, and relation_extraction. Each configuration is designed for specific tasks. For example, the decontextualization configuration includes fields such as context paragraph, dependent sentence, and target independent sentence, and is used for text decontextualization tasks; the ner configuration contains original text, tokens, and NER tags for named entity recognition; the relation_extraction configuration covers sentences, tokenized text, and entity relationship information for relation extraction. The dataset is constructed based on Wikipedia data and supports multiple languages, with the number of training set examples being 10025, except for the default configuration which has 10 training examples.

提供机构:
NobleMind
二维码
社区交流群
二维码
科研交流群
商业服务