conceptnet5/conceptnet5
收藏资源简介:
ConceptNet是一个多语言知识库,代表了人们使用的单词和短语以及它们之间的常识关系。ConceptNet中的知识是从多种资源中收集的,包括众包资源(如Wiktionary和Open Mind Common Sense)、有目的的游戏(如Verbosity和nadya.jp)以及专家创建的资源(如WordNet和JMDict)。该数据集旨在提供从各种来源提取的常识关系的训练数据。数据集是多语言的,支持的语言包括英语、法语、意大利语、德语、西班牙语、俄语、葡萄牙语、日语、荷兰语和中文等。
ConceptNet is a multilingual knowledge base that encapsulates words and phrases in daily use and the common-sense relationships between them. The knowledge contained in ConceptNet is curated from a wide range of sources, including crowdsourced resources such as Wiktionary and Open Mind Common Sense, games with a purpose like Verbosity and nadya.jp, as well as expert-curated resources such as WordNet and JMDict. This dataset is designed to provide training data for common-sense relationships extracted from diverse sources. The dataset supports multiple languages, including English, French, Italian, German, Spanish, Russian, Portuguese, Japanese, Dutch, Chinese and others.
数据集概述
基本信息
- 数据集名称: Conceptnet5
- 许可证: cc-by-4.0
- 语言: de, en, es, fr, it, ja, nl, pt, ru, zh
- 多语言性: 单语种
- 大小类别: 100K<n<1M, 10M<n<100M, 1M<n<10M
- 源数据: 原始数据
- 任务类别: 文本分类
- 任务ID: 多类分类
配置信息
-
conceptnet5
- 特征:
- sentence: string
- full_rel: string
- rel: string
- arg1: string
- arg2: string
- lang: string
- extra_info: string
- weight: float32
- 分割:
- train: 34074917个样本, 11493772756字节
- 下载大小: 1280623369字节
- 数据集大小: 11493772756字节
- 特征:
-
omcs_sentences_free
- 特征:
- sentence: string
- raw_data: string
- lang: string
- 分割:
- train: 898160个样本, 174810230字节
- 下载大小: 72941617字节
- 数据集大小: 174810230字节
- 特征:
-
omcs_sentences_more
- 特征:
- sentence: string
- raw_data: string
- lang: string
- 分割:
- train: 2001735个样本, 341421867字节
- 下载大小: 129630544字节
- 数据集大小: 341421867字节
- 特征:
数据文件
-
conceptnet5
- 数据文件:
- train: conceptnet5/train-*
- 数据文件:
-
omcs_sentences_free
- 数据文件:
- train: omcs_sentences_free/train-*
- 数据文件:
-
omcs_sentences_more
- 数据文件:
- train: omcs_sentences_more/train-*
- 数据文件:




