VC-SLAM Versatile Corpus for Semantic Labeling And Modeling
收藏资源简介:
Benchmark Corpus for semantic labeling and modeling. This corpus contains 101 data sets from different open data portals.<br> Each data set consists of the following data: Raw csv data [rawdata_csv] Large json data sample [json_sample_large] Small json data sample [json_sample_small] Raw data in csv format [rawdata_csv] Raw data samples in csv format [rawdata_csv_samples] Mappings to translate between csv and json files [csv_json_mappings] Textual description / Metadata [descriptions] Semantic model as rdf/ttl [semantic_models] Mappings describing mapping between raw data attributes and concepts from the ontology [mappings] List of attributes that have been ignored during modeling [ignored_attributes] Additionally the corpus contains a target ontology as rdf/ttl [ontology]. The individual data sets are licensed by the licenses specified in the attached Excel sheet (DataSetOverview.xlsx) These data are provided "as is", without any warranties of any kind. The data are provided under the Creative Commons Attribution 4.0 International license.
用于语义标注与建模的基准语料库(Benchmark Corpus)。该语料库涵盖来自不同开放数据门户的101个数据集。<br>每个数据集均包含以下类型的数据:原始CSV格式数据 [rawdata_csv]、大型JSON格式数据样本 [json_sample_large]、小型JSON格式数据样本 [json_sample_small]、CSV格式原始数据 [rawdata_csv]、CSV格式原始数据样本 [rawdata_csv_samples]、用于实现CSV与JSON文件间格式转换的映射文件 [csv_json_mappings]、文本描述与元数据 [descriptions]、采用RDF/TTL格式存储的语义模型 [semantic_models]、用于描述原始数据属性与本体(Ontology)概念间映射关系的映射文件 [mappings]、建模过程中已忽略的属性列表 [ignored_attributes]。<br>此外,该语料库还包含一个采用RDF/TTL格式存储的目标本体(Ontology) [ontology]。各数据集的授权协议详见附件Excel表格《DataSetOverview.xlsx》。本数据集按“现状”提供,不附带任何形式的担保。本数据集采用知识共享署名4.0国际许可协议(Creative Commons Attribution 4.0 International)进行授权。



