osm-polygon-wikidata-sentence-relevance
收藏资源简介:
OSM多边形Wikidata句子相关性数据集是一个从OpenStreetMap(OSM)多边形及其关联的维基百科/Wikivoyage页面提取的规范化句子集合。每个数据行代表一个在特定多边形、语言和内容哈希范围内去重后的句子出现。当前版本仅覆盖阿富汗地区,包含54,462个句子行,涉及161个唯一多边形、154个唯一Wikidata实体、1,680个唯一文档和115种语言。数据主要来源于维基百科(54,152行)和Wikivoyage(310行)。句子经过多语言模型分割、边界修复、文本规范化等处理,并在同一多边形和语言内精确去重。该数据集适用于多语言语料库分析以及研究不同语言下对地点的描述方式,但并非标注的相关性、相似性或分类数据集。数据遵循上游数据源的许可条款,包括OpenStreetMap的ODbL许可和维基媒体的CC BY-SA许可。
The OSM Polygon Wikidata Sentence Relevance Dataset contains normalized sentences extracted from OpenStreetMap (OSM) polygons and their associated Wikipedia/Wikivoyage pages. Each data row represents a deduplicated sentence occurrence, scoped to a specific polygon, language, and content hash. The current version only covers the Afghanistan region (afghanistan shard). The dataset includes 54,462 sentence rows, involving 161 unique polygons, 154 unique Wikidata entities, 1,680 unique documents, and 115 languages. Data is primarily sourced from Wikipedia (54,152 rows) and Wikivoyage (310 rows). Each sentence undergoes processing such as multilingual model segmentation, boundary fixing, text normalization, and exact deduplication within the same polygon and language. The dataset is suitable for multilingual corpus analysis and studying descriptions of places across languages, but it is not an annotated relevance, similarity, or classification dataset. Data follows the licensing terms of upstream sources, including ODbL for OpenStreetMap and CC BY-SA for Wikimedia.





