遇见数据集

Datasets for Reproducing the Experiments in "Early Language Learning via Spreading Activation and Category Exploration in Complex Networks"

收藏
Zenodo2026-06-22 更新2026-06-17 收录
官方服务:

资源简介:

This repository contains the datasets used to reproduce the experiments presented in the work Early Language Learning via Spreading Activation and Category Exploration in Complex Networks (to appear on arXiv). The repository includes preprocessed datasets for four languages: German (de), English (en), Dutch (nl), and Rioplatense Spanish (sp). Network Structure The files free_de, free_en, free_nl, and free_sp are derived from the association norms provided by the Small World of Words (SWOW) project. Specifically, we consider the releases SWOW-DE25, SWOW-EN18, SWOW-NL13, and SWOW-RP22, respectively. The files syn_ant_hier_de, syn_ant_hier_en, syn_ant_hier_nl, and syn_ant_hier_sp contain semantic relations corresponding to synonymy, antonymy, and hypo/hypernymy relations, derived from WordNet resources: NLTK WordNet for English, Dutch, and Spanish, and Open German WordNet. The files phon_de, phon_en, phon_nl, and phon_sp contain phonological relations. Phonetic transcriptions are obtained using the CMUdict library for English and the epitran Python package for German, Dutch, and Spanish. Edges are introduced between pairs of words whose phoneme strings have an edit distance d<=2. The files graph_de, graph_en, graph_nl, and graph_sp are the effective network representations obtained by aggregating the files described above, and directly used in all experiments presented. Ground-Truth Validation Orderings, Age-of-Acquisition values, and lexical categories (in the form of CDIs) are derived from the Wordbank database, which is based on the MacArthur-Bates Communicative Development Inventories (CDIs).

提供机构:
Zenodo
创建时间:
2026-06-15
二维码
社区交流群
二维码
科研交流群
商业服务