遇见数据集

A Procedurally Generated Multi-Domain Text Classification Dataset with Controlled Cross-Domain Lexical Overlap

收藏
Mendeley Data2026-08-09 收录
官方服务:

资源简介:

This dataset is supplementary to the publication titled "From Learned Routing to Bayesian Inference: Interpretable Evidence Combination in Modular Neural Systems. It is a procedurally created corpus for short-text categorization across many domains, where cross-domain lexical ambiguity serves as a dynamic, modifiable production parameter instead of a static or accidental characteristic of the text. The corpus is wholly synthetic in nature, produced from distinct and overlapping domain lexicons using an embedded Python script, and is not derived from any external text repositories, web-scraped materials, or human-subject information. Five sectors (sports, finance, technology, politics, entertainment) serve as the primary basis for comparison; an expansion encompassing twelve sectors is available for scalability assessment. This deposit comprises both the tangible corpus files (the generated text, immediately usable without executing any code) and the generator script that created them, enabling any reader to replicate or augment the identical corpus or new configurations. Corpus files consist of JSON Lines, featuring one document per line, containing the fields "text", "label", and "overlap_rate". Comprehensive information regarding the lexicon, generating specifications, and file organization may be found in the accompanying README. Documents: data/corpora/ (40 files, primary comparison, 5 overlap rates over 8 seeds) and data/scalability_corpora/ (15 files, configurations for 5/8/12 domains across 5 seeds), in addition to code/data_generation.py (the autonomous generator, utilizing only the standard library, devoid of dependencies).

创建时间:
2026-07-27
二维码
社区交流群
二维码
科研交流群
商业服务