impresso-project/impresso-mediaagencies-ner-dataset
收藏资源简介:
Impresso媒体来源数据集是一个用于标记分类的精选数据集,专门针对Impresso历史报纸文本中的新闻机构和广播电台提及进行命名实体识别。v0.1版本数据源自法语/德语的HIPE风格新闻机构注释,已转换为JSONL格式,并经过人工审查以解决当前模型开发/测试中的分歧,同时根据v2.0注释指南进行了更新。当前指南注释每个显式的规范媒体来源组织提及,而不仅仅是来源归属用途。数据集包含训练、验证和测试文件,支持使用Hugging Face的datasets库加载。标签策略使用org.ent.pressagency.<规范ID>和org.ent.radiostation.<规范ID>命名空间,并排除了如unk等被禁止的遗留标签。数据集许可证暂定为other,待最终分发条款确认。
The Impresso Media Source Dataset is a curated dataset for sequence labeling and classification, specifically focused on named entity recognition (NER) of news agency and radio station mentions within Impresso historical newspaper corpora. The v0.1 version of the dataset is derived from HIPE-style annotations of news agencies in French and German, converted to JSONL format, manually reviewed to resolve discrepancies encountered during current model development and testing, and updated in accordance with the v2.0 annotation guidelines. The current guidelines annotate every explicit, canonical mention of media source organizations, rather than only those used for source attribution purposes. The dataset includes training, validation, and test splits, and supports loading via the Hugging Face datasets library. The labeling strategy employs the namespaces `org.ent.pressagency.<canonical ID>` and `org.ent.radiostation.<canonical ID>`, and excludes deprecated legacy labels such as `unk`. The dataset license is provisionally set to `other`, pending final confirmation of distribution terms.




