HPLT/OpenLID-v3
收藏资源简介:
OpenLID-v3数据集是OpenLID-v2的更新版本,覆盖194种语言变体及一个非语言类别,旨在训练高覆盖率的语言识别模型。每个数据实例包含一行文本、由ISO 639-3语言代码和ISO 15924脚本代码组成的语言标签,以及来源标签。数据集仅提供训练分割,适用于提升语言识别精度,特别是对低资源语言和密切相关的语言。使用需注意社会影响(如覆盖未充分服务语言)和潜在偏见(如可能强化语言权力不平衡),许可允许非商业用途开放使用。
OpenLID-v3 is an updated version of the OpenLID-v2 dataset, covering 194 language varieties plus a not-a-language class. It is intended for training high-coverage language identification models. Each entry consists of a line of text, a language label combining an ISO 639-3 language code and an ISO 15924 script code, and a source tag. Only a train split is provided, designed to improve language identification precision, especially for low-resource and closely related languages. Considerations include social impact (e.g., covering under-served languages) and biases (e.g., potential reinforcement of power imbalances), with licensing allowing open use for non-commercial purposes.




