procesaur/cirilica
收藏资源简介:
该数据集名为Ћирилица/Ćirilica,是一个塞尔维亚语语料库,包含原始西里尔字母文档及其对应的拉丁字母版本。它适用于训练机器学习模型和测试转写(西里尔字母与拉丁字母之间转换)解决方案。初始版本包含来自Знање(Znanje)和Википедије(Vikipedije)语料库的大约2亿词,数据规模在1亿到10亿之间。数据集以CC BY-SA 4.0许可证发布,文件格式为JSONL,可通过Hugging Face的datasets库加载使用。
The dataset is named Ћирилица/Ćirilica and is a Serbian language corpus consisting of original Cyrillic script documents and their Latin script counterparts. It is suitable for training models and testing solutions for transliteration (conversion between Cyrillic and Latin scripts). The initial version includes approximately 200 million words from the Znanje and Wikipedia corpora, with a size category of 100M to 1B. It is licensed under CC BY-SA 4.0, stored in JSONL files, and can be loaded via the Hugging Face datasets library.



