遇见数据集

A 10-Million-Row Sinhala Narrative Corpus for Low-Resource NLP: Dataset Construction, Statistical Characterisation, and Ethical Considerations

收藏
Zenodo2026-07-04 更新2026-08-01 收录
官方服务:

资源简介:

sinhala_stories is a crowdsourced corpus of Sinhala-language narrative text comprising 10,949,004 rows (~1.95 GB) in Apache Parquet format, collected through a public Streamlit submission portal integrated with the Hugging Face Hub. Each submission was validated for Sinhala script content, length, language-identification confidence, and heuristic spam/duplicate checks before being merged into the corpus via an automated CI/CD pipeline. Content warning: a substantial portion of the corpus consists of explicit adult narrative content, submitted with minimal moderation beyond character-level validation. Anyone training generative models on this data should apply content filtering and safety alignment before deployment. See the accompanying data note for full ethical considerations. Dataset: https://huggingface.co/datasets/Isuru0x01/sinhala_storiesCollection app source: https://github.com/isuru0x01/Sinhala-Stories-Dataset-Creator

提供机构:
Zenodo
创建时间:
2026-07-04
二维码
社区交流群
二维码
科研交流群
商业服务