A 10-Million-Row Sinhala Narrative Corpus for Low-Resource NLP: Dataset Construction, Statistical Characterisation, and Ethical Considerations
收藏资源简介:
sinhala_stories is a crowdsourced corpus of Sinhala-language narrative text comprising 10,949,004 rows (~1.95 GB) in Apache Parquet format, collected through a public Streamlit submission portal integrated with the Hugging Face Hub. Each submission was validated for Sinhala script content, length, language-identification confidence, and heuristic spam/duplicate checks before being merged into the corpus via an automated CI/CD pipeline. Content warning: a substantial portion of the corpus consists of explicit adult narrative content, submitted with minimal moderation beyond character-level validation. Anyone training generative models on this data should apply content filtering and safety alignment before deployment. See the accompanying data note for full ethical considerations. Dataset: https://huggingface.co/datasets/Isuru0x01/sinhala_storiesCollection app source: https://github.com/isuru0x01/Sinhala-Stories-Dataset-Creator



