遇见数据集

CeyNews: A Sri Lankan Trilingual News Corpus

收藏
Zenodo2026-06-07 更新2026-06-12 收录
官方服务:

资源简介:

CeyNews: A Trilingual Sri Lankan News Corpus This Zenodo release contains the CeyNews dataset, a large-scale trilingual news corpus for Sri Lanka covering Sinhala, Tamil, and English. The dataset includes over 1.02 million news articles collected from three popular Sri Lankan news outlets. The articles span the period from 2013 to 2026 and contain approximately 150 million tokens in total. Each record includes the news article text together with available metadata, including the news source, publication timestamp, headline, category, URL, and language. The dataset has been preprocessed to improve consistency and usability. Preprocessing includes deduplication and Unicode/script normalisation for Sinhala and Tamil text. The dataset is intended to support research in multilingual NLP, low-resource language processing, cross-lingual transfer learning, LLM adaptation, Sri Lankan and South Asian digital humanities, and comparative multilingual news analysis. The dataset can also be used to construct downstream NLP tasks such as news source identification, news category classification, and headline generation.

提供机构:
Zenodo
创建时间:
2026-06-07
二维码
社区交流群
二维码
科研交流群
商业服务