SiTSE
收藏资源简介:
SiTSE数据集是由斯里兰卡莫拉图瓦大学计算机科学与工程系的团队创建的,专门用于僧伽罗语文本简化任务的评估。该数据集包含1000个复杂的僧伽罗语句子及其对应的3000个简化句子,每个复杂句子由三位不同的专家进行简化。数据集的来源是僧伽罗语官方政府文档,这些文档经过精心筛选和处理,确保了数据的质量。创建过程包括多次人工标注和反馈,确保简化句子的准确性和可读性。SiTSE数据集主要用于评估和改进僧伽罗语文本简化系统,旨在提高低资源语言文本处理的公平性和可访问性。
The SiTSE dataset was developed by a research team from the Department of Computer Science and Engineering, University of Moratuwa, Sri Lanka, specifically for evaluating Sinhala text simplification tasks. It contains 1000 complex Sinhala sentences and their corresponding 3000 simplified sentences, with each complex sentence being simplified by three distinct experts. The dataset is sourced from official Sinhala government documents, which have been rigorously screened and processed to ensure high data quality. The creation process includes multiple rounds of manual annotation and post-annotation feedback to guarantee the accuracy and readability of the simplified sentences. The SiTSE dataset is primarily used to evaluate and improve Sinhala text simplification systems, aiming to enhance the fairness and accessibility of low-resource language text processing.

- 1SiTSE: Sinhala Text Simplification Dataset and Evaluation计算机科学与工程系,莫拉图瓦大学,斯里兰卡 · 2024年



