遇见数据集

filbench/cebuano-readability

收藏
Hugging Face2025-03-12 更新2026-01-03 收录
官方服务:

资源简介:

--- dataset_info: features: - name: source dtype: string - name: text dtype: string - name: label dtype: int64 splits: - name: test num_bytes: 429614 num_examples: 350 download_size: 233244 dataset_size: 429614 configs: - config_name: default data_files: - split: test path: data/test-* license: cc-by-4.0 --- Source: https://github.com/imperialite/cebuano-readability > We asked permission from one of the authors to include this dataset to our catalog effort. We copy a portion of the README in this dataset card. # Baseline Readability Assessment Model for Cebuano This repository contains the code and datasets from Bloom, Let's Read Asia, and Department of Education (DepEd) websites used for developing the first ML-based baseline for readability assessment in the Cebuano language described in the paper **A Baseline Readability Model for Cebuano**. Paper Link: https://arxiv.org/abs/2203.17225 All reading materials collected for the study are free-to-download from their original websites and licensed with CC BY 4.0 which means they are also free to distribute in any form given proper citation. Please also add the following citation to your paper/presentation if you use the resources found in this repository: ``` Reyes, L. L. A., & Ibañez, M. A., & Sapinit, R., & Hussien, M., Imperial, J. M. (2022). A Baseline Readability Model for Cebuano. arXiv preprint arXiv:2203.17225. ``` ## Contact If you need any help reproducing the results, please don't hesitate to contact the author through **Joseph Marvin Imperial** <br/> jrimperial@national-u.edu.ph <br/> www.josephimperial.com

数据集信息: 特征字段: - 字段名:source,数据类型:字符串(string) - 字段名:text,数据类型:字符串(string) - 字段名:label,数据类型:64位整型(int64) 划分集: - 划分名称:测试集(test),字节数:429614,样本数:350 下载大小:233244 数据集占用大小:429614 配置项: - 配置名称:默认配置(default) 数据文件: - 划分:测试集(test) 路径:data/test-* 许可证:CC BY 4.0 --- 源地址:https://github.com/imperialite/cebuano-readability > 我们已获得其中一位作者的许可,将此数据集纳入我们的目录构建工作。本数据集卡片中复制了该数据集的部分README文件内容。 # 宿务语(Cebuano)基准可读性评估模型 本仓库包含来自Bloom、Let's Read Asia以及菲律宾教育部(Department of Education, DepEd)官网的代码与数据集,用于开发论文**《宿务语基准可读性模型》**中提及的首个基于机器学习(Machine Learning)的宿务语可读性评估基准模型。 论文链接:https://arxiv.org/abs/2203.17225 本研究收集的所有阅读材料均可从原始官网免费下载,且采用CC BY 4.0许可证授权,这意味着在注明正确引用来源的前提下,可免费以任何形式分发这些材料。 若您使用本仓库中的资源,请在您的论文/演示文稿中添加如下引用信息: Reyes, L. L. A., & Ibañez, M. A., & Sapinit, R., & Hussien, M., Imperial, J. M. (2022). A Baseline Readability Model for Cebuano. arXiv预印本 arXiv:2203.17225. ## 联系方式 若您需要协助复现实验结果,请随时通过以下方式联系作者: **约瑟夫·马文·恩佩里亚尔(Joseph Marvin Imperial)** <br/> jrimperial@national-u.edu.ph <br/> www.josephimperial.com

提供机构:
filbench
二维码
社区交流群
二维码
科研交流群
商业服务