saraiki-english-dataset
收藏资源简介:
该数据集是一个面向教育领域的学生需求识别数据集,包含英文和旁遮普语(Shahmukhi字体)双语文本。每个样本包含一个英文句子描述学生Leonardo在课堂中表现出的阅读流畅性困难,以及对应的旁遮普语翻译。数据集共包含50005个训练样本,可用于多语言文本分类、翻译或教育需求分析等相关任务。数据来源为课堂教师输入,旨在支持对特殊教育需求学生的识别与干预。
This dataset is an education-oriented student needs identification dataset, containing bilingual text in English and Punjabi (Shahmukhi script). Each sample includes an English sentence describing the reading fluency difficulties exhibited by student Leonardo in the classroom, along with its Punjabi translation. The dataset contains a total of 50,005 training samples, which can be used for tasks such as multilingual text classification, translation, or educational needs analysis. The data source is teacher input from the classroom, aiming to support the identification and intervention of students with special educational needs.
Saraiki-English 数据集概述
基本信息
- 许可证:MIT 许可证
- 主要语言:Saraiki(斯科特语,语言代码:skr)
- 数据集类型:双语对齐文本数据集(Saraiki-英语)
数据内容
该数据集包含两条主要特征字段:
- Areas of Need(英语):基于课堂教师反馈,描述学生Leonardo在阅读流利度方面的困难。
- لوڑ دے شعبے(Saraiki语):为上述英语内容的Saraiki语对应翻译,同样描述Leonardo在阅读流利度方面的学习需求。
数据集规模
| 项目 | 数值 |
|---|---|
| 训练集样本数 | 50,005 条 |
| 数据集总大小 | 6,706,006 字节(约6.4 MB) |
| 下载大小 | 3,471,685 字节(约3.3 MB) |
数据划分
- 仅包含一个划分:
train(训练集),所有数据均用于训练。
文件结构
- 数据文件路径:
data/train-*(默认配置为default,使用通配符匹配多个分片文件)。




