telugu-script-sanskrit-text
收藏资源简介:
本数据集是一个泰卢固文文本语料库,包含26,200个训练样本。数据来源于Kaggle上的一个梵文文本语料库,并利用aksharamukha工具将原始梵文文本转写为泰卢固文。数据集的核心特征是一个名为text的字符串字段,用于存储转写后的泰卢固文文本内容。数据集总大小约为623MB,下载大小约为234MB。该数据集适用于自然语言处理任务,如泰卢固文的语言建模、文本生成或与梵文相关的跨语言研究。数据集采用CC BY-SA 4.0许可证发布。
This dataset is a Telugu text corpus containing 26,200 training samples. The data originates from a Sanskrit text corpus on Kaggle, and the original Sanskrit text is transliterated into Telugu using the aksharamukha tool. The core feature of the dataset is a string field named text that stores the transliterated Telugu text content. The total size of the dataset is approximately 623MB, with a download size of about 234MB. It is suitable for natural language processing tasks, such as Telugu language modeling, text generation, or cross-language research related to Sanskrit. The dataset is released under the CC BY-SA 4.0 license.
数据集概述:telugu-script-sanskrit-text
基本信息
- 数据集地址:https://huggingface.co/datasets/harsha-desaraju/telugu-script-sanskrit-text
- 许可证:CC-BY-SA-4.0
数据内容
该数据集包含以泰卢固文转写的梵语文本。原始梵语文本来源于Kaggle上的一个数据集(Sanskrit Text Corpus),并通过aksharamukha工具转写为泰卢固文字。
数据特征
- 特征字段:仅包含一个字段
text,数据类型为字符串(string)
数据规模
- 数据集大小:623,310,076 字节(约594.5 MB)
- 下载大小:233,647,919 字节(约222.8 MB)
数据划分
仅包含一个划分:
| 划分 | 样本数量 | 字节数 |
|---|---|---|
| train | 26,200 | 623,310,076 |
配置
- 默认配置名:default
- 数据文件路径:
data/train-*





