MeghanaKap/english_dravidian_mix_sentences
收藏资源简介:
该数据集是一个多语言文本数据集,包含英语和四种印度语言(卡纳达语kn、马拉雅拉姆语ml、泰米尔语ta、泰卢固语te)的平行文本。每条数据包含以下字段:唯一标识符id、语言代码language、英语文本english、混合本地脚本文本native_script_codemixed、完整本地脚本文本full_native_script、以及罗马化非正式文本romanized_casual。数据集按语言分为四个子集,每个子集约有10万条数据,总数据量超过38万条,适用于多语言自然语言处理任务,如机器翻译、代码转换分析等。
This dataset is a multilingual text dataset containing parallel texts in English and four Indian languages (Kannada kn, Malayalam ml, Tamil ta, Telugu te). Each entry includes the following fields: unique identifier id, language code language, English text english, mixed native script text native_script_codemixed, full native script text full_native_script, and romanized casual text romanized_casual. The dataset is divided into four subsets by language, each with approximately 100,000 entries, totaling over 380,000 entries, suitable for multilingual natural language processing tasks such as machine translation and code-switching analysis.



