遇见数据集

HiPuSi-MT44K: Multi-Engine Machine Translation Outputs for Hindi, Punjabi and Devanagari-Script Sindhi in Six Translation Directions

收藏
Mendeley Data2026-09-08 收录
官方服务:

资源简介:

HiPuSi-MT44K is a multi-engine machine translation output dataset for Hindi, Punjabi, and Devanagari-script Sindhi. It contains 3,000 source sentences and 44,000 machine-generated translation outputs across six translation directions. The resource includes outputs from web-based MT services, conversational large language models, pretrained multilingual neural machine translation systems, and in-house EBMT, SMT, and NMT systems. The source collection is general-domain material compiled during a MeitY, Government of India-sponsored project and manually vetted for grammatical structure and spelling by five native-speaker language experts holding Master's-level or higher qualifications. The dataset is intended for comparative MT evaluation, quality estimation, error analysis, system agreement studies, and human evaluation research.

HiPuSi-MT44K是一款面向印地语、旁遮普语与天城文信德语的多引擎机器翻译输出数据集。该数据集包含3000条源语句,以及覆盖6个翻译方向的44000条机器生成翻译结果。本资源涵盖基于网页的机器翻译(Machine Translation,MT)服务、对话式大语言模型(Large Language Model,LLM)、预训练多语言神经机器翻译系统,以及自研基于实例的机器翻译(Example-Based Machine Translation,EBMT)、统计机器翻译(Statistical Machine Translation,SMT)与神经机器翻译(Neural Machine Translation,NMT)系统生成的翻译输出。该源语语料为通用领域素材,由印度电子与信息技术部(MeitY)赞助的项目编纂而成,并由5名母语适配对应语种、持有硕士及以上学历的语言专家针对语法结构与拼写进行了人工审核。本数据集可用于机器翻译对比评测、质量估计、错误分析、系统一致性研究以及人工评估相关研究。

创建时间:
2026-08-12
二维码
社区交流群
二维码
科研交流群
商业服务