遇见数据集

mamed0v/alpaca-turkmen

收藏
Hugging Face2024-07-06 更新2024-07-22 收录
官方服务:

资源简介:

Turkmen Alpaca数据集是原始Alpaca数据集的土库曼语翻译版本,旨在为土库曼语社区提供指令跟随数据集。数据集包含约52,000个样本,支持多种任务,如开放式生成和问答。翻译通过Google Translate完成,数据集以JSONL格式提供,每个JSON对象包含英文和土库曼语的指令、输入、输出和完整文本。

The Turkmen Alpaca Dataset is a Turkmen translation of the original Alpaca dataset. It contains approximately 52,000 instruction-following samples, including various tasks such as open-ended generation and question-answering. The dataset is provided in JSONL format, with each entry containing both English and Turkmen versions of the instruction, input, output, and full text. The translation was done using Google Translate, and the file acknowledges potential inaccuracies due to this method. The dataset aims to extend accessibility to the Turkmen language community.

提供机构:
mamed0v
二维码
社区交流群
二维码
科研交流群
商业服务