awesome-language-dataset
收藏资源简介:
这是一个用于微调大型语言模型(LLM)的语言数据集合集,收集了多个公开可用的数据集,涵盖AlpacaGPT3.5Customized、中文-英文翻译数据、GPT4all、pCLUE、ShareGPT_Vicuna_unfiltered、Stanford Alpaca多语言版本(英语、粤语、中文、法语、日语)以及Laion OIG等,主题涉及指令生成、对话、翻译和通用语言任务,以表格形式组织,提供数据量、来源链接和简要描述。
This is a language dataset collection for fine-tuning Large Language Models (LLMs). It aggregates multiple publicly available datasets, including AlpacaGPT3.5Customized, Chinese-English translation datasets, GPT4all, pCLUE, ShareGPT_Vicuna_unfiltered, the multilingual version of Stanford Alpaca (covering English, Cantonese, Mandarin, French, and Japanese), and Laion OIG, among others. Its topics cover instruction generation, dialogue, translation and general language tasks. The datasets are organized in a tabular format, with details including dataset size, source links and brief descriptions provided.




