ghananlpcommunity/ghanaian-corpus-generate
收藏资源简介:
这是一个基于534,294个加纳语料块使用Qwen2.5-0.5B-Instruct模型通过vLLM生成的合成内容编写提示数据集。每个数据行包含:source_type(源文档类别,如议会、新闻、学术等)、source(源文档名称或标识符)、page_range(源文档中的页码范围)、text(原始语料块文本,至少4个句子,每个句子包含6个以上唯一字母单词)和generated_question(合成提示,代表内容编写者输入AI助手以生成类似给定语料块文本的指令)。数据集旨在用于微调大型语言模型,以生成基于加纳语境的文本,其中text列提供目标内容,generated_question列提供指令式提示,适用于指令调优或监督微调。
Synthetic content-writer prompts generated from 534,294 Ghanaian corpus chunks using Qwen2.5-0.5B-Instruct via vLLM. Each row contains: source_type (category of the source document, e.g., parliament, news, academic), source (name or identifier of the source document), page_range (page range within the source), text (the original corpus chunk text with a minimum of 4 sentences, each with 6+ unique alphabetic words), and generated_question (a synthetic prompt representing what a content writer would type into an AI assistant to produce text similar to the given chunk). The dataset is intended for fine-tuning LLMs to generate text grounded in Ghanaian context, with the text column providing target content and generated_question providing instruction-style prompts, suitable for instruction-tuning or supervised fine-tuning.




