mahiyama/amagasaki-qna
收藏资源简介:
这是一个基于日本兵库县尼崎市市民FAQ语料库(1,786条)构建的日语问答检索学习数据集。数据集通过LLM合成查询和Hard Negative Mining技术构建,主要用于corpus特化的微调验证。数据集包含多种配置(pairs、triplets、n-tuples),每种配置有不同的查询和正负样本组合。数据集的构建过程详细描述了从生成合成查询到最终上传的完整流程,并提供了数据集的统计信息和限制。
This is a Japanese QA retrieval learning dataset built on a corpus of citizen FAQs (1,786 items) from Amagasaki City, Hyogo Prefecture. The dataset is constructed using LLM synthetic queries and Hard Negative Mining techniques, primarily for corpus-specialized fine-tuning validation. The dataset includes various configurations (pairs, triplets, n-tuples), each with different combinations of queries and positive/negative samples. The construction process of the dataset is described in detail, from generating synthetic queries to the final upload, and includes statistical information and limitations of the dataset.




