crypto-education-en-golden-set
收藏资源简介:
Crypto Education Golden Set 是一个用于评估检索增强生成(RAG)系统在加密货币和区块链教育内容上表现的黄金标准数据集。该数据集包含497个问答对,覆盖了249个文档(总语料库包含3,487个文档)。数据语言为英语,旨在评估检索器的召回率/精确度、答案的忠实度以及端到端RAG系统的质量。数据集包含以下字段:问题(question)、参考答案(answer)、源文档URL(source_url)、源文档标题(source_title)、主题类别(topic)和问题类型(question_type)。问题类型分布包括事实性(44.9%)、程序性(15.3%)、关键词搜索(13.5%)、非正式(13.1%)、比较(9.9%)和拼写错误(3.4%)等,以模拟真实用户查询。主题分布涵盖DeFi/NFT/Web3(37.8%)、钱包/安全(17.5%)、挖矿/质押/共识(10.9%)等。数据集采用分层抽样方法生成,并通过LLM(Claude)基于文档内容生成问答对,确保语言多样性和内容相关性。
Crypto Education Golden Set is a gold-standard dataset for evaluating the performance of retrieval-augmented generation (RAG) systems on cryptocurrency and blockchain educational content. This dataset comprises 497 question-answer pairs spanning 249 documents, while the total corpus consists of 3,487 documents. All data is in English, and it is designed to assess the recall and precision of the retriever, the faithfulness of generated answers, and the overall quality of end-to-end RAG systems. The dataset includes the following fields: question, reference answer, source document URL, source document title, topic category, and question type. The distribution of question types includes factual (44.9%), procedural (15.3%), keyword search (13.5%), informal (13.1%), comparative (9.9%), and typo (3.4%), among others, to simulate real-world user queries. The topic distribution covers DeFi/NFT/Web3 (37.8%), wallet/security (17.5%), mining/staking/consensus (10.9%), and other thematic categories. The dataset was constructed using stratified sampling, and the question-answer pairs were generated by LLM (Claude) based on the content of the source documents to ensure linguistic diversity and content relevance.
Crypto Education Golden Set 数据集概述
数据集基本信息
- 数据集名称: Crypto Education Golden Set
- 主要用途: 用于在加密货币和区块链教育内容上对检索增强生成(RAG)系统进行基准测试的金标准评估数据集。
- 语言: 英语
- 许可证: MIT
- 数据规模: 小于1K(n<1K)
- 任务类别: 问答、文本检索
数据内容
- 问答对总数: 497
- 覆盖的文档数: 249(源自总语料库的3,487个文档)
- 数据列:
question: 用户提出的英文问题。answer: 参考答案(2-3句话)。source_url: 语料库中源文档的URL。source_title: 源文档的标题。topic: 主题类别。question_type: 问题的语言风格类型。
问题类型分布
数据集旨在模拟真实的用户查询,分布如下:
- factual (事实性): 223个,占44.9%。例如:“What is X?”, “How does X work?”
- procedural (程序性): 76个,占15.3%。例如:“How do I X?”, “What steps are needed?”
- keyword (关键词): 67个,占13.5%。例如:搜索风格的“bitcoin mining energy”。
- informal (非正式): 65个,占13.1%。例如:口语化的“is defi safe to use”。
- comparison (比较): 49个,占9.9%。例如:“Whats the difference between X and Y?”
- typo (拼写错误): 17个,占3.4%。例如:“whats the differnce betwen...”
主题分布
- defi_nft_web3: 188个,占37.8%
- wallets_security: 87个,占17.5%
- mining_staking_consensus: 54个,占10.9%
- tokens_stablecoins: 39个,占7.8%
- exchanges_trading: 36个,占7.2%
- blockchain_projects: 35个,占7.0%
- core_concepts: 26个,占5.2%
- smart_contracts: 25个,占5.0%
- taxes_regulation: 7个,占1.4%
生成方法
- 分层抽样: 按源分布比例抽样249份文档,并根据字数(短/中/长)进行分层。
- 大语言模型生成: 使用Claude为每份抽样文档生成2个基于文档内容的问答对。
- 语言多样性: 分配不同问题类型以模拟真实用户行为。
- 去重: 移除重复的问题。
评估目的
用于评估检索器的召回率/精确度、答案忠实度以及端到端RAG系统的质量。
相关数据集
- 语料库:
kskada/crypto-education-en-corpus— 包含3,487个教育文档。 - 语料库链接: https://huggingface.co/datasets/kskada/crypto-education-en-corpus
引用格式
@dataset{konovalov2026crypto_golden, title={Crypto Education Golden Set for RAG Evaluation}, author={Konovalov, Kirill}, year={2026}, publisher={HuggingFace}, url={https://huggingface.co/datasets/kskada/crypto-education-en-golden-set} }




