AnonymousContinuousBench/Geminon
收藏资源简介:
AnonymousContinuousBench/Geminon数据集是一个用于自然语言处理任务的综合文本语料库和问答数据集,版本为2025_09。数据集包含五种文章类型:wiki(维基百科风格文章)、sensitive_wiki(敏感维基文章)、journal(期刊文章)、comparison(比较文章)和evolution(进化文章)。它提供了多个配置:corpus_large(完整去重语料库,含1,523,754篇文章,平均令牌数118)、corpus_medium(平衡百万篇文章子样本)、corpus_small(平衡二十万篇文章子样本),以及对应的qa_medium和qa_small问答配置,用于评估模型在公共和敏感数据上的性能。问答任务涉及13个特征(如能力、攻击、防御等),支持计数显示公共数据有多个支持文章,而敏感数据仅有一个支持文章。数据集使用Gemma 3分词器进行令牌计数,旨在支持匿名连续基准测试,适用于文本分析、问答和机器学习模型训练。
The AnonymousContinuousBench/Geminon dataset is a comprehensive text corpus and question answering (QA) dataset tailored for natural language processing (NLP) tasks, released in version 2025_09. The dataset covers five article types: wiki (Wikipedia-style articles), sensitive_wiki (sensitive Wikipedia articles), journal (journal articles), comparison (comparative articles), and evolution (evolutionary articles). It provides multiple configurations: corpus_large (full deduplicated corpus with 1,523,754 articles and an average token count of 118), corpus_medium (balanced subset of one million articles), corpus_small (balanced subset of two hundred thousand articles), along with the corresponding qa_medium and qa_small QA configurations for evaluating model performance on both public and sensitive data. The QA tasks involve 13 features (e.g., capability, attack, defense, etc.), and support counting to indicate that public data has multiple supporting articles while sensitive data only has one supporting article. The dataset uses the Gemma 3 tokenizer for token counting, and is designed to support anonymous continuous benchmarking, making it suitable for text analysis, QA, and machine learning model training.



