登录后查看消息通知
搜索
常见问题
消息
登录
首页
/
数据集
/
Cognition performance comparison on MME benchmark.
Cognition performance comparison on MME benchmark.
收藏
Figshare
2025-08-11 更新
2026-04-28 收录
大模型评估
认知能力评估
数据链接:
https://figshare.com/articles/dataset/Cognition_performance_comparison_on_MME_benchmark_/29884275
数据链接
链接失效反馈
官方服务:
问题咨询
购买咨询
在线客服
NEW
资源简介:
Cognition performance comparison on MME benchmark.
应用场景:
创建时间:
2025-08-11
相关数据集
rebase_gemma-4-E4B-it_rg_cognition_ns128_md4_bt0_1_seed73_rg_cognition__v0
生成模型
认知能力评估
该数据集包含一个测试集,共12,800个样本,主要用于评估生成模型或问答系统的性能。数据集包含多个特征字段,如问题文本(question)、生成ID(generation_id)、生成内容(generation)、令牌数量(num_tokens)、奖励分数(reward)、问题索引(question_index)、目标文本(target)、任务类型(task)、价值函数预测(vf_predicti
Hugging Face
2026-05-07 更新
5
0
young and older participants.txt
认知能力评估
年龄相关认知
Raw RT and error data of young and older adults participanting in the experiment. Also full R code.
Figshare
2022-02-21 更新
4
0
Comparison of participants’ understanding and objective numeracy and graph literacy proficiency for each trial arm.
认知能力评估
学生数字素养
Comparison of participants’ understanding and objective numeracy and graph literacy proficiency for each trial arm.
Figshare
2021-07-23 更新
2
0
AlignmentResearch/StrongREJECT-test
大模型评估
有害指令检测
--- dataset_info: features: - name: clf_label dtype: class_label: names: {} - name: proxy_clf_label dtype: class_label: names: {} - name: instructions d
Hugging Face
2024-07-26 更新
4
0
jkminder/model-raising-reflection-end-eval
大模型评估
宪章引导反思生成
这是一个用于评估的保留集,专门针对基于宪章指导的预训练反思,放置在文档末尾(reflection_end)。每个数据行包含一个dolma3网络文档以及配对的第一人称和第三人称反思,这些反思引用了文档实质性涉及的宪章部分([X.Y])。数据集使用冻结的生产管道(Qwen3.5-35B-A3B-FP8模型,提示generator_reflection_v7.md,宪章为ModelRaisingCons
Hugging Face
2026-05-28 更新
3
0
© 2023-2026 上海数据发展科技有限责任公司 版权所有
沪ICP备17003045号-15
沪公网安备31010402336585号
热门搜索
社区交流群
科研交流群
商业服务
数据资源
寻源服务
数据采集
标注服务
数据产品
代理销售
数据领域
凭证登记
数据产品
介绍推广