登录后查看消息通知
搜索
常见问题
消息
登录
首页
/
数据集
/
True values of μ and q for each arm in simulations.
True values of μ and q for each arm in simulations.
收藏
Figshare
2023-07-12 更新
2026-04-28 收录
多臂老虎机模拟
强化学习基准
数据链接:
https://figshare.com/articles/dataset/True_values_of_i_i_and_i_q_i_for_each_arm_in_simulations_/23669662
数据链接
链接失效反馈
官方服务:
问题咨询
购买咨询
在线客服
NEW
资源简介:
True values of μ and q for each arm in simulations.
应用场景:
创建时间:
2023-07-12
相关数据集
NetHack Learning Dataset(NLD)
强化学习基准
顺序决策任务
最近在代理开发方面的突破,以解决具有挑战性的顺序决策问题,例如Go [50],StarCraft [58] 或DOTA [3],都依赖于模拟环境和大规模数据集。但是,由于开源数据集的稀缺性以及与它们一起工作的高昂计算成本,阻碍了这项研究的进展。在这里,我们介绍了NetHack学习数据集 (NLD),这是来自流行的NetHack游戏的轨迹的大型且高度可扩展的数据集,这对于当前的方法来说是极具挑战性的
OpenXLab
7
0
Parameters to the dynamics of the pole-cart system.
倒立摆控制
强化学习基准
Parameters to the dynamics of the pole-cart system.
Figshare
2015-12-02 更新
5
0
penfever/rl__48GPU_shaped_32b__swe_rebench_patched_oracle__Qwen3-32B
强化学习基准
代码生成与修复
--- dataset_info: features: - name: conversations list: - name: content dtype: string - name: role dtype: string - name: agent dtype: string - name: model dtype
Hugging Face
2026-03-22 更新
3
0
DCAgent2/terminal_bench_2_g1_weighted_100k_32b_cont_20260428_002638
智能体终端操作
强化学习基准
--- dataset_info: features: - name: conversations list: - name: content dtype: string - name: role dtype: string - name: agent dtype: string - name: model dtype
Hugging Face
2026-04-28 更新
5
0
AutumnBench Public Benchmark
模型评估
强化学习基准
AutumnBench Dataset Release This dataset accompanies the paper “Benchmarking World-Model Learning with Environment-Level Queries.” It contains the benchmark release for AutumnBench, including task p
Zenodo
2026-04-10 更新
3
0
© 2023-2026 上海数据发展科技有限责任公司 版权所有
沪ICP备17003045号-15
沪公网安备31010402336585号
热门搜索
社区交流群
科研交流群
商业服务
数据资源
寻源服务
数据采集
标注服务
数据产品
代理销售
数据领域
凭证登记
数据产品
介绍推广