LMAct
收藏资源简介:
LMAct是一个用于长上下文模仿学习的基准数据集,由谷歌深度思维创建。该数据集包含六个决策任务,涵盖了从零样本到多样本的学习场景,旨在评估大型多模态模型在长上下文环境中的决策能力。数据集内容包括井字棋、国际象棋、Atari游戏、网格世界导航、填字游戏和模拟猎豹控制等任务。数据集的创建过程涉及专家策略的生成和多种状态表示的优化。LMAct的应用领域主要集中在测试和提升大型模型在复杂决策任务中的表现,旨在解决模型在长上下文环境中进行有效决策的问题。
LMAct is a benchmark dataset for long-context imitation learning, created by Google DeepMind. This dataset comprises six decision-making tasks covering learning scenarios ranging from zero-shot to few-shot settings, aiming to evaluate the decision-making capabilities of large multimodal models in long-context environments. The dataset includes tasks such as tic-tac-toe, chess, Atari games, grid world navigation, crossword puzzles, and simulated cheetah control. The development of LMAct involves the generation of expert policies and the optimization of multiple state representations. The primary application scenarios of LMAct focus on testing and enhancing the performance of large models in complex decision-making tasks, with the goal of addressing the challenges of effective decision-making by models in long-context environments.

- 1LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations谷歌深度思维 · 2024年



