TAC (Travel Agent Compassion) Benchmark
收藏资源简介:
TAC(旅行代理同情心)基准数据集是由同情对齐机器学习、感知未来等多个机构联合创建的首个用于评估AI智能体在代理行为中是否避免涉及动物剥削的基准。该数据集包含12个手工编写的旅行预订场景,覆盖六类动物剥削行为,并通过数据增强扩展至48个样本以控制价格、评分和位置等混淆变量,旨在模拟真实旅行代理决策环境。其创建过程基于动物福利分类框架,通过静态数据库和工具调用接口构建,每个场景设计包含至少一个剥削性选项和一个伦理替代选项。该数据集主要应用于前沿AI模型评估领域,旨在揭示AI智能体在隐式动物福利推理方面的行为差距,并为AI治理和系统性风险框架提供实证依据。
The TAC (Travel Agent Compassion) benchmark dataset is the first joint benchmark developed by multiple institutions including Compassion-Aligned Machine Learning, Perceive Future, and other institutions, designed to evaluate whether AI Agents avoid animal exploitation in their agentic behaviors. This dataset contains 12 manually crafted travel booking scenarios covering six categories of animal exploitation behaviors, and is expanded to 48 samples via data augmentation to control confounding variables such as price, rating and location, aiming to simulate real travel agent decision-making environments. Grounded in the animal welfare classification framework, the dataset is constructed using static databases and tool invocation interfaces, with each scenario designed to include at least one exploitative option and one ethical alternative. This dataset is mainly applied in the field of cutting-edge AI model evaluation, with the objectives of revealing the behavioral gaps of AI Agents in implicit animal welfare reasoning and providing empirical evidence for AI governance and systemic risk frameworks.




