AdamJovine/LISTEN-benchmark
收藏资源简介:
LISTEN多目标选择基准是一个用于评估大型语言模型(LLM)在用户偏好引导下从具有多个数值和分类属性的大规模候选集中选择最佳选项能力的基准数据集。该数据集包含四个场景:考试安排(4,938个候选,基于冲突、聚类和延迟惩罚评分)、芝加哥到纽约航班(903个候选,包含结构化中转/时长/价格列)、伊萨卡到雷斯顿航班(216个候选,涉及多种承运商/中转组合)和耳机产品(77个候选,包含技术规格如驱动单元大小、电池、主动降噪和评论)。每个场景都提供了候选集、文档化的偏好陈述和人工策划的可接受“获胜”行集合。此外,数据集还包括五个独立的人类重新排名数据(rerank_h1至rerank_h5),用于在论文中衡量LISTEN与人类评分者之间的一致性。数据集主要用于通过归一化平均排名(NAR)等指标评估LLM在多目标选择任务中的性能,并支持学术研究。
The LISTEN Multi-Objective Selection Benchmark is a benchmark for evaluating how well a large language model (LLM) can elicit a users preferences and select a top option from a large candidate set with many numeric and categorical attributes. The benchmark contains four scenarios: Exam Scheduling (4,938 candidates, scored on conflict, clustering, and lateness penalties), Flights CHI→NYC (903 candidates, with structured layover/duration/price columns), Flights Ithaca→Reston (216 candidates, featuring varied carrier/stop combinations), and Headphones (77 candidates, with technical specs such as driver size, battery, ANC, and reviews). Each scenario includes a candidate set, a documented preference statement, and a small human-curated set of acceptable winning rows (human_sol). It also includes five independent human re-rankings (rerank_h1 to rerank_h5) used in the paper to measure agreement between LISTEN and human raters. The dataset is primarily used to assess LLM performance in multi-objective selection tasks through metrics like Normalized Average Rank (NAR) and supports academic research.





