遇见数据集

driftlense1/driftlense

收藏
Hugging Face2026-05-25 更新2026-05-31 收录
官方服务:

资源简介:

DriftLens: Questions数据集是DriftLens框架的组成部分,旨在评估大型语言模型在不可验证、开放式问题上的推理鲁棒性。这些问题没有单一客观标准答案,仅依赖于合理的推理链。数据集通过从职业、伦理、金融、法律、医疗和日常困境等八个公共来源收集候选问题,并经过两阶段筛选(包括LLM评分和人工标注)构建而成。它包含两个版本:filtered(严格筛选,422行,标注者一致同意)和raw(宽松筛选,1061行,标注者部分同意)。每个问题包含问题文本、标签、理由和领域标签等字段。数据集专门用于推理鲁棒性评估,不提供参考答案,适用于生成基线响应和扰动响应以计算漂移分数。

This dataset accompanies the DriftLens framework for measuring reasoning robustness in large language models on unverifiable, open-ended questions — prompts where there is no single objective ground truth, only a plausible chain of reasoning. It is constructed from eight public sources across domains like career, ethics, finance, legal, medical, and daily dilemmas, filtered through a two-stage process involving LLM rating and human annotation. Two releases are available: filtered (422 rows with unanimous human agreement) and raw (1,061 rows with relaxed agreement). Each question includes fields such as question text, label, reason, and domains. The dataset is intended for reasoning-robustness evaluation pipelines, without reference answers, and supports generating baseline and perturbed responses to compute drift scores.

提供机构:
driftlense1
二维码
社区交流群
二维码
科研交流群
商业服务