遇见数据集

mr233/TokenHD-training-data

收藏
Hugging Face2026-05-14 更新2026-05-31 收录
官方服务:

资源简介:

TokenHD训练数据是一个用于训练TokenHD检测器的token级幻觉标注数据集。每个条目包含由两个批评者模型(gpt-4.1和o4-mini)通过自适应集成生成的软token级标签,这些数据已经过最终处理,可直接用于训练脚本。数据集涵盖数学和代码两个领域:数学推理部分包含82,191个样本,来自math_train、big_math、nv_ace和gemini_math_train;代码推理部分包含41,880个样本,来自gpt-4o-mini和gemini-2.0-flash。数据模式包括问题文本、模型原始回答、正确性标签、token ID列表、每个token的软幻觉分数(0到1之间)以及领域标识。

TokenHD Training Data is a token-level hallucination annotation dataset used to train TokenHD detectors. Each entry contains soft token-level labels produced by an adaptive ensemble of two critic models (gpt-4.1 and o4-mini). This is the final processed data — ready to use directly with the training script. The dataset covers two domains: math reasoning with 82,191 samples from math_train, big_math, nv_ace, and gemini_math_train, and code reasoning with 41,880 samples from gpt-4o-mini and gemini-2.0-flash. The schema includes problem text, raw model response, correctness label, token IDs list, soft hallucination scores per token (in [0, 1]), and domain identifier.

提供机构:
mr233
二维码
社区交流群
二维码
科研交流群
商业服务