VERIFASTSCORE synthetic datasets
收藏资源简介:
VERIFASTSCORE数据集是用于训练和评估VERIFASTSCORE模型的人工合成数据集。该数据集由大约9600个提示-回复对组成,其中每个回复都被分解为可验证的声明,并使用VERISCORE管道根据检索到的证据进行验证。数据集旨在解决长文本事实性评估的效率和实用性问题,通过将声明分解和验证过程整合到一个模型调用中,实现了显著的性能提升。VERIFASTSCORE模型在保持高事实精度和召回率的同时,减少了推理时间,提高了事实性评估的效率和可解释性。该数据集的发布旨在促进未来事实性研究,并支持在大型评估和训练场景中的应用。
The VERIFASTSCORE dataset is a synthetic dataset designed for training and evaluating the VERIFASTSCORE model. It consists of approximately 9,600 prompt-response pairs, where each response is decomposed into verifiable claims and validated using the VERISCORE pipeline against retrieved evidence. This dataset aims to address the efficiency and practicality issues of long-text factual evaluation, and achieves significant performance improvements by integrating the claim decomposition and validation processes into a single model invocation. The VERIFASTSCORE model reduces inference time while maintaining high factual accuracy and recall, thereby enhancing the efficiency and interpretability of factual evaluation. The release of this dataset aims to facilitate future factual research and support applications in large-scale evaluation and training scenarios.

- 1VeriFastScore: Speeding up long-form factuality evaluation马里兰大学, Lambda Labs · 2025年



