rajatagarwal457/gdelt-forecast-freeform
收藏资源简介:
GDELT-Forecast Free-form数据集包含924个自由形式的预测问题,这些问题来自GDELT 2.0语料库(2025年8月至2026年4月)中的新闻文章聚类。每个问题都配有原始种子事件文章、检索到的前5篇证据文章(日期严格早于问题创建日期)和经过验证的正确答案。数据集用于训练和评估基于大型语言模型(LLM)的预测模型,特别是在非二元问题上。答案类型包括名称(402个)、数字(370个)、自由形式(140个)和日期(12个)。数据集构建过程与gdelt-forecast-binary类似,但在问题生成阶段使用了不同的答案类型。数据集遵循严格的预测立场,要求在使用时尊重问题开始日期的截止点。由于答案表达的多样性,精确匹配不适用于大多数行,建议使用模糊判断(如GPT-4o)来评分预测结果。
The GDELT-Forecast Free-form dataset contains 924 free-form forecasting questions generated from clusters of news articles in the GDELT 2.0 corpus (Aug 2025 – Apr 2026). Each question is paired with the original seed-event articles, top-5 retrieved evidence articles dated strictly before the question creation date, and a verified ground-truth answer. The dataset is intended for training and evaluating LLM-based forecasting models on non-binary questions. Answer types include name (402), number (370), free_form (140), and date (12). The dataset was built using a similar pipeline as gdelt-forecast-binary but with non-yes/no answer types assigned during question generation. It adheres to a strict forecasting posture, requiring respect for the question_start_date cutoff. Due to high answer phrasing variance, exact-match scoring is not recommended; instead, a fuzzy judge (e.g., GPT-4o) should be used to score predictions.




