StakeBench
收藏资源简介:
StakeBench是由布里斯托大学研究团队构建的一个专注于市场承诺语言理解的双平台预测市场评估框架。该数据集整合了来自Polymarket和Manifold两个预测市场的560,876条评论,覆盖2,261个已结算市场,包含18个主题与平台的组合,数据规模庞大且具有时间序列特征。数据创建过程通过公开API采集评论线程,并基于交易历史重建用户持仓记录,将语言表达与可验证的持仓行为、后续交易动作及市场赔率轨迹进行关联,形成无人工标注的监督信号。该数据集旨在解决金融自然语言处理中语言理解与市场实际承诺脱节的问题,通过四个渐进式诊断任务评估模型对市场承诺信号、持仓方向、未来行动及集体赔率变化的识别能力,为金融文本的战略性分析提供实证基础。
StakeBench is a dual-platform prediction market evaluation framework focused on market commitment language understanding, developed by the research team at the University of Bristol. This dataset integrates 560,876 comments from two prediction markets, Polymarket and Manifold, covering 2,261 settled markets and spanning 18 topic-platform combinations, exhibiting large-scale volume and inherent time-series characteristics. The data creation process collects comment threads via public APIs, reconstructs users' position records based on transaction histories, and associates linguistic expressions with verifiable position behaviors, subsequent trading actions and market odds trajectories, thereby generating manually-unannotated supervision signals. This dataset aims to address the disconnect between language understanding and actual market commitments in financial natural language processing. It evaluates models' ability to recognize market commitment signals, position directions, future actions and collective odds changes through four progressive diagnostic tasks, providing an empirical foundation for strategic analysis of financial texts.




