遇见数据集

Replication Code and Data for 'Chronologically Consistent Large Language Models'

收藏
Mendeley Data2026-09-09 收录
官方服务:

资源简介:

This dataset contains the data, code, and documentation for replicating the results in “Chronologically Consistent Large Language Models,” forthcoming in the Journal of Financial Economics. The package reproduces all tables and figures from the included derived series and provides the code for the underlying asset-pricing and language-model pipelines. Dow Jones Newswires and CRSP are proprietary and therefore are not redistributed; synthetic stand-ins with identical schemas are provided so that the data-processing and asset-pricing pipeline can be run end-to-end without access to the proprietary data. The package also documents the public pretraining corpora and released ChronoBERT and ChronoGPT model checkpoints. See README.md for data provenance, computational requirements, and detailed reproduction instructions.

创建时间:
2026-09-01
二维码
社区交流群
二维码
科研交流群
商业服务