遇见数据集

macrolens/MacroLens

收藏
Hugging Face2026-05-20 更新2026-05-31 收录
官方服务:

资源简介:

MacroLens是一个用于宏观经济情景下上下文金融推理的基准语料库,涵盖4,416只美国小盘和微盘股票(2021年1月4日至2026年3月31日)。它统一了七个任务:上下文时间序列预测(T1)、公开估值(T2)、财务报表生成(T3)、情景调节收益预测(T4)、私营公司估值(T5)、自然语言描述的生成器评估(T6)以及房地产估值(T7)。每个实例包含一个131个数值/141列的时间点面板,包括价格、46.8百万XBRL会计事实、53个宏观经济序列、文件最近性和衍生比率,以及可选的宏观经济情景对象(1,130个事件,49种类型)和SEC文件与金融新闻上下文。数据严格按时间点对齐,确保所有观察在预测时间戳前公开可用。数据集结构包括每日、每周、每月和房地产数据子集,提供训练和测试分割。数据来源包括SEC EDGAR、FRED、EIA、yfinance、RentCast和 curated宏观经济事件。许可证为CC-BY-4.0。

MacroLens is a benchmarking corpus for contextual financial reasoning under macroeconomic scenarios across 4,416 U.S. small- and micro-cap equities (2021-01-04 — 2026-03-31). It unifies seven tasks over a single point-in-time panel: contextual time-series forecasting (T1), public valuation (T2), financial-statement generation (T3), scenario-conditioned return forecasting (T4), private-company valuation (T5), generator evaluation from natural-language descriptions (T6), and real-estate valuation (T7). Every instance carries a 131-numeric / 141-column point-in-time panel (prices, 46.8M XBRL accounting facts, 53 macroeconomic series, filing recency, derived ratios), an optional macroeconomic scenario object (1,130 events across 49 types), and optional SEC filings and financial-news context. Temporal alignment is strictly point-in-time, ensuring all observations were publicly available by the prediction timestamp. The dataset structure includes daily, weekly, monthly, and real-estate subsets, with train and test splits. Data sources include SEC EDGAR, FRED, EIA, yfinance, RentCast, and curated macroeconomic events. License is CC-BY-4.0.

提供机构:
macrolens
二维码
社区交流群
二维码
科研交流群
商业服务