ADHI-Vel/sp500-edgar-10k
收藏资源简介:
--- license: mit tags: - nlp pretty_name: SP500 EDGAR 10-K Filings --- # Dataset Card for SP500-EDGAR-10K ## Dataset Description - **Homepage:** - **Repository:** - **Paper:** - **Leaderboard:** - **Point of Contact:** ### Dataset Summary This dataset contains the annual reports for all SP500 historical constituents from 2010-2022 from SEC EDGAR Form 10-K filings. It also contains n-day future returns of each firm's stock price from each filing date. ## Dataset Structure ### Data Fields [More Information Needed] ### Data Splits [More Information Needed] ## Dataset Creation ### Source Data #### Initial Data Collection and Normalization 10-K filings data was collected and processed using `edgar-crawler` available <a href='https://github.com/nlpaueb/edgar-crawler'>here.</a> Return data was computed manually from other price data sources. ### Annotations #### Annotation process N/A #### Who are the annotators? N/A ### Personal and Sensitive Information N/A ## Considerations for Using the Data ### Social Impact of Dataset N/A ### Discussion of Biases The firms in the dataset are constructed from historical SP500 membership data, removing survival biases. ### Other Known Limitations N/A ### Licensing Information MIT
license: MIT协议 tags: - 自然语言处理(NLP) pretty_name: 标普500 EDGAR 10-K文件 # SP500-EDGAR-10K 数据集卡片 ## 数据集概述 - **主页:** - **仓库:** - **论文:** - **排行榜:** - **联系方式:** ### 数据集概况 本数据集涵盖了2010年至2022年间所有标普500(S&P 500)历史成分股公司的美国证券交易委员会(SEC)EDGAR平台10-K表格(Form 10-K)年度申报文件,同时包含每家公司自各申报日期起的N日未来股价收益率数据。 ## 数据集结构 ### 数据字段 [需补充更多信息] ### 数据划分 [需补充更多信息] ## 数据集构建 ### 源数据 #### 初始数据收集与标准化 10-K表格申报文件数据通过`edgar-crawler`工具收集并处理,该工具可在此<a href='https://github.com/nlpaueb/edgar-crawler'>获取</a>。股价收益率数据则通过其他股价数据源手动计算得到。 ### 注释信息 #### 注释流程 无适用 #### 注释人员 无适用 ### 个人与敏感信息 无适用 ## 数据使用注意事项 ### 数据集的社会影响 无适用 ### 偏差说明 本数据集基于标普500历史成分股数据构建,已剔除生存偏差(survival bias)。 ### 其他已知局限性 无适用 ### 许可信息 MIT协议



