ECL
收藏资源简介:
ECL数据集是由根特大学创建的一个多模态数据集,包含来自企业10K文件的文本和数值数据,以及相关的二元破产标签。该数据集由170,139份10K文件组成,这些文件来自18,582家不同的公司,平均每家公司有9.16年的数据。数据集的创建过程涉及从EDGAR、CompuStat和LoPucki Bankruptcy Research Database三个现有数据源中收集数据,并通过特定的标签策略进行标记。ECL数据集主要用于破产预测研究,旨在通过分析公司的财务和业务状况,预测其未来一年的破产风险。
The ECL dataset is a multimodal dataset created by Ghent University, which contains textual and numerical data from corporate 10-K filings, along with corresponding binary bankruptcy labels. It consists of 170,139 10-K filings sourced from 18,582 distinct companies, with an average of 9.16 years of data per company. The dataset was constructed by collecting data from three existing data sources: EDGAR, CompuStat, and the LoPucki Bankruptcy Research Database, followed by labeling via a dedicated labeling strategy. The ECL dataset is primarily intended for bankruptcy prediction research, aiming to forecast the one-year-ahead bankruptcy risk of firms by analyzing their financial and operational conditions.

- 1From Numbers to Words: Multi-Modal Bankruptcy Prediction Using the ECL Dataset根特大学 · 2024年



