10k-fraud-detection
收藏资源简介:
Description: We provide (to the best of our knowledge) the first publicly available data set of (specific sections of) 10-K reports alongside labels indicating fraudulent behavior. To enable other researchers to replicate or extend our work, we provide transparent descriptions of the data scraping and labeling processes, release our source code, and make the data available. In our paper "Do Companies Reveal Their Own Fraud? - A Novel Data Set for Fraud Detection Based on 10-K Reports", we provide baseline results for a given data split motivated by specific temporal characteristics of fraud. Scientific Work: This data set accompanies the following research article:Amin, Moustafa & Aßenmacher, Matthias (2025) Do Companies Reveal Their Own Fraud? - A Novel Data Set for Fraud Detection Based on 10-K Reports". In Proceedings of the 10th Workshop on Financial Technology and Natural Language Processing (FinNLP), Suzhou, China (Hybrid). Association for Computational Linguistics.



