遇见数据集

CRDI Corporate Fraud Multimodal Dataset (N=276)

收藏
Zenodo2026-07-17 更新2026-08-01 收录
官方服务:

资源简介:

Calibrated semi-synthetic company-profile dataset (N=276; 50 fraud, 226 legitimate) used to train and evaluate the Corporate Reality Distortion Index (CRDI) multimodal fraud-risk model, described in the paper "CRDI Multimodal AI Framework for Corporate Fraud Risk Estimation Using Geospatial, Auditory, and Linguistic Signals." Each company profile includes geospatial shell-company risk (from ResNet-18/Places365 scene classification of HQ imagery), vocal-stress biomarkers (jitter, shimmer, pitch variance, pause rate), and a linguistic semantic-evasion score (from Chain-of-Thought LLM analysis of earnings-call transcripts), alongside company identity, sector, size category, and fraud type/label. Provenance: Company identities, sectors, and fraud/legitimate labels are real and drawn from public sources (SEBI/RBI/Enforcement Directorate filings, BSE/NSE listings). The geospatial, auditory, and linguistic feature values are calibrated semi-synthetic: sampled from Gaussian distributions parameterized on statistics reported in forensic-phonetics and deception-detection literature, rather than extracted directly from raw audio/video/text for every company. 40% of fraud cases are modeled as competent liars with low vocal stress and credible headquarters, producing realistic class overlap. See the included README.md for full column definitions. This dataset is a research proof-of-concept and should not be used to make real-world determinations about the fraud status of any named company. Source code: https://github.com/HarshitK2814/Corporate-Fraud-Detection-using-GenAI

提供机构:
Zenodo
创建时间:
2026-07-17
二维码
社区交流群
二维码
科研交流群
商业服务