遇见数据集

Replication Code and Data: Emoji-Based Stylometric Features for Pairwise Authorship Verification in Social Media

收藏
Zenodo2026-02-26 更新2026-05-26 收录
官方服务:

资源简介:

<p>This archive contains the complete, standalone replication code for the paper:</p> <p><strong>Etli, Y., Özdil, A. &amp; Aşırdizer, M. (2026). Emoji-Based Stylometric Features for Pairwise Authorship Verification in Social Media: A Novel Approach for Digital Forensics. <em>Forensic Science International: Digital Investigation.</em></strong></p> <h4>What this study does</h4><p>We propose <strong>20 novel emoji-based stylometric features</strong> (covering distributional, vocabulary, behavioral, sequential, positional, and temporal dimensions of emoji usage) and evaluate them for pairwise authorship verification on social media text. These emoji features are compared against 27 classic stylometric features and a 47-feature hybrid model using a Random Forest classifier across 49 experimental configurations (7 profile sizes × 7 test sizes).</p> <h4>What this archive contains</h4><ul> <li>9 standalone Python scripts (00–08) covering the full pipeline: data preparation, descriptive statistics, emoji-only model, classic baseline, hybrid model, statistical significance tests (McNemar + paired t-test), ROC/confusion matrix visualization, calibration metrics, and author-disjoint validation</li> <li>Machine-readable feature definitions (JSON) for all 47 features</li> <li>Centralized configuration file with all experiment parameters</li></ul> <h4>Dataset</h4><p>Experiments were conducted on <strong>4,924 unique Reddit authors</strong> across 17 subreddits, filtered from the publicly available <a href="https://huggingface.co/datasets/HuggingFaceGECLM/REDDIT_comments">HuggingFace REDDIT_comments</a> corpus (accessed 15 January 2024). Raw comment data is not redistributed due to Reddit's Terms of Service; a reproduction script is included.</p> <h4>Key results</h4><ul> <li>Emoji-only model: 78.4% mean accuracy (97.4% at P500/T500) with only 20 features</li> <li>Classic model: 87.1% mean accuracy with 27 features</li> <li>Hybrid model: 87.7% mean accuracy (up to 100% at P500/T500) with 47 features</li></ul> <h4>Reproducibility</h4><p>All scripts are fully standalone with argparse CLI interfaces. All random operations use a fixed seed (42). Random Forest hyperparameters are fixed across all experiments (no tuning). Scripts 02–04 support checkpoint-based resumption for long-running experiments.</p>

提供机构:
Zenodo
创建时间:
2026-02-26
二维码
社区交流群
二维码
科研交流群
商业服务