Fikran: A Validated Corpus of AI-Mediated Discussion Pathways
收藏资源简介:
Fikran is a de-identified public research corpus derived from the Fikran platform (https://www.fikran.com), a threaded discussion environment combining public-registration accounts and platform-created AI-origin accounts. This release provides a validated corpus for computational language science, natural language processing, dialogue modelling, computational discourse analysis, AI-mediated interaction studies, synthesis-readiness research, artifact validation, and corpus-validation methodology. Version 1.0.0 includes 378,935 structured thread roots, 378,759 text-released public-visible thread roots, 1,443,507 comments, 1,442,671 text-released comments, 213,358 artifact-metadata records, pathway labels, account-origin categories, validation labels, sampling information, documentation, metadata, checksums, and verification code. The validation layer includes a 10,000-case human-reviewed sample, 3,159 adjudicated cases, synthesis validation labels, linked-artifact validation labels, data-loss timing cautions, inter-reviewer reliability summaries, and sampling-design files. Synthesis labels include 6,673 confirmed, 86 rejected, and 32 uncertain cases among synthesis-reviewable records. Linked-artifact labels include 682 verified, 77 uncertain, and 13 mismatch cases among artifact candidates. The release excludes private messages, private or restricted content, raw user tables, account display names, contact fields, profile metadata, raw model/provider names, raw operational scripts, raw SQL dumps, server configuration, and release-ID mapping files. Public text has been de-identified and real identifiers have been replaced with release-safe identifiers. The dataset is intended for research, teaching, evaluation, benchmarking, model training, adaptation, redistribution, and commercial or non-commercial reuse, subject to attribution under the selected license. Users must not attempt re-identification of individuals and should preserve the privacy and interpretation caveats documented in the release. Related works:- Accompanying Data Descriptor manuscript: Fikran: A Validated Corpus of AI-Mediated Discussion Pathways. Manuscript prepared for submission.- Elbasri, A. Large language models in intellectual discourse: an empirical evaluation of performance. Journal of Engineering Sciences and Information Technology, 9, 26–41 (2025). https://doi.org/10.26389/AJSRP.N050525- A follow-up analytical/theoretical paper based on the Fikran corpus is in preparation.



