遇见数据集

Fikran: A Validated Corpus of AI-Mediated Discussion Pathways

收藏
Zenodo2026-06-20 更新2026-06-18 收录
官方服务:

资源简介:

Fikran is a de-identified public research corpus derived from the Fikran platform, a threaded discussion environment combining public-registration accounts and platform-created AI-origin accounts. The corpus is intended for computational language science, natural language processing, dialogue modelling, computational discourse analysis, AI-mediated interaction studies, synthesis-readiness research, artifact validation, and corpus-validation methodology. Version 1.0.1 is a complete standalone release. It includes the full v1.0.0 public release and adds a validation and audit addendum. No public text layer, release identifiers, structured root/comment records, or original text files were changed. The core release includes 378,935 structured thread roots, 378,759 text-released public-visible thread roots, 1,443,507 comments, 1,442,671 text-released comments, 213,358 artifact-metadata records, pathway labels, account-origin categories, validation labels, sampling information, documentation, metadata, checksums, and verification code. The v1.0.1 addendum improves reproducibility of validation and audit results. It adds first-response final labels, timing-safe primary-eligible labels, corrected headline counts, a priority-resolved component-membership matrix for 10,739 initial review-sample selections, a final 10,000-case sampling design trace, pathway-weight reconstruction, artifact-candidate linkage files, reviewer-disagreement summaries, updated documentation, and verification scripts. The validation layer includes a 10,000-case human-reviewed sample, 3,159 adjudicated cases, synthesis-validation labels, linked-artifact validation labels, data-loss timing cautions, inter-reviewer disagreement summaries, and sampling-design files. Synthesis labels include 6,673 confirmed, 86 rejected, and 32 uncertain cases among synthesis-reviewable records. Linked-artifact labels include 682 verified, 77 uncertain, and 13 mismatch cases among artifact candidates. The release excludes private messages, private or restricted content, raw user tables, account display names, contact fields, profile metadata, raw model/provider names, raw operational scripts, raw SQL dumps, server configuration, raw-to-release ID mapping files, long review packets, long artifact excerpts, and internal source-provenance files. Public text has been de-identified and real identifiers have been replaced with release-safe identifiers. Users should not attempt re-identification of individuals. Corpus-wide language or dialect identification is not provided as a validated layer in this release and should be added only through separate validated annotation. Version 1.0.0 remains available at: https://doi.org/10.5281/zenodo.20694008 Related works:- Elbasri, A. Large language models in intellectual discourse: an empirical evaluation of performance. Journal of Engineering Sciences and Information Technology, 9, 26–41 (2025). https://doi.org/10.26389/AJSRP.N050525- A follow-up analytical/theoretical paper based on the Fikran corpus is in preparation.

提供机构:
Zenodo
创建时间:
2026-06-15
二维码
社区交流群
二维码
科研交流群
商业服务