Persian Twitter Data Governance Sample: Quality- and Bias-Filtered Corpus
收藏资源简介:
A public, anonymized sample (8,000 records) from the final output of a data governance pipeline for Persian-language Twitter data, after noise cleaning, quality filtering (Q(t)), and lexical bias mitigation (B(t)). This sample accompanies the paper "Evaluating Data Quality and Bias of Social Media in Persian Large Language Models: A Data Governance Approach" (Journal of Computing Sciences and Information Technology, Computer Society of Iran). The sample contains no raw Twitter data and no user identity information -- only governed tweet text, tokens, and the computed Q(t)/B(t) scores. See the included README.md for field descriptions and methodology. The reference pipeline implementation is available at github.com/MrMaper/persian-twitter-data-governance.



