Stratified comparison dataset for a source-critical audit of YouTube discourse on BaʿAlawi religious authority in Indonesia
收藏资源简介:
This dataset contains the stratified comparison corpus and validation materials underlying a source-critical (Quellenkritik) audit of a digital humanities study on the religious authority of BaʿAlawi ḥabāʾib in Indonesia. It was constructed as an independent counterfactual sample to test the representativeness of a previously published Zenodo dataset (DOI: 10.5281/zenodo.14275842), not as an estimate of public opinion. The corpus comprises 670 top-level YouTube comments (667 unique comment texts after whitespace normalization) collected from five videos stratified into three narrative categories: pro-ḥabāʾib (300 comments, 2 videos), contra/critical (250 comments, 2 videos), and neutral panel-debate (120 comments, 1 video). Comments were harvested with the youtube-comment-downloader Python utility, restricted to top-level comments by filtering out reply identifiers containing a period. Sentiment was classified with the multilingual XLM-RoBERTa model tabularisai/multilingual-sentiment-analysis using a five-class scheme (Very Negative, Negative, Neutral, Positive, Very Positive) identical to the audited study, with softmax confidence scores reported. A 75-comment manual validation subsample (50 stratified-random plus 25 low-confidence edge cases) was annotated by the first author; inter-rater agreement with the model was Cohen's κ unweighted = 0.086, linear-weighted = 0.144 (95% bootstrap CI [−0.029, 0.303]), quadratic-weighted = 0.205. The workbook includes the following sheets: audit_check (key-figure verification), Sheet1 (the 670-comment corpus with sentiment labels and confidence), channel_mapping (video-to-channel reference), manual_validation_75 (manual annotation), annotator_guide (five-class operational definitions), kappa_summary (confusion matrix and agreement metrics), calibration_inverse (inverse-matrix calibration of the true distribution), thematic_coding_150 (Braun & Clarke thematic coding of 150 comments, 21 unique codes), summarization_150 (per-comment summaries of the same 150 comments), and clustering_670 (TF-IDF + KMeans clustering, k=20, random_state=42, over all 670 comments). All author identifiers in the comment data have been anonymized (replaced with sequential pseudonyms such as @user_0001). The dataset is released to enable independent replication of the comparative analysis. Sentiment labels are automated model outputs with known limitations (see κ values) and should be interpreted as comparative signals, not absolute ground truth.



