VN-MSASar: Multilingual Aspect-Based Sentiment and Sarcasm Analysis
收藏资源简介:
<p>VN-MSASar is an aspect-based sentiment analysis (ABSA) dataset for the hotel and restaurant domain in Vietnam, covering four languages (Vietnamese, English, Chinese, Japanese). Version 3 contains 76,785 (sentence, aspect) pairs drawn from 18,609 reviews over nine service aspects, each with an ordinal sentiment label on a [-3,+3] scale and a binary sarcasm label.</p> <p>Every row is a real review. The 7,659 synthetic, template-generated sentences that formed the <code>augmented</code> source in version 1 have been removed. Splitting moved from the sentence to the review, so no review appears in two splits, and near-duplicate leakage was audited under a two-part criterion (TF-IDF character 3–5-gram cosine, or character-4-gram containment) that reports zero flagged pairs between train and either evaluation split. Results computed on version 1 are not comparable to version 3.</p> <p>The deposit includes both the aspect-pair version (<code>vn_msa_sar_v3</code>) and the derived multi-label aspect-detection version (<code>vn_hotel_res_aspect_v3</code>) with matching train/validation/test splits, the annotation codebook, the split-construction and leakage-audit scripts, and the collection scripts.</p> <p><strong>Reporting note.</strong> The <code>other</code> source is a deliberate over-sample of sarcastic reviews. Primary results should be reported on the natural slice <code>test[source != "other"]</code>; the full test set is for sarcasm analysis, where the larger positive class gives tighter confidence intervals.</p>



