VAST
收藏资源简介:
VAST数据集由哥伦比亚大学创建,专注于零样本立场检测,涵盖广泛的主题和词汇变异。数据集包含大量主题,如政治、教育和公共卫生,并捕捉了人类可能真实描述同一主题的多种表达方式。创建过程涉及从ARC语料库中提取特定主题,并通过众包收集立场标签。VAST数据集适用于开发零样本和少样本立场检测模型,旨在解决模型在真实世界中对广泛主题的泛化能力评估问题。
The VAST dataset was developed by Columbia University, focusing on zero-shot stance detection and covering a broad spectrum of topics and lexical variations. It encompasses numerous topics including politics, education, and public health, and captures the diverse authentic expressions that humans may employ when describing the same subject. The dataset's creation involved extracting specific topics from the ARC corpus and collecting stance labels via crowdsourcing. The VAST dataset is suitable for developing zero-shot and few-shot stance detection models, and aims to address the challenge of evaluating a model's generalization capabilities across a wide range of topics in real-world scenarios.




