Webis Clickbait Spoiling Corpus 2022
收藏资源简介:
<strong>Webis Clickbait Spoiling Corpus 2022</strong> The Webis Clickbait Spoiling Corpus 2022 (Webis-Clickbait-22) contains 5,000 spoiled clickbait posts crawled from Facebook, Reddit, and Twitter.<br> This corpus supports the task of clickbait spoiling, which deals with generating a short text that satisfies the curiosity induced by a clickbait post. This dataset contains the clickbait posts and manually cleaned versions of the linked documents, and extracted spoilers for each clickbait post.<br> Additionally, the spoilers are categorized into three types: short phrase spoilers, longer passage spoilers, and multiple non-consecutive pieces of text. This dataset contains the clickbait posts and manually cleaned versions of the linked documents, and extracted spoilers for each clickbait post.<br> Additionally, the spoilers are categorized into three types: short phrase spoilers, longer passage spoilers, and multiple non-consecutive pieces of text. The test set of this dataset was used for the SemEval-2023 clickbait spoiling task. You can re-execute and adopt the software submissions made through for this SemEval task, please see the instructions and overview of approaches in TIRA. <strong>Overview</strong> The dataset comes with predefined train/validation/test splits: training.jsonl: 3,200 posts for training validation.jsonl: 800 posts for validation test.jsonl: 1,000 posts for testing The test set was used for the SemEval-2023 clickbait spoiling task. This shared task was organized with TIRA.io and participants submitted Docker software during the task. Please see the instructions in TIRA to re-execute or modify the approaches.



