Chinese Text Semantic Matching Dataset
收藏资源简介:
This dataset is designed for Chinese text semantic matching. It incorporates the original ATEC dataset and additional data we collected ourselves. The content covers everyday life, the financial sector, daily conversations, idioms and proverbs, web slang, and other semantic-matching scenarios.Size: 70,860 KBTraining pairs: 735,956Test-set size: 5,137 KBTest pairs: 59,523All data were gathered automatically from the open web, so the topical span is very broad.Format: A tabular file with three columnsColumn A: first text to be matchedColumn B: second text to be matchedColumn C: label (1 = match, 0 = no match)Reference dataset cited:ATEC: Alibaba Taobao E-commerce Click-Through Rate Prediction Dataset. Alibaba Group. 20 Oct 2023. Available at: https://www.atecup.cn/ods. https://www.atecup.cn/ods



