官方服务:
资源简介:
A CCL reader (Corpus2) with MWE detection.
一款适配Corpus2语料库的CCL阅读器,集成多词表达式(Multi-Word Expression,MWE)检测功能
应用场景:
相关数据集
A French corpus annotated for multiword expressions with adverbial function
该数据集是由巴黎东大学研究团队构建的法语多词表达式副词功能标注语料库,旨在支持信息检索与抽取以及句法分析研究。语料库包含两个来源:法国国民议会2006年10月3-4日会议完整记录(AS)和儒勒·凡尔纳小说《八十天环游地球》(JV),总计约168,846个词符和37,856个词型,标注了4,247个多词副词实例。数据集通过结合6,800条目的句法语义词典与有限状态转换器自动标注,并经过人工校验流程构
arXiv2026-06-03 更新110
Dataset of "Knowledge-based Sense Disambiguation of Multiword Expressions in Requirements Documents"
This is the dataset used in the paper "Knowledge-based Sense Disambiguation of Multiword Expressions in Requirements Documents" at AIRE'21 In this paper, we explore the use of a multiword expression d
Zenodo2021-08-06 更新40
PANACEA Labour Multi Word Italian Lexicon
The Labour MW Italian Lexicon is a lexicon of noun-noun multiword expressions automatically /nextracted from a 70Mio word web crawled corpus in the labour law domain. The...
B2FIND10
CoAM: Corpus of All-Type Multiword Expressions
CoAM数据集是由奈良先端科学技术大学院大学和Resolve Research共同创建的多词表达(MWE)识别数据集,包含1300个句子。该数据集旨在解决现有MWE识别数据集标注不一致、类型单一或规模有限的问题。数据集通过多步骤构建过程,包括人工标注、人工审查和自动化一致性检查,确保数据质量。数据集中的MWE被标记为不同类型(如名词、动词等),以便进行细粒度的错误分析。数据集的应用领域包括机器翻译
arXiv2024-12-24 更新60
Spanish results of the PARSEME shared task edition 1.1 by MWE category.
Only the three categories including VNMWEs are shown in the Table. Apart from the average score obtained in each category, the lowest and highest scores are also shown between brackets.
Figshare2020-08-27 更新20



