遇见数据集

SenWiCh: Sense-Annotated Sentences for WSD and WiC in Low-Resource Languages

收藏
Zenodo2026-03-17 更新2026-05-26 收录
官方服务:

资源简介:

SenWiCh is a multilingual dataset of sense-annotated sentences designed to support research in Word Sense Disambiguation (WSD) and Word-in-Context (WiC) tasks for low-resource languages. It covers fourteen typologically and geographically diverse languages: Azerbaijani, Bangla, Bulgarian, Cantonese, Kannada, Korean, Mandarin, Marathi, Polish, Punjabi, Swahili, Telugu, Urdu, and Vietnamese. Each dataset contains sentences featuring polysemous words, annotated for word senses using a semi-automatic method that combines embedding-based clustering with interactive manual annotation. This approach ensures broad semantic coverage and efficient sense labeling across diverse linguistic contexts. The release includes three main components: wic_files.zip: Contains WiC-formatted JSON files ({lang}_train.data, {lang}_dev.data, {lang}_test.data) for each language, suitable for binary classification-based training and evaluation. wsd_files.zip: Provides full sets of sense-annotated sentences for each language in CSV format ({lang}.csv), supporting fine-grained WSD tasks. sense_lists.zip: Includes CSV files ({lang}.csv) listing the polysemous lemmas, their associated senses, and unique sense IDs, useful for lexical resource analysis and alignment. The dataset creation process is described in our paper "SenWiCh: Sense-Annotation of Sentences in Low-Resource Languages for WiC using Hybrid Methods" (SIGTYP 2025, ACL). [arXiv link]If you use any of the datasets or the tool, please cite (SIGTYP 2025, ACL). To support dataset creation and adaptation to new languages, we also release the interactive annotation tool used in this process: projecting_sentences. This work was partially supported by the research program Change is Key!, funded by Riksbankens Jubileumsfond (grant number M21-0021).

提供机构:
Zenodo
创建时间:
2026-02-27
二维码
社区交流群
二维码
科研交流群
商业服务