AMSunda: A Novel Dataset for Sundanese Information Retrieval
收藏官方服务:
资源简介:
The AMSunda dataset was introduced as the first resource designed explicitly for fine-tuning and evaluating embedding models in the Sundanese language. AMSunda dataset consists of two dataset types: (1) triplet data containing a query passage, a positive, and a negative response aimed for fine-tuning embedding models, and (2) BEIR-compatible data structured for evaluating embedding models on retrieval tasks.
提供机构:
Zenodo创建时间:
2025-05-01



