遇见数据集

AMSunda: A Novel Dataset for Sundanese Information Retrieval

收藏
Zenodo2025-05-23 更新2026-05-26 收录
官方服务:

资源简介:

The AMSunda dataset was introduced as the first resource designed explicitly for fine-tuning and evaluating embedding models in the Sundanese language. AMSunda dataset consists of two dataset types: (1) triplet data containing a query passage, a positive, and a negative response aimed for fine-tuning embedding models, and (2) BEIR-compatible data structured for evaluating embedding models on retrieval tasks.

提供机构:
Zenodo
创建时间:
2025-05-01
二维码
社区交流群
二维码
科研交流群
商业服务