MFAQ
收藏资源简介:
MFAQ是一个公开的多语言FAQ数据集,由CLiPS研究中心在安特卫普大学创建。该数据集从网络收集了约613万条FAQ对,涵盖21种不同语言,显著大于现有的FAQ检索数据集。MFAQ数据集面临内容重复和话题分布不均的挑战,但通过采用与Dense Passage Retrieval类似的设置,并测试多种双编码器,发现基于XLM-RoBERTa的多语言模型表现最佳。数据集的应用领域包括FAQ检索,旨在通过自动回答最常见问题,优化用户服务体验,如邮件、聊天机器人或搜索栏。
MFAQ is a publicly available multilingual FAQ dataset created by the CLiPS Research Center at the University of Antwerp. This dataset collects approximately 6.13 million FAQ pairs from the web across 21 distinct languages, and is significantly larger than existing FAQ retrieval datasets. The MFAQ dataset faces challenges including content duplication and uneven topic distribution. However, by adopting a setup similar to that of Dense Passage Retrieval and testing multiple dual encoders, it was found that multilingual models based on XLM-RoBERTa achieve the best performance. The application scenarios of the dataset include FAQ retrieval, which aims to optimize user service experiences such as those via emails, chatbots, or search bars by automatically answering frequently asked questions.




