mteb/TMD
收藏资源简介:
TMD(Temporal Memory Dataset)是一个用于大规模文本嵌入基准测试(MTEB)的数据集,专注于基于LMEB(长时程记忆嵌入基准)的对话记忆检索任务。它评估在对话中日期、会话和相对时间检索的性能,涉及社交和口语领域。数据集包括多个配置,如content_time_qs、date_span_time_qs等,每个配置包含语料库(文本和标题)、查询、相关性评分和排名数据,用于文本检索和多选问答任务。数据来源于KaLM-Embedding/LMEB,语言为英语,为单语种,标注由衍生而来,适用于文本检索任务类别。
TMD (Temporal Memory Dataset) is an MTEB (Massive Text Embedding Benchmark) dataset based on the LMEB (Long-horizon Memory Embedding Benchmark) dialogue-memory retrieval task, evaluating date, session, and relative-time retrieval over dialogues in social and spoken domains. It includes multiple configurations such as content_time_qs, date_span_time_qs, etc., each with corpus (text and title), queries, relevance scores, and ranked data for text retrieval and multiple-choice QA tasks. The dataset is derived from KaLM-Embedding/LMEB, in English, monolingual, and falls under the text-retrieval task category.



