kev-KOH/time-embed-bge-m3
收藏资源简介:
该数据集是用于FlagEmbedding/BGE-M3微调的韩语LMS时间查询数据集。目标是使语义上等价的韩语相对时间表达在嵌入空间中接近,同时区分看起来相似但时间不同的表达。数据集包含精心策划的LMS时间模板和确定性日历焦点层,用于处理硬案例,如精确持续时间偏移、日历周期、滚动窗口、相邻日历周期和条件重合边界。LLM生成的模板可能被用作提议,但除非通过意图特定的确定性验证,否则不会作为最终训练数据。排除产生格式错误字符串的截止时间模板,例如重复的까지后缀。数据集格式为JSONL,每行包含查询语句、正例和负例。
This dataset is a FlagEmbedding/BGE-M3 fine-tuning dataset for Korean LMS temporal queries. The goal is to make semantically equivalent Korean relative time expressions close in embedding space while separating similar-looking but temporally different expressions. The root files now point to the v1.7 training dataset. v1.7 keeps the curated LMS temporal templates and adds a deterministic calendar-focus layer for hard cases such as exact duration offsets, calendar periods, rolling windows, adjacent calendar periods, and conditional coincidence boundaries. LLM-generated templates may be used as proposals, but they are not trusted as final training data unless they pass intent-specific deterministic validation. Deadline templates that create malformed strings such as duplicate `까지` suffixes are excluded.





