pythonformer/Trajectory-Stitching-Test-7M
收藏资源简介:
该数据集通过高度优化的管道在双H100 NVL GPU集群上构建,采用两阶段算法自动拼接节点,无需外部LLM调用。它从源数据集中提取高信息关键词,包括大写单词和长单词,并使用子代理种子生成研究轨迹。通过倒排索引实现局部关键词链接,确保全局唯一性,再通过BGE-Large模型进行语义连贯性过滤(相似度阈值0.55),最后注入元认知提示以模拟AI研究计划。数据集旨在创建独特、语义连贯的研究轨迹。
This dataset is built via a highly optimized pipeline running on a dual-H100 NVL GPU cluster, using a two-pass algorithm for autonomous stitching without external LLM calls. It extracts high-information keywords, including capitalized words and long words, from the source dataset and employs sub-agent seeding to generate research trajectories. Local keyword linkage is achieved through an inverted index, ensuring global uniqueness, followed by semantic coherence filtering using the BGE-Large model (similarity threshold 0.55). Finally, meta-cognitive prompts are injected to emulate AI research planning. The dataset aims to create unique and semantically coherent research trajectories.




