遇见数据集

pythonformer/Trajectory-Stitching-Test-7M

收藏
Hugging Face2026-05-06 更新2026-05-31 收录
官方服务:

资源简介:

该数据集通过高度优化的管道在双H100 NVL GPU集群上构建,采用两阶段算法自动拼接节点,无需外部LLM调用。它从源数据集中提取高信息关键词,包括大写单词和长单词,并使用子代理种子生成研究轨迹。通过倒排索引实现局部关键词链接,确保全局唯一性,再通过BGE-Large模型进行语义连贯性过滤(相似度阈值0.55),最后注入元认知提示以模拟AI研究计划。数据集旨在创建独特、语义连贯的研究轨迹。

This dataset is built via a highly optimized pipeline running on a dual-H100 NVL GPU cluster, using a two-pass algorithm for autonomous stitching without external LLM calls. It extracts high-information keywords, including capitalized words and long words, from the source dataset and employs sub-agent seeding to generate research trajectories. Local keyword linkage is achieved through an inverted index, ensuring global uniqueness, followed by semantic coherence filtering using the BGE-Large model (similarity threshold 0.55). Finally, meta-cognitive prompts are injected to emulate AI research planning. The dataset aims to create unique and semantically coherent research trajectories.

提供机构:
pythonformer
二维码
社区交流群
二维码
科研交流群
商业服务