ORBIT-curated astronomy dataset
收藏资源简介:
ORBIT-curated astronomy dataset是由伊利诺伊大学厄巴纳-香槟分校的研究团队创建的高质量天文学领域数据集,包含100亿个Tokens。该数据集从FineWeb-Edu数据集中筛选而来,结合了学术文本和网络教育内容,旨在为大语言模型提供深度和广度兼具的天文学知识。数据集的创建过程采用了嵌入式相似性匹配和BERT回归模型进行过滤,确保了数据的质量和相关性。该数据集主要用于提升大语言模型在天文学领域的性能,解决通用模型在专业领域知识不足的问题。
ORBIT-curated astronomy dataset is a high-quality domain-specific astronomy dataset created by the research team from the University of Illinois Urbana-Champaign, which contains 10 billion Tokens. This dataset is curated from the FineWeb-Edu corpus, combining academic texts and online educational content, with the objective of supplying Large Language Models (LLMs) with astronomical knowledge that balances both depth and breadth. During the dataset construction process, embedding-based similarity matching and BERT regression models were employed for filtering, ensuring the quality and relevance of the data. This dataset is primarily utilized to enhance the performance of LLMs in the astronomy domain, addressing the issue of inadequate professional domain knowledge in general-purpose AI models.




