dnagpt/omnigene4-mm-corpus
收藏资源简介:
OmniGene-4-MM统一语料库是一个多模态训练数据集,用于OmniGene-4-MM的第1至第3阶段。每个数据行是一个JSON对象,包含messages(聊天格式)、images(相对图像路径列表)和modality字段。视觉数据行引用了来自多个源数据集的图像,包括Vis-CheBI20(化学结构识别)、PubMedVision(生物医学文献图像)、HPA10M(人类蛋白质图谱显微镜图像)、ChartQA(图表问答)以及内部合成的生物医学视觉任务。该语料库旨在支持生物信息学、多模态和视觉语言任务,涵盖生物学、蛋白质和DNA等领域。用户可以通过下载源视觉数据集并运行特定脚本来重现训练过程。
The OmniGene-4-MM unified corpus is a multi-modal training dataset used for OmniGene-4-MM Stages 1–3. Each row is a JSON object with messages (chat-format), images (list of relative image paths), and modality field. Vision rows reference images from source datasets including Vis-CheBI20 (chemical structure recognition), PubMedVision (biomedical literature images), HPA10M (Human Protein Atlas microscopy images), ChartQA (chart question answering), and synthetic biomedical visual tasks (project-internal). It is designed for bioinformatics, multimodal, and vision-language tasks, covering areas such as biology, protein, and DNA. Users can reproduce the training by downloading the source vision datasets and running a specific script.




