遇见数据集

Text embeddings from the MultiCaRe dataset

收藏
Zenodo2026-05-22 更新2026-05-26 收录
官方服务:

资源简介:

The MultiCaRe dataset contains multi-modal data from over 70,000 open access and de-identified case reports from PubMed Central. The full dataset includes metadata, clinical cases, image captions and more than 130,000 images, but this subset contains only the textual clinical cases and their embeddings (using "gemini-embedding-001" embeddings model, with 768 dimensions). The license of the dataset as a whole is CC BY-NC-SA. However, its individual contents may have less restrictive license types (CC BY, CC BY-NC, CC0). The license information and the citation data of each article can be found in the metadata.parquet file.

提供机构:
Zenodo
创建时间:
2025-04-30
二维码
社区交流群
二维码
科研交流群
商业服务