遇见数据集

MilvusDB database dumps of DBpedia URIs and short abstracts.

收藏
Zenodo2025-10-04 更新2026-05-26 收录
官方服务:

资源简介:

the extracted *.db files are sqlite databases. they can be loaded in milvusDB for further processing. due to the use of binary blobs for storage, they are unlikely to be useful in contexts other than milvusDB. dbpediaAbstractsAsVectors.tar.gz: - uncompressed: 36.1 GiB- URIs and up to 1024 characters of all english abstracts as text and embedded with LaBSE. German abstracts for object type "thing" are also included.- intended use: vector-based semantic search dbpediaAbstractsNoStopwords.tar.gz:- uncompressed: 8.9 GiB- URIs and up to 1024 characters of all english abstracts as text with stopwords removed. German abstracts for object type "thing" are also included. Stopwords were removed in the respective languages. - intended use: full text search with keyword matchin. - NB: due to some technical issues, the index is not being saved in the DB and needs to be created. This takes around 24h. Both datasets hold around 8.6 million records. 00_EMBEDDED_catsAndAbstracts_from_0_to_1000.json: - sample record of the first 1000 english dbpedia entries in json format.

提供机构:
Zenodo
创建时间:
2025-10-04
二维码
社区交流群
二维码
科研交流群
商业服务