Caselaw_Access_Project_embeddings
收藏资源简介:
这是一个为Caselaw Access Project创建的嵌入数据集,由用户Endomorphosis生成。每个法律案例条目通过IPFS/multiformats哈希处理,可通过IPFS/filecoin网络检索。数据集的嵌入由三个模型生成,分别是thenlper/gte-small、Alibaba-NLP/gte-large-en-v1.5和Alibaba-NLP/gte-Qwen2-1.5B-instruct,它们的上下文长度和维度各不相同。嵌入被分成了4096个集群,每个集群的质心和内容ID都提供了。建议在客户端搜索嵌入时,先查询质心,再检索最接近的gte-small集群,然后查询该集群。
This is an embedding dataset created for the Caselaw Access Project, generated by user Endomorphosis. Each legal case entry is hashed via IPFS/multiformats and retrievable through the IPFS/Filecoin network. The embeddings of the dataset are generated by three models: thenlper/gte-small, Alibaba-NLP/gte-large-en-v1.5, and Alibaba-NLP/gte-Qwen2-1.5B-instruct, each with distinct context lengths and embedding dimensionalities. The embeddings are partitioned into 4096 clusters, with the centroid and content ID of each cluster provided. When searching for embeddings on the client side, it is recommended to first query the centroids, retrieve the closest gte-small cluster, and then query that cluster.




