遇见数据集

muhamparlak/turkish-law-bge-m3-embeddings

收藏
Hugging Face2026-05-16 更新2026-05-31 收录
官方服务:

资源简介:

该数据集包含土耳其法律先例(最高法院/Yargıtay和国务委员会/Danıştay)以及成文立法(法典和条款)的预计算嵌入。数据以Parquet格式存储,专门针对使用“多索引”架构的RAG(检索增强生成)流水线进行优化。数据集分为两个子集:法律先例(emsal_kararlar)和立法(mevzuat)。法律先例子集源自公开的法院记录,通过Langchain的RecursiveCharacterTextSplitter进行分块处理,分块大小为3000字符,重叠500字符,以保持语义段落和句子。立法子集则包括土耳其核心法典和法律条款,未进行分块,但通过上下文丰富格式以提高向量表示。嵌入使用BAAI/bge-m3模型生成,向量维度为1024,最大序列长度为2048,处理在NVIDIA A100 GPU上以torch.bfloat16精度完成。数据集架构包括文档ID、分块索引、文本、来源、元数据和嵌入向量等字段,便于法律信息检索和分析应用。

This dataset contains pre-computed embeddings for Turkish legal precedents (Supreme Court/Yargıtay and Council of State/Danıştay) and statutory legislation (codes and articles). It is formatted in Parquet and specifically optimized for RAG (Retrieval-Augmented Generation) pipelines using a Multi-Index architecture. The dataset is divided into two subsets: legal precedents (emsal_kararlar) and legislation (mevzuat). The legal precedents subset is derived from public court records, chunked using Langchains RecursiveCharacterTextSplitter with a chunk size of 3000 characters and overlap of 500 characters to maintain semantic paragraphs and sentences. The legislation subset includes Turkish core codes and laws, without chunking but enriched with contextual formatting for better vector representations. Embeddings are generated using the BAAI/bge-m3 model, with a vector dimension of 1024, max sequence length of 2048, and processed on an NVIDIA A100 GPU in torch.bfloat16 precision. The dataset schema includes fields such as document ID, chunk index, text, source, metadata, and embedding vectors, facilitating legal information retrieval and analysis applications.

提供机构:
muhamparlak
二维码
社区交流群
二维码
科研交流群
商业服务