Chat2Find/Chat2Find-Corpus
收藏资源简介:
--- license: mit language: - si - ta - en tags: - conversational - trilingual - sri-lanka - chat2find pretty_name: Chat2Find Corpus size_categories: - 100K<n<1M --- # Chat2Find Corpus The **Chat2Find Corpus** is a high-quality trilingual conversational dataset derived from real-world interactions on the [Chat2Find](https://chat2find.com) platform. It contains approximately **255 Million tokens** in **Sinhala (සිංහල)**, **Tamil (தமிழ்)**, and **English**, including significant instances of **Singlish** and **Tanglish** (transliterated and code-mixed speech). This makes it an ideal resource for Continual Pre-training (CPT) and Supervised Fine-Tuning (SFT) of Large Language Models in low-resource and multilingual contexts. --- ## New: Chat2Find Reasoning & Tool Datasets We have just released the instruction-tuned and reasoning subsets, designed to teach foundation models complex logic, chain-of-thought, and API tool-calling in Sinhala, Tamil, and English. * **[Free Preview (5,000 Records)](https://huggingface.co/datasets/Chat2Find/Chat2Find-Instruct-Reasoning-Sample):** A robust sample containing heavily mixed single-turn instructions and deep multi-turn agentic workflows. * **[Full Dataset (279k Records / 1.8 GB)](https://huggingface.co/datasets/Chat2Find/Chat2Find-Instruct-Reasoning-Dataset):** The massive, complete instruction dataset available under a commercial/advanced research license. --- ## Dataset Details - **Origin:** Real conversations and tool-assisted interactions from Chat2Find. - **Languages:** Trilingual (Sinhala, Tamil, English) + **Singlish** & **Tanglish**. - **Format:** JSON Lines (`.jsonl`). - **Volume:** Approximately 255 Million tokens (~279,248 records) - **Content:** Information seeking, role-playing, and cultural/local knowledge specific to the South Asian region (primarily Sri Lanka and India). ## Key Features - **Non-Synthetic:** Unlike many large-scale datasets, this corpus originates from real platform usage, capturing natural language patterns, code-switching, and local nuances. - **Trilingual & Code-Mixed:** Seamless transitions between Sinhala, Tamil, and English. Naturally includes Singlish and Tanglish as used by the community. - **Foundation for Future Models:** This corpus is the primary training data for Chat2Find's upcoming open-weights model suite. - **Metadata:** Each record includes a `source` tag and a unique `record_id`. ## Intended Use Specifically designed for: 1. **Continual Pre-training (CPT):** To enhance the multilingual capabilities of base models (e.g., Qwen, Llama, Gemma). 2. **Domain Adaptation:** Improving model performance on Sri Lankan/South Asian cultural, logistical, and linguistic queries. 3. **Research:** Exploring code-switching and multilingual information retrieval. ## Dataset Structure The dataset is partitioned into multiple JSON Lines (`.jsonl`) files located in the `data/` directory for better accessibility and streaming. Each line is a JSON object: ```json { "text": "...", "metadata": { "source": "chat2find.com", "record_id": 12345 } } ``` ## Licensing This dataset is released under the **MIT License**. ## Upcoming Model Releases (Coming Soon) Chat2Find is actively developing a suite of models trained on this corpus. These will be released with **Open Weights** to the community: 1. **Chat2Find Base:** A foundational trilingual model. 2. **Chat2Find Instruct:** Optimised for following complex instructions in Sinhala, Tamil, and English. 3. **Chat2Find Reasoning:** A high-logic model designed for complex problem-solving and chain-of-thought reasoning in a multilingual context. **Stay tuned to this repository and chat2find.com for updates.**
--- license: MIT许可证 language: - 僧伽罗语(si) - 泰米尔语(ta) - 英语(en) tags: - 对话式 - 三语 - 斯里兰卡 - chat2find pretty_name: Chat2Find语料库 size_categories: - 100K<n<1M --- # Chat2Find语料库 **Chat2Find语料库**是源自[Chat2Find](https://chat2find.com)平台真实交互场景的高质量三语对话数据集。该数据集包含约**2.55亿个Token**,覆盖**僧伽罗语(සිංහල)**、**泰米尔语(தமிழ்)**与**英语**,同时包含大量**新加坡式英语(Singlish)**与**泰米尔语转写混用语(Tanglish)**实例,非常适合在低资源多语言场景下对大语言模型(Large Language Model, LLM)进行持续预训练(Continual Pre-training, CPT)与监督微调(Supervised Fine-Tuning, SFT)。 --- ## 新增:Chat2Find推理与工具数据集 我们刚刚发布了指令微调与推理子集,旨在帮助基础模型掌握僧伽罗语、泰米尔语与英语环境下的复杂逻辑、思维链(Chain-of-Thought)与API工具调用能力。 * **[免费预览(5000条数据)](https://huggingface.co/datasets/Chat2Find/Chat2Find-Instruct-Reasoning-Sample):** 包含大量混合式单轮指令与深度多轮AI智能体(AI Agent)工作流的优质样本集。 * **[完整数据集(27.9万条数据 / 1.8 GB)](https://huggingface.co/datasets/Chat2Find/Chat2Find-Instruct-Reasoning-Dataset):** 体量庞大的完整指令数据集,可通过商业/高级研究许可获取。 --- ## 数据集详情 - **来源**:Chat2Find平台的真实对话与工具辅助交互内容。 - **语言**:三语(僧伽罗语、泰米尔语、英语)+ 新加坡式英语(Singlish)与泰米尔语转写混用语(Tanglish)。 - **格式**:JSON Lines(`.jsonl`)格式。 - **规模**:约2.55亿个Token(约279248条数据) - **内容**:面向南亚地区(主要为斯里兰卡与印度)的信息查询、角色扮演与本地化文化知识相关内容。 ## 核心特性 - **非合成生成**:与多数大规模数据集不同,该语料库源自平台真实使用场景,完整保留了自然语言模式、语码转换与地域语言细节。 - **三语与语码混合**:支持僧伽罗语、泰米尔语与英语间的无缝切换,天然包含社区常用的新加坡式英语(Singlish)与泰米尔语转写混用语(Tanglish)。 - **下一代模型训练基石**:该语料库是Chat2Find即将推出的开源权重模型套件的核心训练数据。 - **元数据**:每条数据均包含`source`标签与唯一的`record_id`。 ## 预期用途 专为以下场景设计: 1. **持续预训练(CPT)**:增强基础模型(如Qwen、Llama、Gemma)的多语言能力。 2. **领域适配**:提升模型在斯里兰卡/南亚文化、物流与语言查询场景下的性能。 3. **研究**:探索语码转换与多语言信息检索相关课题。 ## 数据集结构 为便于访问与流式读取,数据集被拆分为多个存储于`data/`目录下的JSON Lines(`.jsonl`)文件。 每行均为一个JSON对象: json { "text": "...", "metadata": { "source": "chat2find.com", "record_id": 12345 } } ## 许可证 本数据集采用**MIT许可证**发布。 ## 即将发布的模型套件(即将上线) Chat2Find正在开发基于该语料库训练的模型套件,将以开源权重形式向社区发布: 1. **Chat2Find Base**:基础三语模型。 2. **Chat2Find Instruct**:针对僧伽罗语、泰米尔语与英语下的复杂指令遵循任务优化的模型。 3. **Chat2Find Reasoning**:面向多语言场景下复杂问题求解与思维链推理的高逻辑能力模型。 **敬请关注本仓库与chat2find.com获取最新动态。**




