TW-GSAT-Chinese
收藏资源简介:
该数据集是一个台湾学科能力测验的中文考科资料库,包含了考试题目和相关内容,适用于自然语言处理任务如token分类和问题回答。数据集以纯文本形式存在,没有图片,便于训练。所有信息都包含在'text_pre_format'字段中,可直接用于预训练。同时,数据集还提供了适用于TW-TextForge工具的格式。数据集遵循Apache 2.0开源许可,可用于商业、研究和私人使用。
This dataset is a Chinese subject question bank for the Taiwan Scholastic Ability Test. It contains exam questions and related materials, suitable for natural language processing (NLP) tasks including token classification and question answering (QA). The dataset is stored in plain text format without any images, which facilitates model training. All relevant information is contained within the 'text_pre_format' field and can be directly used for pre-training. Additionally, the dataset provides a format compatible with the TW-TextForge tool. The dataset is licensed under the Apache 2.0 open-source license, allowing commercial, research and private use.
台灣本土語言模型語料庫:台灣學科能力測驗-中文考科
基本資訊
- 許可證: Apache 2.0
- 語言: 繁體中文 (zh)
- 資料集大小: 778781 bytes
- 下載大小: 481068 bytes
- 樣本數量: 347
- 任務類別: 標記分類 (token-classification)、問答 (question-answering)
- 規模分類: n<1K
資料特徵
- 欄位:
- year (int64): 年份
- id (int64): 識別碼
- question_type (string): 問題類型
- article (string): 文章內容
- question (string): 問題
- A (string): 選項A
- B (string): 選項B
- C (string): 選項C
- D (string): 選項D
- E (string): 選項E
- grading_criteria (float64): 評分標準
- answer (string): 答案
- answer_rate (float64): 答對率
- text_pre_format (string): 預處理文本格式
- text_pre_tw_textforge_format (string): TW-TextForge專用格式
- references (string): 參考資料
資料集特色
- 純文字內容,無圖片,降低訓練難度。
text_pre_format包含所有資訊,可直接用於預訓練。text_pre_tw_textforge_format專用於TW-TextForge產生題目分析。
法律聲明
- 根據中華民國著作權法第9條,依法令舉行的考試試題不具著作權。
- 資料集已調整題目敘述,使其更適合NLP任務。
- 學測題目引用內容可能仍有著作權保護,本資料集不收集此類內容。
- 使用時需遵守Apache 2.0許可證。




