遇见数据集

wisenut-nlp-team/llama_ko_smr

收藏
Hugging Face2024-04-30 更新2024-06-12 收录
官方服务:

资源简介:

--- dataset_info: - config_name: art features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 23253173 num_examples: 15627 download_size: 12801716 dataset_size: 23253173 - config_name: artifact_science features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 362643834 num_examples: 89531 download_size: 167429211 dataset_size: 362643834 - config_name: beauty_and_health features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 11495982 num_examples: 19203 download_size: 6174548 dataset_size: 11495982 - config_name: briefing features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string - name: tpye dtype: string splits: - name: train num_bytes: 84092000 num_examples: 36000 download_size: 26138279 dataset_size: 84092000 - config_name: c_event features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string - name: tpye dtype: string splits: - name: train num_bytes: 70105743 num_examples: 31166 download_size: 21295859 dataset_size: 70105743 - config_name: culture features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string - name: tpye dtype: string splits: - name: train num_bytes: 35908844 num_examples: 23700 download_size: 11289413 dataset_size: 35908844 - config_name: daily_and_occupation features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 14495402 num_examples: 22982 download_size: 7769431 dataset_size: 14495402 - config_name: edit features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string - name: tpye dtype: string splits: - name: train num_bytes: 41226597 num_examples: 18000 download_size: 13617131 dataset_size: 41226597 - config_name: editorial features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 204950743 num_examples: 63768 download_size: 117562937 dataset_size: 204950743 - config_name: education features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 8992532 num_examples: 14759 download_size: 4846739 dataset_size: 8992532 - config_name: enter features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string - name: tpye dtype: string splits: - name: train num_bytes: 77007245 num_examples: 36092 download_size: 24622632 dataset_size: 77007245 - config_name: etc features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 13009615 num_examples: 7597 download_size: 6696866 dataset_size: 13009615 - config_name: event features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 13632825 num_examples: 24006 download_size: 7160232 dataset_size: 13632825 - config_name: fm_drama features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string - name: tpye dtype: string splits: - name: train num_bytes: 65279567 num_examples: 36000 download_size: 20994133 dataset_size: 65279567 - config_name: food_and_drink features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 18831258 num_examples: 33957 download_size: 9768013 dataset_size: 18831258 - config_name: fs_drama features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string - name: tpye dtype: string splits: - name: train num_bytes: 62984894 num_examples: 36004 download_size: 20000234 dataset_size: 62984894 - config_name: his_cul features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string - name: tpye dtype: string splits: - name: train num_bytes: 30609601 num_examples: 18000 download_size: 10628675 dataset_size: 30609601 - config_name: history features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string - name: tpye dtype: string splits: - name: train num_bytes: 48220219 num_examples: 25766 download_size: 14665043 dataset_size: 48220219 - config_name: housing_and_living features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 29295812 num_examples: 50827 download_size: 15854030 dataset_size: 29295812 - config_name: law features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 59837947 num_examples: 27333 download_size: 29960383 dataset_size: 59837947 - config_name: leisure features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 23140399 num_examples: 39654 download_size: 12420477 dataset_size: 23140399 - config_name: life_science features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 35720463 num_examples: 7802 download_size: 17482630 dataset_size: 35720463 - config_name: literature features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string - name: tpye dtype: string splits: - name: train num_bytes: 51905166 num_examples: 21600 download_size: 18123605 dataset_size: 51905166 - config_name: minute features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string - name: tpye dtype: string splits: - name: train num_bytes: 149240389 num_examples: 61200 download_size: 41433544 dataset_size: 149240389 - config_name: narration features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string - name: tpye dtype: string splits: - name: train num_bytes: 24511774 num_examples: 18742 download_size: 7720190 dataset_size: 24511774 - config_name: nature_science features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 31775215 num_examples: 10862 download_size: 12939961 dataset_size: 31775215 - config_name: news_r features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string - name: tpye dtype: string splits: - name: train num_bytes: 161506493 num_examples: 48600 download_size: 52108494 dataset_size: 161506493 - config_name: newspaper features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 778034038 num_examples: 274105 download_size: 453662932 dataset_size: 778034038 - config_name: paper features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 669171434 num_examples: 324174 download_size: 354490940 dataset_size: 669171434 - config_name: paper2 features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string - name: tpye dtype: string splits: - name: train num_bytes: 40000149 num_examples: 18000 download_size: 13367455 dataset_size: 40000149 - config_name: patent features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 6932303601 num_examples: 312600 download_size: 2398178917 dataset_size: 6932303601 - config_name: patent_section features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 499358509 num_examples: 151000 download_size: 239316958 dataset_size: 499358509 - config_name: public features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string - name: tpye dtype: string splits: - name: train num_bytes: 40666888 num_examples: 18000 download_size: 12762114 dataset_size: 40666888 - config_name: relationships features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 45706612 num_examples: 80022 download_size: 24000637 dataset_size: 45706612 - config_name: shopping features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 17079513 num_examples: 29586 download_size: 9159776 dataset_size: 17079513 - config_name: social_science features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 186311981 num_examples: 129870 download_size: 96285745 dataset_size: 186311981 - config_name: speech features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string - name: tpye dtype: string splits: - name: train num_bytes: 162899290 num_examples: 72000 download_size: 48896868 dataset_size: 162899290 - config_name: technology_science features: - name: instruction dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 37930287 num_examples: 26907 download_size: 19950147 dataset_size: 37930287 - config_name: wisenut features: - name: instruction dtype: string - name: input dtype: string - name: title dtype: string - name: output dtype: string - name: lenght dtype: string splits: - name: train num_bytes: 440353415 num_examples: 228728 download_size: 145508702 dataset_size: 440353415 configs: - config_name: art data_files: - split: train path: art/train-* - config_name: artifact_science data_files: - split: train path: artifact_science/train-* - config_name: beauty_and_health data_files: - split: train path: beauty_and_health/train-* - config_name: briefing data_files: - split: train path: briefing/train-* - config_name: c_event data_files: - split: train path: c_event/train-* - config_name: culture data_files: - split: train path: culture/train-* - config_name: daily_and_occupation data_files: - split: train path: daily_and_occupation/train-* - config_name: edit data_files: - split: train path: edit/train-* - config_name: editorial data_files: - split: train path: editorial/train-* - config_name: education data_files: - split: train path: education/train-* - config_name: enter data_files: - split: train path: enter/train-* - config_name: etc data_files: - split: train path: etc/train-* - config_name: event data_files: - split: train path: event/train-* - config_name: fm_drama data_files: - split: train path: fm_drama/train-* - config_name: food_and_drink data_files: - split: train path: food_and_drink/train-* - config_name: fs_drama data_files: - split: train path: fs_drama/train-* - config_name: his_cul data_files: - split: train path: his_cul/train-* - config_name: history data_files: - split: train path: history/train-* - config_name: housing_and_living data_files: - split: train path: housing_and_living/train-* - config_name: law data_files: - split: train path: law/train-* - config_name: leisure data_files: - split: train path: leisure/train-* - config_name: life_science data_files: - split: train path: life_science/train-* - config_name: literature data_files: - split: train path: literature/train-* - config_name: minute data_files: - split: train path: minute/train-* - config_name: narration data_files: - split: train path: narration/train-* - config_name: nature_science data_files: - split: train path: nature_science/train-* - config_name: news_r data_files: - split: train path: news_r/train-* - config_name: newspaper data_files: - split: train path: newspaper/train-* - config_name: paper data_files: - split: train path: paper/train-* - config_name: paper2 data_files: - split: train path: paper2/train-* - config_name: patent data_files: - split: train path: patent/train-* - config_name: patent_section data_files: - split: train path: patent_section/train-* - config_name: public data_files: - split: train path: public/train-* - config_name: relationships data_files: - split: train path: relationships/train-* - config_name: shopping data_files: - split: train path: shopping/train-* - config_name: social_science data_files: - split: train path: social_science/train-* - config_name: speech data_files: - split: train path: speech/train-* - config_name: technology_science data_files: - split: train path: technology_science/train-* - config_name: wisenut data_files: - split: train path: wisenut/train-* --- ## [문서요약 텍스트](https://www.aihub.or.kr/aihubdata/data/view.do?currMenu=115&topMenu=100&aihubDataSe=data&dataSetSn=97) - subset: law - length: 27.3k - subset: newspaper - length: 274k - subset: editorial - length: 63.8k ## [도서자료 요약](https://www.aihub.or.kr/aihubdata/data/view.do?currMenu=115&topMenu=100&aihubDataSe=data&dataSetSn=93) * subset: art - length: 15.6k * subset: technology_science - length: 26.9k * subset: social_science - length: 130k * subset: etc - length: 7.6k ## [논문자료 요약](https://www.aihub.or.kr/aihubdata/data/view.do?currMenu=115&topMenu=100&aihubDataSe=data&dataSetSn=90) * subset: paper - length: 324k * subset: patent - length: 313k * subset: patent_section - length: 151k ## [방송 콘텐츠 대본 요약 데이터](https://www.aihub.or.kr/aihubdata/data/view.do?currMenu=115&topMenu=100&aihubDataSe=data&dataSetSn=591) * subset: fm_drama - length: 36k * subset: fs_drama - length: 36k * subset: history - length: 25.8k * subset: culture - length: 23.7k * subset: enter - length: 36k * subset: c_event - length: 31.1k ## [요약문 및 레포트 생성 데이터](https://www.aihub.or.kr/aihubdata/data/view.do?currMenu=115&topMenu=100&aihubDataSe=data&dataSetSn=582) * subset: news_r - length: 48.6k * subset: briefing - length: 36k * subset: his_cul - length: 18k * subset: paper2 - length: 18k * subset: minute - length: 61.2k * subset: edit - length: 18k * subset: public - length: 18k * subset: speech - length: 72k * subset: literature - length: 21.6k * subset: narration - length: 18.7k ## [한국어 대화 요약](https://www.aihub.or.kr/aihubdata/data/view.do?currMenu=115&topMenu=100&aihubDataSe=data&dataSetSn=117) * subset: relationships - length: 80k * subset: beauty_and_health - length: 19.2k * subset: shopping - length: 29.5k * subset: education - length: 14.7k * subset: food_and_drink - length: 33.9k * subset: leisure - length: 39.6k * subset: daily_and_occupation - length: 22.9k * subset: housing_and_living - length: 50.8k * subset: event - length: 24k ## [기술과학 요약 데이터](https://www.aihub.or.kr/aihubdata/data/view.do?currMenu=115&topMenu=100&aihubDataSe=data&dataSetSn=71532) * subset: life_science - length: 7.8k * subset: artifact_science - length: 89.5k * subset: nature_science - length: 10.8k

提供机构:
wisenut-nlp-team
原始信息汇总

数据集概述

本数据集包含多个子集,每个子集针对不同的主题和领域,具有各自的特征和数据规模。以下是各子集的详细信息:

1. 艺术 (art)

  • 特征: instruction, input, output
  • 训练集: 15627个样本,总大小23253173字节
  • 下载大小: 12801716字节

2. 文物科学 (artifact_science)

  • 特征: instruction, input, output
  • 训练集: 89531个样本,总大小362643834字节
  • 下载大小: 167429211字节

3. 美容与健康 (beauty_and_health)

  • 特征: instruction, input, output
  • 训练集: 19203个样本,总大小11495982字节
  • 下载大小: 6174548字节

4. 简报 (briefing)

  • 特征: instruction, input, output, tpye
  • 训练集: 36000个样本,总大小84092000字节
  • 下载大小: 26138279字节

5. C事件 (c_event)

  • 特征: instruction, input, output, tpye
  • 训练集: 31166个样本,总大小70105743字节
  • 下载大小: 21295859字节

6. 文化 (culture)

  • 特征: instruction, input, output, tpye
  • 训练集: 23700个样本,总大小35908844字节
  • 下载大小: 11289413字节

7. 日常生活与职业 (daily_and_occupation)

  • 特征: instruction, input, output
  • 训练集: 22982个样本,总大小14495402字节
  • 下载大小: 7769431字节

8. 编辑 (edit)

  • 特征: instruction, input, output, tpye
  • 训练集: 18000个样本,总大小41226597字节
  • 下载大小: 13617131字节

9. 社论 (editorial)

  • 特征: instruction, input, output
  • 训练集: 63768个样本,总大小204950743字节
  • 下载大小: 117562937字节

10. 教育 (education)

  • 特征: instruction, input, output
  • 训练集: 14759个样本,总大小8992532字节
  • 下载大小: 4846739字节

11. 进入 (enter)

  • 特征: instruction, input, output, tpye
  • 训练集: 36092个样本,总大小77007245字节
  • 下载大小: 24622632字节

12. 其他 (etc)

  • 特征: instruction, input, output
  • 训练集: 7597个样本,总大小13009615字节
  • 下载大小: 6696866字节

13. 事件 (event)

  • 特征: instruction, input, output
  • 训练集: 24006个样本,总大小13632825字节
  • 下载大小: 7160232字节

14. FM戏剧 (fm_drama)

  • 特征: instruction, input, output, tpye
  • 训练集: 36000个样本,总大小65279567字节
  • 下载大小: 20994133字节

15. 食品与饮料 (food_and_drink)

  • 特征: instruction, input, output
  • 训练集: 33957个样本,总大小18831258字节
  • 下载大小: 9768013字节

16. FS戏剧 (fs_drama)

  • 特征: instruction, input, output, tpye
  • 训练集: 36004个样本,总大小62984894字节
  • 下载大小: 20000234字节

17. 历史与文化 (his_cul)

  • 特征: instruction, input, output, tpye
  • 训练集: 18000个样本,总大小30609601字节
  • 下载大小: 10628675字节

18. 历史 (history)

  • 特征: instruction, input, output, tpye
  • 训练集: 25766个样本,总大小48220219字节
  • 下载大小: 14665043字节

19. 住房与生活 (housing_and_living)

  • 特征: instruction, input, output
  • 训练集: 50827个样本,总大小29295812字节
  • 下载大小: 15854030字节

20. 法律 (law)

  • 特征: instruction, input, output
  • 训练集: 27333个样本,总大小59837947字节
  • 下载大小: 29960383字节

21. 休闲 (leisure)

  • 特征: instruction, input, output
  • 训练集: 39654个样本,总大小23140399字节
  • 下载大小: 12420477字节

22. 生命科学 (life_science)

  • 特征: instruction, input, output
  • 训练集: 7802个样本,总大小35720463字节
  • 下载大小: 17482630字节

23. 文学 (literature)

  • 特征: instruction, input, output, tpye
  • 训练集: 21600个样本,总大小51905166字节
  • 下载大小: 18123605字节

24. 分钟 (minute)

  • 特征: instruction, input, output, tpye
  • 训练集: 61200个样本,总大小149240389字节
  • 下载大小: 41433544字节

25. 叙述 (narration)

  • 特征: instruction, input, output, tpye
  • 训练集: 18742个样本,总大小24511774字节
  • 下载大小: 7720190字节

26. 自然科学 (nature_science)

  • 特征: instruction, input, output
  • 训练集: 10862个样本,总大小31775215字节
  • 下载大小: 12939961字节

27. 新闻R (news_r)

  • 特征: instruction, input, output, tpye
  • 训练集: 48600个样本,总大小161506493字节
  • 下载大小: 52108494字节

28. 报纸 (newspaper)

  • 特征: instruction, input, output
  • 训练集: 274105个样本,总大小778034038字节
  • 下载大小: 453662932字节

29. 论文 (paper)

  • 特征: instruction, input, output
  • 训练集: 324174个样本,总大小669171434字节
  • 下载大小: 354490940字节

30. 论文2 (paper2)

  • 特征: instruction, input, output, tpye
  • 训练集: 18000个样本,总大小40000149字节
  • 下载大小: 13367455字节

31. 专利 (patent)

  • 特征: instruction, input, output
  • 训练集: 312600个样本,总大小6932303601字节
  • 下载大小: 2398178917字节

32. 专利部分 (patent_section)

  • 特征: instruction, input, output
  • 训练集: 151000个样本,总大小499358509字节
  • 下载大小: 239316958字节

33. 公共 (public)

  • 特征: instruction, input, output, tpye
  • 训练集: 18000个样本,总大小40666888字节
  • 下载大小: 12762114字节

34. 关系 (relationships)

  • 特征: instruction, input, output
  • 训练集: 80022个样本,总大小45706612字节
  • 下载大小: 24000637字节

35. 购物 (shopping)

  • 特征: instruction, input, output
  • 训练集: 29586个样本,总大小17079513字节
  • 下载大小: 9159776字节

36. 社会科学 (social_science)

  • 特征: instruction, input, output
  • 训练集: 129870个样本,总大小186311981字节
  • 下载大小: 96285745字节

37. 演讲 (speech)

  • 特征: instruction, input, output, tpye
  • 训练集: 72000个样本,总大小162899290字节
  • 下载大小: 48896868字节

38. 技术科学 (technology_science)

  • 特征: instruction, input, output
  • 训练集: 26907个样本,总大小37930287字节
  • 下载大小: 19950147字节

39. Wisenut (wisenut)

  • 特征: instruction, input, title, output, lenght
  • 训练集: 228728个样本,总大小440353415字节
  • 下载大小: 145508702字节

以上数据集提供了丰富的文本数据,适用于多种研究和应用场景。

搜集汇总
数据集介绍
wisenut-nlp-team/llama_ko_smr 数据集图片
构建方式
在自然语言处理领域,高质量文本摘要数据集的构建是提升模型生成能力的关键基石。wisenut-nlp-team/llama_ko_smr 数据集通过整合韩国人工智能中心(AI Hub)多个权威语料库资源,精心构建而成。其数据来源涵盖文档摘要文本、图书资料摘要、论文资料摘要、广播内容脚本摘要、摘要及报告生成数据、韩国语对话摘要以及技术科学摘要数据等七大类别。每个子数据集均采用统一的指令(instruction)、输入(input)与输出(output)三元组结构进行格式化,部分子集还额外包含类型(tpye)或标题(title)等字段,以确保数据的高质量与一致性。整个数据集包含超过40个不同领域的子配置,如法律、报纸、专利、科学等,总计约数百万条样本,为韩国语文本摘要任务提供了丰富多样的训练素材。
特点
该数据集的核心特点在于其极致的领域多样性与规模宏大性。从艺术、文化、历史等人文学科,到生命科学、技术科学、自然科学等理工领域,再到日常对话、购物、教育等生活场景,数据集覆盖了极其广泛的知识范畴。其中,专利子集以超过69亿字节和31.2万条样本的规模独占鳌头,报纸与论文子集也各自拥有超过27万和32万条样本,为模型提供了海量的学习实例。此外,数据集在结构上保持了高度的统一性,所有子集均采用标准的三元组格式,便于研究者直接用于序列到序列模型的训练。部分子集还提供了类型标签,为进行细粒度的摘要风格控制或领域自适应学习提供了可能。这种兼具广度与深度的设计,使得该数据集成为训练通用型韩国语文本摘要模型的理想选择。
使用方法
研究者可通过Hugging Face Datasets库便捷地加载该数据集。使用时,需指定具体的配置名称(config_name)以获取特定领域的子集,例如加载法律子集可调用 load_dataset('wisenut-nlp-team/llama_ko_smr', 'law')。每个样本包含instruction(任务指令)、input(待摘要文本)与output(参考摘要)三个核心字段,可直接用于构建模型训练的输入与标签。对于包含tpye字段的子集,研究者可依据该字段进行数据筛选或条件生成任务的训练。由于所有子集仅提供训练集(train)划分,建议使用者在加载后自行划分验证与测试集。数据加载完成后,可通过标准的PyTorch或TensorFlow数据管道进行批处理,配合分词器将文本转换为模型可接受的输入格式,从而高效地开展韩国语文本摘要模型的微调与评估工作。
背景与挑战
背景概述
随着自然语言处理领域对多任务与跨领域理解能力的日益重视,高质量、大规模的指令微调数据集成为推动语言模型发展的关键基石。在此背景下,由韩国Wisenut NLP团队于2023年前后构建的llama_ko_smr数据集应运而生,其核心研究问题在于如何系统性地整合来自多个权威韩语语料库的摘要与生成任务,以提升模型在多样化场景下的文本理解和生成能力。该数据集汇集了韩国人工智能中心(AI Hub)的九大公开项目资源,涵盖法律、新闻、学术论文、广播剧本、对话记录及科技文献等逾38个子领域,总计超过200万条样本。其影响力体现在为韩语大语言模型的指令跟随与摘要生成提供了统一的训练基准,显著促进了相关领域研究的进展。
当前挑战
该数据集所面临的挑战首先体现在领域覆盖的广度与深度平衡上。由于整合了法律、科技、人文等跨度极大的文本类型,不同子集在语言风格、专业术语及摘要长度上存在显著差异,模型需具备适应多元语域的能力,这对泛化性提出了极高要求。构建过程中,数据来源于异构的原始语料库,如AI Hub的文档摘要、图书摘要、论文摘要及广播脚本等,各来源的数据格式、标注规范和噪声水平不一,统一清洗、对齐与质量控制成为瓶颈。此外,部分子集如专利和报纸数据量庞大(分别达312k与274k条),而生命科学等子集仅7.8k条,长尾分布可能导致模型对低频领域的学习不足,易产生偏见或遗忘现象。
常用场景
经典使用场景
在自然语言处理与文本摘要研究的交汇点上,llama_ko_smr数据集凭借其跨越法律、新闻、论文、专利、广播剧本及对话等多领域的庞大语料库,成为训练和评估韩语抽象式与抽取式摘要模型的经典基准。其子集涵盖从27.3k的法律文本到324k的学术论文摘要,研究者可针对不同文本类型精细调优模型,尤其适用于探索跨领域泛化能力与长文本压缩策略。
衍生相关工作
基于该数据集,学术界衍生了一系列经典工作,包括基于BART和T5的韩语摘要模型微调框架、融合领域知识的提示学习摘要方法,以及面向长文档的分层编码-解码架构。部分研究还将其与多语言摘要数据集结合,探索跨语言知识迁移。此外,数据集的多样性子集催生了针对特定领域(如法律、科技)的专用摘要模型,并推动了韩语摘要评测基准的标准化进程。
数据集最近研究
最新研究方向
该数据集聚焦于韩语文本摘要任务的指令微调,整合了来自韩国AI Hub的多个高质量子集,涵盖法律、新闻、论文、专利、广播脚本、对话及科技文献等多元领域。当前前沿研究方向集中于利用此类大规模、多领域指令数据提升大型语言模型在韩语环境下的摘要生成能力与领域适应性,尤其关注零样本泛化与跨领域迁移学习。随着生成式AI在信息压缩与知识管理中的广泛应用,该数据集为构建能够精准提炼长文档、专利及学术论文核心内容的韩语模型提供了关键训练资源,对推动韩语自然语言处理技术落地具有显著意义。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务