BCCard/bc-finance-llm-benchmark
收藏资源简介:
BC Card Finance LLM Benchmark是一个专门用于评估韩国金融领域大型语言模型(LLM)性能的基准数据集。该数据集由BC卡与延世大学DSL通过产学合作项目(S2026 LLMOps项目)共同开发。数据集包含800个韩语问答对,涵盖金融领域,特别是BC卡FAQ和通用金融主题。数据格式为问题与真实答案对,包含序号、ID、大类分类(如BC卡FAQ、通用等)、金融细分主题、问题类型(如单一推理、多重推理等)、问题文本、真实答案和来源标签等列。该数据集主要用于LLM-as-Judge自动评估流程,通过比较模型生成的回答与提供的真实答案来评估模型性能。数据集采用Apache 2.0许可证。
BC Card Finance LLM Benchmark is a benchmark dataset specifically designed for evaluating the performance of large language models (LLMs) in the Korean financial domain. It is an outcome of the industry-academia collaboration between BC Card and Yonsei University DSL (S2026 LLMOps project). The dataset consists of 800 Korean question-answer pairs covering the financial domain, particularly BC Card FAQs and general finance topics. The format is Question/Ground Truth pairs, with columns including serial number, question ID, broad category (e.g., BC Card FAQ, general), financial sub-topic, question type (e.g., single reasoning, multiple reasoning), question text, ground truth (reference answer), and source tag. It is intended for use in LLM-as-Judge automatic evaluation pipelines to compare model responses against the ground truth. The dataset is licensed under Apache 2.0.




