Vocabulary Tests LLMs
收藏资源简介:
Vocabulary evaluation of LLMs This repository contains the results for the different vocabulary tests run on LLM tools/models presented in the paper: "The continued usefulness of vocabulary tests for evaluating large language models" currently published in PLOS ONE: https://doi.org/10.1371/journal.pone.0308259 The name of the files correspond to the vocabulary tests for which results are presented in Tables 3-6 in the paper (note that questions for the TOEFL test are not public). In each file, the first column has the question posed to the LLM tool/model, followed by the correct answer. The rest of the columns correspond to the answers of the different LLM tools/models evaluated. The results and percentages are summarized at the bottom of the file after the last test items. The models evaluated are: Model Link Llama 2 7b https://huggingface.co/meta-llama/Llama-2-7b-chat Llama 2 13b https://huggingface.co/meta-llama/Llama-2-13b-chat Llama 2 70b https://huggingface.co/meta-llama/Llama-2-70b-chat Mistral 7b v0.1 https://huggingface.co/mistralai/Mistral-7B-v0.1 GPT 3.5 turbo 0613 https://platform.openai.com/docs/models/gpt-3-5-turbo GPT 4 0613 https://platform.openai.com/docs/models/gpt-4-turbo-and-gpt-4 Bard To cite our work: @article{10.1371/journal.pone.0308259, doi = {10.1371/journal.pone.0308259}, author = {Martínez, Gonzalo AND Conde, Javier AND Merino-Gómez, Elena AND Bermúdez-Margaretto, Beatriz AND Hernández, José Alberto AND Reviriego, Pedro AND Brysbaert, Marc}, journal = {PLOS ONE}, publisher = {Public Library of Science}, title = {Establishing vocabulary tests as a benchmark for evaluating large language models}, year = {2024}, month = {12}, volume = {19}, url = {https://doi.org/10.1371/journal.pone.0308259}, pages = {1-17}, number = {12}, }
大语言模型词汇评测 本仓库收录了现已正式发表于《PLOS ONE》期刊的论文《用于评估大语言模型的词汇测试的持续应用价值》(DOI:https://doi.org/10.1371/journal.pone.0308259)中,针对多款大语言模型(Large Language Model,以下简称LLM)工具与模型开展的各类词汇测试的结果。 数据集文件的命名与论文中表3至表6所呈现的词汇测试项目一一对应(需注意:托福(TOEFL)测试的试题未对外公开)。 每个文件的第一列为向待测大语言模型工具/模型提出的测试问题,紧随其后的列为正确答案,剩余列则为各被评估大语言模型工具/模型所输出的作答结果。在每个文件的最后一道测试题之后,会汇总本次测试的整体结果与正确率数据。 本次评估的模型及对应链接如下: 1. Llama 2 7b:https://huggingface.co/meta-llama/Llama-2-7b-chat 2. Llama 2 13b:https://huggingface.co/meta-llama/Llama-2-13b-chat 3. Llama 2 70b:https://huggingface.co/meta-llama/Llama-2-70b-chat 4. Mistral 7b v0.1:https://huggingface.co/mistralai/Mistral-7B-v0.1 5. GPT 3.5 turbo 0613:https://platform.openai.com/docs/models/gpt-3-5-turbo 6. GPT 4 0613:https://platform.openai.com/docs/models/gpt-4-turbo-and-gpt-4 7. Bard:未提供公开链接 若需引用本研究,请采用以下参考文献格式: @article{10.1371/journal.pone.0308259, doi = {10.1371/journal.pone.0308259}, author = {Martínez, Gonzalo AND Conde, Javier AND Merino-Gómez, Elena AND Bermúdez-Margaretto, Beatriz AND Hernández, José Alberto AND Reviriego, Pedro AND Brysbaert, Marc}, journal = {PLOS ONE}, publisher = {Public Library of Science}, title = {Establishing vocabulary tests as a benchmark for evaluating large language models}, year = {2024}, month = {12}, volume = {19}, url = {https://doi.org/10.1371/journal.pone.0308259}, pages = {1-17}, number = {12}, }



