遇见数据集

On the Importance of Design Choices in Transformer Decoder-Only Large Language Models

收藏
Zenodo2026-05-26 更新2026-06-05 收录
官方服务:

资源简介:

This repository contains raw data and experimental results. At the top level, you will find three folders: data/ – Contains the core datasets used in the project, including both raw and cleaned versions. results/ – Contains experimental results and their visualization. data/ df_long_accuracy.csvContains structured meta-features and evaluation results for large language models across multiple benchmarks and evaluation settings. Each row corresponds to a model–dataset–evaluation setting combination, linking scale, architectural, and training meta-features with an observed performance value. df_long_accuracy_final.csvThe cleaned and preprocessed version of the dataset. The notebooks used to generate this file can be found in the notebooks/ directory on the software link below. This file is used in downstream experiments. Other files The remaining files in this directory consist of visualizations derived from the above data and generated using the notebooks. results/ uncorr_15 importance.csvContains meta-feature importance results generated using fANOVA. Each row represents the importance of a single meta-feature or pair of meta-features for a given task (evaluation setting–dataset combination). marginals/Contains marginal plots for each of the meta-features and individual tasks. clustering/Contains clustering results of tasks based on fANOVA embeddings. scatterplots_raw/Contains scatterplots of raw meta-features vs accuracy. Other filesThe remaining files in this directory consist of visualizations derived from the above data and generated using the notebooks. all_18 Contains the results for the full feature portfolio. Software: https://anonymous.4open.science/r/llm-fanova-analysis-v2-FAE1/

提供机构:
Zenodo
创建时间:
2026-05-26
二维码
社区交流群
二维码
科研交流群
商业服务