--- pretty_name: Evaluation run of bigcode/starcoder2-7b dataset_summary: "Dataset automatically created during the evaluation run of model\ \ [bigcode/starcoder2-7b](https://huggingface.co/bigcode/
This dataset was created to provide comparative statistics on various general-purpose and pre-trained financial large language models (LLMs) based on their financial capabilities. We also emphasize on
This repository archives the official replication dataset, processing source code, and empirical assets for the paper titled "The Automated Consensus: A Hierarchical Evaluation of Task-Sensitive Seman
--- pretty_name: Evaluation run of jisukim8873/mistral-7B-alpaca-case-0-2 dataset_summary: "Dataset automatically created during the evaluation run of model\ \ [jisukim8873/mistral-7B-alpaca-case-0-
--- pretty_name: Evaluation run of Qwen/Qwen2.5-Math-7B dataset_summary: "Dataset automatically created during the evaluation run of model\ \ [Qwen/Qwen2.5-Math-7B](https://huggingface.co/Qwen/Qwen2