witcheer/rtx-5090-benchmarks
收藏资源简介:
该数据集是NVIDIA RTX 5090 LLM基准测试的集合,专注于在NVIDIA RTX 5090 32GB GPU上对量化大型语言模型(LLM)进行速度和质量的评估。数据集包含质量基准测试(如MMLU、ARC-Challenge、HellaSwag、GSM8K和HumanEval),使用自定义评估器通过llama-server聊天完成进行生成式评估,并采用分层抽样(种子=42)和禁用思考模式。速度基准测试测量提示处理(并行批处理令牌吞吐量,上下文长度从128到16384)和文本生成(顺序自回归令牌吞吐量,128令牌),所有模型完全GPU卸载。数据集还包括硬件配置(如GPU、CPU、RAM)和工具(llm-bench-rig)的详细信息,以及关键发现(如MoE模型与密集模型的性能比较)。数据以CSV文件(benchmarks.csv)形式提供,适用于模型比较和推理优化研究。
This dataset is a collection of NVIDIA RTX 5090 LLM Benchmarks, focusing on speed and quality evaluations of quantized large language models (LLMs) on an NVIDIA RTX 5090 32GB GPU. It includes quality benchmarks (e.g., MMLU, ARC-Challenge, HellaSwag, GSM8K, and HumanEval) conducted via generative evaluation through llama-server chat completions using custom evaluators, with 50% stratified sampling (seed=42) and thinking disabled. Speed benchmarks measure prompt processing (parallel batched token throughput at context lengths from 128 to 16384) and text generation (sequential autoregressive token throughput at 128 tokens), with all models fully GPU-offloaded. The dataset also provides hardware specifications (e.g., GPU, CPU, RAM) and tooling details (llm-bench-rig), along with key findings (such as performance comparisons between MoE and dense models). Data is provided in CSV format (benchmarks.csv), suitable for model comparison and inference optimization research.




