YuvrajSingh9886/jetson-non-reasoning-benchmark-maxn
收藏资源简介:
该数据集是一个名为Jetson非推理基准测试—MAXN超级模式(nvpmodel模式2)的基准测试数据集,专门用于评估在NVIDIA Jetson Orin Nano 8GB边缘AI设备上运行的小型大语言模型(LLM)的推理性能。数据集包含多个模型(如gemma3-1b、lfm2.5-1.2b、llama3.2-1b、qwen2.5-0.5b、qwen3-0.6b等)在不同量化设置(如Q4_K_M、Q8_0)下的性能指标,测试了不同的输入序列长度(128、512、1024、2048词元)和生成序列长度(64、128、256词元)。性能指标包括首词元时间(TTFT)、词元到词元时间(T2T)、输入延迟(ITL)、每秒词元数(Tok/s)、请求延迟、预填充词元速度、功耗(瓦特)以及每焦耳词元数(Tok/J)等。数据集旨在为边缘AI场景中的LLM推理提供详细的性能基准,帮助优化模型部署和能效。
This dataset is a benchmark dataset named Jetson Non-Inference Benchmark — MAXN Super Mode (nvpmodel Mode 2), specifically designed to evaluate the inference performance of small Large Language Models (LLMs) running on the NVIDIA Jetson Orin Nano 8GB edge AI device. It includes performance metrics of multiple models (e.g., gemma3-1b, lfm2.5-1.2b, llama3.2-1b, qwen2.5-0.5b, qwen3-0.6b, etc.) under different quantization settings (e.g., Q4_K_M, Q8_0), with tests covering varying input sequence lengths (128, 512, 1024, 2048 tokens) and output sequence lengths (64, 128, 256 tokens). Performance metrics include Time to First Token (TTFT), Token-to-Token (T2T) latency, Input Latency (ITL), Tokens per Second (Tok/s), request latency, prefill token throughput, power consumption (in Watts), and Tokens per Joule (Tok/J), among others. This dataset aims to provide detailed performance benchmarks for LLM inference in edge AI scenarios, assisting with model deployment optimization and energy efficiency improvement.




