官方服务:
资源简介:
Shortcut list of the tight GPU kernel.
应用场景:
创建时间:
2015-12-02
相关数据集
深度学习交错共享及调度技术的测试数据集
本数据集包括深度学习交错共享及调度技术测试所需的核心任务轨迹数据和相关测试代码。任务轨迹数据涵盖深度学习任务的关键属性,如开始时间、持续时间、模型类型、批大小和GPU使用数量等。测试数据集包括测试用任务轨迹。测试用任务轨迹来源于微软公开的Philly Trace数据集,包含了任务开始时间、任务结束时间、任务模型类型、任务使用的GPU卡数、模型执行信息等在内的测试用任务轨迹。数据量 2MB
国家基础学科公共科学数据中心110
Accelerate the Parameterization of Unified Microphysics Across Scales (PUMAS) on the graphics processing unit (GPU) with directive-based methods
Code used to produce results of paper titled "Accelerate the Parameterization of Unified Microphysics Across Scales (PUMAS) on the graphics processing unit (GPU) with directive-based methods" by Sun e
NIAID Data Ecosystem80
Optimizing Sparse Matrix-Matrix Multiplication for the GPU supplementary data
This contains the matrices for the SpGEMM tests presented in "Optimizing Sparse Matrix-Matrix Multiplication for the GPU", by Steven Dalton, Nathan Bell, and Luke N. Olson. Each A matrix from Table
NIAID Data Ecosystem90
Full-DPP Wave64 Primitives for CDNA3: Eliminating LDS-Mediated Cross-Lane Communication on AMD Instinct MI300X
Research artifacts for the paper: Eliminating LDS-Mediated Cross-Lane Communication on CDNA3. Includes header-only HIP library (wave_reduce_dpp, wave_scan_dpp), Flash Attention kernels (fa_naive, fa_d
Zenodo2026-06-15 更新70



