遇见数据集

Extracted CU features and their descriptions.

收藏
Figshare2026-02-06 更新2026-04-28 收录
官方服务:

资源简介:

In Versatile Video Coding (VVC), the partition patterns for coding units (CUs) have significant impact on the encoding efficiency. Determining the optimal CU partition is particularly time-consuming due to the calculation and comparison of rate-distortion costs for all possible partition patterns, especially during the ternary tree (TT) partitioning in intra coding. In this paper, a fast decision mechanism is proposed for TT partitioning based on image feature analysis to skip the complex rate-distortion calculation. Firstly, the correlation between the image structural features and the TT partition patterns is investigated based on experimental analysis and the most relevant features are selected for the subsequent prediction of optimal TT partition patterns. Secondly, we devise an efficient scheme for representing and extracting the selected features, further optimizing the extraction process to minimize computational complexity. Comprehensive datasets for partition pattern prediction are constructed based on these refined features. Finally, these datasets serve as the foundation for training and optimizing a predictive model, which is designed to achieve an optimal trade-off between prediction accuracy and model complexity. The predictive model is seamlessly incorporated into the VVC Test Model (VTM), facilitating efficient feature extraction prior to the Rate-Distortion Optimization (RDO) process for intra prediction and optimal partition pattern selection. By leveraging the prediction results, the model effectively determines whether TT partitioning can be bypassed, thereby streamlining the decision-making process and enhancing overall coding efficiency. Experimental results demonstrate that in comprehensive performance evaluations of time-saving metrics and Bjøntegaard Delta Bit Rate (BDBR), the proposed mechanism significantly outperforms existing lightweight neural network algorithms. Our decision mechanism effectively preserves coding quality while substantially accelerating the video coding process.

在通用视频编码(Versatile Video Coding, VVC)中,编码单元(coding units, CUs)的划分模式对编码效率具有显著影响。由于需要对所有可能的划分模式计算并比较率失真代价,确定最优CU划分的过程耗时极长,尤其在帧内编码的三元树(ternary tree, TT)划分阶段。本文提出一种基于图像特征分析的TT划分快速决策机制,以跳过复杂的率失真计算步骤。首先,本文通过实验分析研究了图像结构特征与TT划分模式之间的相关性,并筛选出最相关的特征用于后续最优TT划分模式的预测。其次,本文设计了高效的所选特征表示与提取方案,进一步优化提取流程以降低计算复杂度;基于这些精细化特征构建了划分模式预测的综合数据集。最后,以该数据集为基础训练并优化预测模型,旨在实现预测精度与模型复杂度之间的最优权衡。该预测模型被无缝集成至VVC测试模型(VVC Test Model, VTM)中,可在帧内预测与最优划分模式选择的率失真优化(Rate-Distortion Optimization, RDO)流程前高效完成特征提取。通过利用预测结果,该模型可有效判定是否可跳过TT划分,从而简化决策流程并提升整体编码效率。实验结果表明,在耗时指标与本特加尔德尔比特率(Bjøntegaard Delta Bit Rate, BDBR)的综合性能评估中,所提机制的表现显著优于现有轻量级神经网络算法。本文提出的决策机制在有效保留编码质量的同时,大幅加速了视频编码流程。

创建时间:
2026-02-06
二维码
社区交流群
二维码
科研交流群
商业服务