遇见数据集

MFCCs Feature Scaling Images for Multi-class Human Action Analysis : A Benchmark Dataset

收藏
Mendeley Data2024-03-27 更新2024-06-28 收录
官方服务:

资源简介:

his dataset comprises an array of Mel Frequency Cepstral Coefficients (MFCCs) that have undergone feature scaling, representing a variety of human actions. Feature scaling, or data normalization, is a preprocessing technique used to standardize the range of features in the dataset. For MFCCs, this process helps ensure all coefficients contribute equally to the learning process, preventing features with larger scales from overshadowing those with smaller scales. In this dataset, the audio signals correspond to diverse human actions such as walking, running, jumping, and dancing. The MFCCs are calculated via a series of signal processing stages, which capture key characteristics of the audio signal in a manner that closely aligns with human auditory perception. The coefficients are then standardized or scaled using methods such as MinMax Scaling or Standardization, thereby normalizing their range. Each normalized MFCC vector corresponds to a segment of the audio signal. The dataset is meticulously designed for tasks including human action recognition, classification, segmentation, and detection based on auditory cues. It serves as an essential resource for training and evaluating machine learning models focused on interpreting human actions from audio signals. This dataset proves particularly beneficial for researchers and practitioners in fields such as signal processing, computer vision, and machine learning, who aim to craft algorithms for human action analysis leveraging audio signals.

本数据集包含经过特征缩放处理的梅尔频率倒谱系数(Mel Frequency Cepstral Coefficients,MFCCs)集合,可表征多样化的人类行为动作。特征缩放(亦称数据归一化)是一种数据预处理技术,用于将数据集中各特征的取值范围统一标准化。针对梅尔频率倒谱系数而言,该处理步骤可确保所有系数在模型学习过程中发挥均等的作用,避免取值范围更大的特征掩盖取值范围更小的特征。本数据集对应的音频信号源自多样化的人类行为动作,包括行走、奔跑、跳跃与舞蹈。该数据集的梅尔频率倒谱系数通过一系列信号处理步骤计算得到,能够以贴合人类听觉感知的方式捕捉音频信号的关键特征。随后通过最小-最大缩放(MinMax Scaling)或标准化(Standardization)等方法对这些系数进行标准化或缩放处理,从而统一其取值范围。每个经归一化处理的梅尔频率倒谱系数向量均对应一段音频信号片段。本数据集经过精心设计,可用于基于听觉线索的人类动作识别、分类、分割与检测等任务。其为训练与评估基于音频信号解读人类行为动作的机器学习模型提供了核心资源。本数据集尤其有益于信号处理、计算机视觉与机器学习等领域的研究人员与从业者,助力其开发基于音频信号的人类动作分析算法。

创建时间:
2024-01-23
二维码
社区交流群
二维码
科研交流群
商业服务