遇见数据集

Audiovisual Analysis of Journal Digital

收藏
Zenodo2026-01-02 更新2026-05-26 收录
官方服务:

资源简介:

Automated extraction and analysis of audio-visual co-occurrences in video data using deep learning models. ## Overview This project analyzes temporal co-occurrences of audio and visual features in videos: - Audio extraction : Uses MIT's Audio Spectrogram Transformer (AST) to classify audio content - Visual extraction : Uses Moondream vision model for object detection in frames - Pairing analysis : Identifies audio-visual co-occurrences with temporal alignment - Temporal analysis : Tracks changes in audio-visual patterns across years ## Structure - src.zip : The source code for the library - results.zip : The .csv outputs - moon_results.tar.gz : Results from the Visual extraction pipeline - audio_results.tar.gz : Results forom the Audio extraction pipeline

提供机构:
Zenodo
创建时间:
2026-01-02
二维码
社区交流群
二维码
科研交流群
商业服务