遇见数据集

Italian TV Series Sample - On-Screen Gender Metrics

收藏
Zenodo2026-09-29 更新2026-10-01 收录
官方服务:

资源简介:

This dataset documents gender-related on-screen measures for the Italian TV Series Sample (https://doi.org/10.5281/zenodo.17191258). Episodes are identified through the same naming convention used in the sample, which combines broadcast year, distributor category (Rai, Mediaset, or Other), genre classification (Soap Opera, Drama, or Comedy), IMDb code, series title, season and episode number. The two files can therefore be joined directly, and the sample weighting coefficients can be applied to the on-screen measures. Off-screen cast and crew credits for the same sample are available in another record (https://doi.org/10.5281/zenodo.19388435). For each episode, the dataset reports female and male speaking time (raw values in seconds and percentage share of total speech), female and male face occurrences (raw counts and percentage share of total detected faces), and the estimated mean age of female and male faces. Speech measures were extracted with inaSpeechSegmenter (Doukhan et al., 2018), an open-source framework that segments audio into speech, music and noise and classifies speech segments by speaker gender. Face measures and age estimates were extracted with inaFaceAnalyzer (Doukhan et al., 2024), an open-source toolkit for face detection and gender and age classification in video, applied at a sampling rate of one frame per second. Both tools infer a binary gender category from acoustic or visual features, and the resulting values should therefore be read as automated estimates of perceived gender rather than as self-declared identities. Here is a detailed description of all variables (columns) contained in the dataset: episode_name: unique episode identifier, shared with the Italian TV Series Sample. It is formed by broadcast year, distributor category (R = Rai, M = Mediaset, O = Other), genre (S = Soap Opera, D = Drama, C = Comedy), IMDb episode code, series title, season number and episode number, separated by dots (e.g. 2000.M.C.tt1656930.Casa_Vianello.8.8). female_speech_raw: total speaking time attributed to female voices, in seconds. male_speech_raw: total speaking time attributed to male voices, in seconds. total_speech_raw: total speaking time detected in the episode (female + male), in seconds. female_speech_percent: share of total speaking time attributed to female voices (%). male_speech_percent: share of total speaking time attributed to male voices (%). female_face_raw: number of face occurrences classified as female, counted on frames sampled at one frame per second. male_face_raw: number of face occurrences classified as male, counted on frames sampled at one frame per second. total_face_raw: total number of face occurrences detected in the episode (female + male). female_face_percent: share of detected face occurrences classified as female (%). male_face_percent: share of detected face occurrences classified as male (%). female_mean_age: mean estimated age, in years, of face occurrences classified as female. male_mean_age: mean estimated age, in years, of face occurrences classified as male. Technical Note: This CSV file uses UTF-8 encoding to preserve Italian characters and special symbols. If characters appear incorrectly when opening the file (e.g., "è" displays as "è"), please verify that your software is configured to read UTF-8 encoding. Questions or Issues? For inquiries about this dataset, please use the "Contact" button on this Zenodo record or open a discussion in the Italian TV Series Sample community. References Doukhan, D., Carrive, J., Vallet, F., Larcher, A., & Meignier, S. (2018). An open-source speaker gender detection framework for monitoring gender equality. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 5214–5218). IEEE. Doukhan, D., Dodson, L., Conan, M., Pelloin, V., Clamouse, A., Lepape, M., ... & Coulomb-Gully, M. (2024). Gender representation in TV and radio: Automatic information extraction methods versus manual analyses. arXiv preprint arXiv:2406.10316.

提供机构:
Zenodo
创建时间:
2026-09-29
二维码
社区交流群
二维码
科研交流群
商业服务