european-gender-drama-jina-embeddings
收藏资源简介:
该数据集包含83,000个训练样本,每个样本代表一个戏剧或类似表演中的语音片段。每个样本具有结构化信息,包括说话者身份(speaker)、所属戏剧或剧目名称(play)、说话者性别(gender)、语音片段的文本内容(speech_chunk)、唯一标识符(unique_id)、文本的向量化嵌入表示(embedding,为浮点数列表)、语言信息(language)以及年份(year,为整数)。数据集总大小约为842MB,以训练集形式组织,适用于戏剧文本分析、说话者属性研究、自然语言处理嵌入应用或跨年份/语言的语言特征分析相关任务。
This dataset contains 83,000 training samples, where each sample represents a speech segment from a play or similar performance. Each sample includes structured information covering speaker identity, the name of the associated play or performance, speaker gender, textual content of the speech segment (speech_chunk), unique identifier (unique_id), vectorized embedding representation of the text (embedding, a list of floating-point numbers), language information (language), and year (year, an integer). The total size of the dataset is approximately 842 MB, and it is organized as a training set. It is suitable for tasks including dramatic text analysis, speaker attribute research, natural language processing embedding applications, or linguistic feature analysis across years and languages.




