Multimodal C4 (mmc4)是由University of California, Santa Barbara等机构开发的开放式十亿规模图像与文本交错数据集。该数据集通过线性分配算法,利用CLIP特征将图像嵌入到c4文本数据集中,覆盖烹饪、旅行、技术等多个日常主题。经过过滤不适当内容后,数据集包含1.012亿文档,其中交错嵌入了5.71亿图像和430亿英语词汇。mmc4旨在支持复杂的视
This dataset contains the data collected from romantic partners on Prolific for our research project on empathic accuracy and interpersonal emotion regulation. All identifying information was removed
Without formal avenues for claims-making or political participation, refugees must find their own means of securing services from state and non-state providers. This paper asks why some refugee commun
Survival and height growth records for individual plants of Quercus agrifolia (n=4281) and Q. lobata (n=4268) from a long term experiment in mixed oak savanna at Sedgwick Reserve, Santa Barbara County