遇见数据集

2016 Imageclef Webupv Collection

收藏
Zenodo2020-09-20 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

===============================================================================<br> Changelog:<br> 2018-05-05 Test data ground truth released.<br> <br> 2016-04-24 Test data for main subtask 3 is released. 2016-03-17 Test data for teaser tasks is released. 2016-03-16 A Bug was found in the Visual Feature files, please redownload if you use them 2016-02-17 Fixed the mixxing xml files in scaleconcept16_data_textual.webpages.tar.gz 2016-02-15 Added input document files for the development set of the teaser tasks.<br> * DevData/TeaserTasks/scaleconcept16.teaser_dev_input_documents.tar.gz 2016-02-15 Fixed some newline formatting issues in<br> * Features/scaleconcept16.teaser.TrainTestSplit.v20160215.tar.gz<br> Please download the latest version. 2016-02-12 Teaser development set: There were 2 duplicate webpages and 2<br> near-duplicate images in the previous release. Based on user<br> feedback, we have updated the dataset -- which now only includes<br> 3337 image-webpage pairs. Please download the latest versions of<br> the following to reflect these minor updates.<br> * DevData/TeaserTasks/scaleconcept16.teaser_dev_id.v20160212.tar.gz<br> * DevData/TeaserTasks/scaleconcept16.teaser_dev_data_textual.scofeat.v20160212.gz<br> * DevData/TeaserTasks/scaleconcept16.teaser2_dev_groundtruth.v20160212.txt<br> =============================================================================== This document describes the ScaleConcept dataset compiled for the ImageCLEF 2016<br> Scalable Concept Image Annotation challenge. The data mentioned here indicates<br> what is ready for download. However, upon request or depending on feedback<br> from the participants, additional data may be released. The following is the directory structure of the collection, and below there<br> is a brief description of what each compressed file contains. <br> Directory structure<br> ------------------- .<br> |<br> |--- README.txt<br> |--- scaleconcept16.agreement.txt.tar.gz<br> |--- scaleconcept16.concepts.tar.gz<br> |<br> |--- Features/<br> | |<br> | |--- scaleconcept16_ImgID.txt_Mod.tar.gz<br> | |--- scaleconcept16_ImgToTextID.tar.gz<br> | |--- scaleconcept16_TextID.txt_Mod.tar.gz<br> | |--- scaleconcept16.teaser.TrainTestSplit.*.tar.gz <br> | |<br> | |--- Textual/<br> | | |<br> | | |---scaleconcept16_data_textual.scofeat.tar.gz <br> | | |---scaleconcept16_data_textual.webpages.zip <br> | |<br> | |--- Visual/<br> | |<br> | |--- scaleconcept16_data_visual_gist.dfeat.gz<br> | |--- scaleconcept16_data_visual_sift_1000.sfeat.gz<br> | |--- scaleconcept16_data_visual_rgbsift_1000.sfeat.gz<br> | |--- scaleconcept16_data_visual_opponentsift_1000.sfeat.gz<br> | |--- scaleconcept16_data_visual_colorhist.sfeat.gz<br> | |--- scaleconcept16_data_visual_getlf.sfeat.gz<br> | |--- scaleconcept16_data_visual_vgg16-relu7.dfeat.gz<br> | |--- scaleconcept16_images.zip<br> |<br> |--- DevData/<br> | |<br> | |--- MainSubTasks/<br> | | |<br> | | |--- scaleconcept16.dev.visual.bbox.*.tar.gz<br> | | |--- scaleconcept16.dev.textdesc.*.tar.gz<br> | | |--- scaleconcept16.subtask3.dev.input_bbox.*.gz<br> | | |--- scaleconcept16.subtask3.dev.textdesc.*.gz<br> | |<br> | |--- TeaserTasks/<br> | | |<br> | | |--- scaleconcept16.teaser_dev_data_textual.scofeat.*.gz<br> | | |--- scaleconcept16.teaser_dev_data_visual_colorhist.sfeat.gz<br> | | |--- scaleconcept16.teaser_dev_data_visual_csift_1000.sfeat.gz<br> | | |--- scaleconcept16.teaser_dev_data_visual_getlf.sfeat.gz<br> | | |--- scaleconcept16.teaser_dev_data_visual_gist.sfeat.gz<br> | | |--- scaleconcept16.teaser_dev_data_visual_opponentsift_1000.sfeat.gz<br> | | |--- scaleconcept16.teaser_dev_data_visual_rgbsift_1000.sfeat.gz<br> | | |--- scaleconcept16.teaser_dev_data_visual_sift_1000.sfeat.gz<br> | | |--- scaleconcept16.teaser_dev_data_visual_vgg16-relu7.dfeat.gz<br> | | |--- scaleconcept16.teaser_dev_id.*.tar.gz<br> | | |--- scaleconcept16.teaser_dev_images.zip<br> | | |--- scaleconcept16.teaser_dev_pages.zip<br> | | |--- scaleconcept16.teaser2_dev_groundtruth.txt<br> |<br> |--- TestData/<br> | |<br> | |--- concepts.lst<br> | |--- scaleconcept16_subtask1_test.lst<br> | |--- scaleconcept16_subtask2_test.lst<br> | |--- scaleconcept16_subtask3_test.lst<br> | |--- scaleconcept16_subtask3_test.input_bbox.txt<br> | |--- scaleconcept16_teaser1_test_image_collection.lst<br> | |--- scaleconcept16_teaser1_test.lst<br> | |--- scaleconcept16_teaser2_test.lst<br> | |--- scaleconcept16_teaser_test_input_documents.tar.gz <br> Contents of files<br> ----------------- * scaleconcept16.concepts.tar.gz -&gt; scaleconcept16.concepts.txt List of 251 concepts for the 2016 challenge. File format:<br> wordnet-offset \t category-word.pos.## \t list,of,synonyms,separated,by,commas \t defintiion -&gt; scaleconcept16.concepts_hierarchy.txt The hierarchy structure of the 'general level' categories. File format:<br> *category \t *parent-category \t definition.<br> For example, *mammal is the child of *animal. '#' represents the root node. -&gt; scaleconcept16.concepts_to_parents.txt List of 'general level' category parent(s) for each 251 concept. A concept may<br> have multiple parents (separated by commas). File format:<br> category \t *parent1,*parent2 * Features/scaleconcept16_ImgID.txt_Mod.tar.gz IDs of the images in the dataset. <br> * Features/scaleconcept16_TextID.txt_Mod.tar.gz IDs of the webpages in the dataset. <br> * Features/scaleconcept16_ImgToTextID.tar.gz IDs of images that appear on corresponding web pages <br> * Features/scaleconcept16.teaser.TrainTestSplit.*.tar.gz For Teasers 1 and 2: IDs of images and webpages, split into approximately<br> 300K for training and exactly 200K for testing.<br> The 200K test data cannot be explored during training for both teaser tasks. * Features/Textual/scaleconcept16_data_textual.scofeat.tar.gz The processed text extracted from the webpages near where the images<br> appeared. Each line corresponds to one image, having the same order<br> as the data_iids.txt list. The lines start with the image ID,<br> followed by the number of extracted unique words and the<br> corresponding word-score pairs. The scores were derived taking into<br> account 1) the term frequency (TF), 2) the document object model<br> (DOM) attributes, and 3) the word distance to the image. The scores<br> are all integers and for each image the sum of scores is always<br> &lt;=100000 (i.e. it is normalized). <br> * Features/Textual/scaleconcept16_data_textual.webpages.tar.gz Contains all of the webpages which referenced the images in the<br> dataset set after being converted to valid xml. In total there are<br> 525766 files, since each image can appear in more than one page, and<br> there can be several versions of same page which differ by the<br> method of conversion to xml. To avoid having too many files in a<br> single directory (which is an issue for some types of partitions),<br> the files are found in subdirectories named using the first two<br> characters of the RID, thus the paths of the files after extraction<br> are of the form: ./scaleconcept16_data_textual.webpages/{RID:0:2}/{RID}.{CONVM}.xml.gz * Features/Visual/scaleconcept16_images.zip Contains thumbnails (maximum 640 pixels of either width or height)<br> of the images in jpeg format. To avoid having too many files in a<br> single directory (which is an issue for some types of partitions),<br> the files are found in subdirectories named using the first two<br> characters of the image ID, thus the paths of the files after extraction<br> are of the form: ./scaleconcepts16_images/{IID:0:2}/{IID}.jpg <br> * Features/Visual/scaleconcept16_*.{s|d}feat.gz The visual features in a simple ASCII text format either in sparse<br> (*.sfeat.gz files) or dense (*.dfeat.gz files). The first<br> line of the file indicates the number of vectors (N) and the<br> dimensionality (DIMS). Then each line corresponds to one vector.<br> For the dense features each line has exactly DIMS values separated<br> by spaces, i.e., the format is: N DIMS<br> Val(1,1) Val(1,2) ... Val(1,DIMS)<br> Val(2,1) Val(1,2) ... Val(2,DIMS)<br> ...<br> Val(N,1) Val(N,2) ... Val(N,DIMS) For the sparse features, each line starts with the number of non-zero<br> elements and is followed by dimension-value pairs, being the first<br> dimension 0, i.e., the format is: N DIMS<br> nz1 Dim(1,1) Val(1,1) ... Dim(1,nz1) Val(1,nz1)<br> nz2 Dim(2,1) Val(2,1) ... Dim(2,nz2) Val(2,nz2)<br> ...<br> nzN Dim(N,1) Val(N,1) ... Dim(N,nzN) Val(N,nzN) The order of the features is the same as in the list data_iids.txt. The procedure to extract the SIFT based features in this<br> subdirectory was conducted as follows. Using the ImageMagick<br> software, the images were first rescaled to having a maximum of 240<br> pixels, of both width and height, while preserving the original<br> aspect ratio, employing the command: convert {IMGIN}.jpg -resize '240&gt;x240&gt;' {IMGOUT}.jpg Then the SIFT features where extracted using the ColorDescriptor<br> software from Koen van de Sande<br> (http://koen.me/research/colordescriptors). As configuration we<br> used, 'densesampling' detector with default parameters, and a hard<br> assignment codebook using a spatial pyramid as<br> 'pyramid-1x1-2x2'. The number in the file name indicates the size of<br> the codebook. All of the vectors of the spatial pyramid are given in<br> the same line, thus keeping only the first 1/5th of the dimensions<br> would be like not using the spatial pyramid. The codebook was<br> generated using 1.25 million randomly selected features and the<br> k-means algorithm. The GIST features were extracted using the<br> LabelMe Toolbox. The images where first resized to 256x256 ignoring<br> original aspect ratio, using 5 scales, 6 orientations and 4<br> blocks. The other features colorhist and getlf, are both color<br> histogram based extracted using our own implementation. <br> * Features/Visual/scaleconcept16_data_visual_vgg16-relu7.dfeat.gz Contains the 4096 dimensional activations of the relu7 layer of Oxford<br> VGG's 16-layer CNN model, extracted using the Berkeley Caffe library.<br> More details can be found at https://github.com/BVLC/caffe/wiki/Model-Zoo. * DevData/MainSubTasks/scaleconcept16.dev.viusal.bbox.*.tar.gz Development set ground truth localised annotations for sub task 1. The format for the development set of annotated bounding boxes of<br> the concepts is &lt;image_ID&gt; &lt;seq&gt; &lt;Concept&gt; &lt;confidence&gt; &lt;xmin&gt; &lt;ymin&gt; &lt;xmax&gt; &lt;ymax&gt; The development set contains 1,979 images. The bounding boxes may enclose<br> single instances (a single tree) or grouped instances (e.g. group of trees),<br> depending on the context. The annotations are not exhaustive: the emphasis<br> is on concepts that are interesting enough to be described in the image,<br> although background objects are also optionally annotated by our annotators<br> in many cases. Also note that a person might not be annotated if the<br> annotator could not decide whether the person is a man/woman/boy/girl. <br> * DevData/MainSubTasks/scaleconcept16.dev.textdesc.*.tar.gz Development set ground truth textual description annotations of images for<br> Subtask 2 The format is: &lt;image_ID&gt; \t &lt;text_description_seq&gt; \t &lt;textual_description&gt; The development set contains 2,000 images with 5 to 51 textual descriptions<br> per image (mean: 9.492, median: 8). Please note that the sentences contain a<br> mix of both American and British English spelling variants (e.g. color vs<br> colour) -- we have decided to retain this variation in the annotations to<br> reflect the challenge of real-world English spelling variants. Basic<br> spell-correction has been performed on the textual descriptions, but we cannot<br> guarantee that they are completely free from spelling or grammatical error. <br> * DevData/MainSubTasks/scaleconcept16.subtask3.dev.input_bbox.*.gz Input bounding boxes for Subtask 3. This is a selected subset of 500 development<br> images from scaleconcept16.dev.visual.bbox above (please refer to above for file format). <br> * DevData/MainSubTasks/scaleconcept16.subtask3.dev.textdesc.*.gz Annotated textual descriptions for 500 development images, to be used to evaluate<br> the content selection ability of the text generation system in the clean track of<br> SubTask 3. The format is the same as the original scaleconcept16.dev.textdesc<br> file, except that we further annotated textual terms with their corresponding<br> input bounding boxes, for example [[[dogs|0,4]]] in a textual description refers<br> to the two instances of dogs with the bounding box id 0 and 4 in<br> scaleconcept16.subtask3.dev.input_bbox. Note that not all descriptions from the original scaleconcept16.dev.textdesc are<br> used in this version, and as such the sequence numbers of the descriptions may not<br> necessarily be contiguous as we retained the sequence numbers from the original<br> file for consistency. * DevData/TeaserTasks/scaleconcept16.teaser_dev_id.*.tar.gz -&gt; scaleconcept16.teaser_dev.ImgID.txt<br> IDs of 3339 images for the development set of both teaser tasks. Note that 2 images<br> are near-duplicates and will thus not be used in this dataset. We have left them<br> intact to avoid having participants re-download the visual features. -&gt; scaleconcept16.teaser_dev.TextID.txt<br> IDs of 3337 webpage documents for the development set of both teaser tasks. -&gt; scaleconcept16.teaser_dev.ImgToTextID.txt<br> IDs of 3337 images that appear on corresponding web pages. <br> * DevData/TeaserTasks/scaleconcept16.teaser_dev_input_documents.tar.gz -&gt; scaleconcept16.teaser_dev.docID.txt<br> IDs of 3337 input text documents for the development set of both teaser tasks. -&gt; docs/{docID}<br> The text for each input document. These should be used as input for both teaser tasks.<br> * DevData/TeaserTasks/scaleconcept16.teaser_dev_images.zip Contains 3339 images for the development set of both teaser tasks in jpeg format. <br> * DevData/TeaserTasks/scaleconcept16.teaser_dev_pages.zip Contains 3337 webpages for the development set of both teaser tasks after<br> converting to valid xml format, each compressed as a gzip file. <br> * DevData/TeaserTasks/scaleconcept16.teaser_dev_data_textual.scofeat.*.gz The processed text extracted from 3337 webpages near where the images appeared.<br> Please refer to Features/Textual/scaleconcept16_data_textual.scofeat.tar.gz<br> above for more details. <br> * DevData/TeaserTasks/scaleconcept16_teaser_dev_data_visual_*.{s|d}feat.gz The visual features for the 3339 images for the development set of both teaser tasks.<br> Please refer to Features/Visual/scaleconcept16_*.{s|d}feat.gz above for more details. <br> * DevData/TeaserTasks/scaleconcept16.teaser2_dev_groundtruth.*.txt The GPS coordinates for 3337 documents from the development set for Teaser Task 2 (Geolocation).<br> File format:<br> WebpageID latitude longitude <br> * TestData/concepts.lst List of 251 concepts for the main subtasks 1, 2, and 3.<br> File format: wordnet-offset \t category-word.pos.## <br> * TestData/scaleconcept16_subtask{1|2}_test.lst<br> <br> List of 510,123 images to annotate for main subtasks 1 and 2. Both files are identical. <br> * TestData/scaleconcept16_subtask3_test.lst<br> <br> List of 450 images to annotate for main subtask 3. <br> * TestData/scaleconcept16_subtask3_test.input_bbox.txt List of bounding boxes for 450 test images, to be used as input for main subtask 3.<br> The format is the same as the development set. <br> * TestData/scaleconcept16_teaser1_test_image_collection.lst The collection of 200,000 test images to be used for Teaser Task 1 (text illustration) <br> * TestData/scaleconcept16_teaser1_test.lst The list of IDs for 180,000 text documents to be used as input for Teaser Task 1.<br> Please note that the IDs are *not* the same as the webpage IDs provided in the 500K corpus.<br> The task is to provide a ranked list of the top 100 images (from the 200,000 test image collection above)<br> for each input text document. <br> * TestData/scaleconcept16_teaser2_test.lst The list of IDs for 180,000 text documents to be used as input for Teaser Task 2 (identical to teaser task 1).<br> The task is to provide the latitude and longitude for each input text document. <br> * TestData/scaleconcept16_teaser_test_input_documents.tar.gz The input text documents to be used for Teaser Tasks 1 and 2.<br> These should be used as input for both teaser tasks. <br> * TestData/scaleconcept16_groundtruth.zip Ground truth for the test set. Contact<br> ------- For further questions, please contact:<br> Andrew Gilbert &lt;a.gilbert@surrey.ac.uk&gt;

提供机构:
Zenodo
创建时间:
2018-05-05
二维码
社区交流群
二维码
科研交流群
商业服务