Overhead Depth Images People Detection (GOTPD)
收藏资源简介:
1.0 Introduction The GOTPD (Geintra Overhead ToF People Detection dataset) is a multimodal database (depth, and infrared data) ,of recordings from a Kinect 2 camera located in overhead position, monitoring people movements under it, and it was designed to fulfill the following objectives: Allow evaluation and fine tuning of the ToF data adquisition system in the GEINTRA research group from UAH. Allow the evaluation of people detection, headgear accesories identification and human activity detection algorithms based on data generated by ToF cameras (Including depth and infrared) placed in overhead position. Provide quality data to the research community in people detection and identification tasks. The people detection task (and the data provided) can also be extended to practical applications such as video-surveillance, access control, people flow analysis, behaviour analysis or event capacity management. To give you an idea on what to expect, you can have a look at the following videos we prepared from the data: sample_videos_1, sample_videos_2, sample_videos_3 and sample_videos_4 2.0 Database Info You can find all the detailed database info in the GOTDP-readme.pdf file, the info provide here is based in this detailed info GOTPD is composed of 48 sequences comprising a broad variety of conditions, with scenarios comprising: Single and multiple persons Single and multiple non-persons (such as chairs) Persons with and without accessories (hats, caps) Persons with different complexity, height, hair color, and hair configuration Persons actively moving and performing additional actions (such as using their mobile phones, moving their fists up and down, moving their arms, etc.). The actual video footage is over 28 minutes, with sequence lengths ranging from 4 seconds to 2.63 minutes. Suplemental data(23 additional sequences) that we have used to extend the GOTPD data to face DNN based aproaches can we found at this additional link(we stored it there as kaggle does not support datasets over 20GB): http://www.depeca.uah.es/media/files/GOTPD+_Upgraded.zip The depth information (distance to the camera plane) in stored in plain binary(.z16) form, with each pixel distance represented in millimeters as a (little endian) signed integer of two bytes. Its values range from 0 to 4500. File naming conventions: To ease adapting the experimental setup for specific tasks, we have designed a (verbose) naming conven- tion for the file names. Each file is named following this structure: seq-PXX -MYY -AUUUU -GXX -CWW -SVVV, where: • PXX : Number of persons in the scene. XX is the maximum number of people than can be seen simultaneously in the scene. Note that there may be multiple users recorded in a given sequence, but at most XX will be seen at the same time. • MYY : Movement information. YY is written in decimal but is meant to refer to a bitmask, with the following convention: – 00 N/A 4– 01 static – 02 mostly regular around scene – 04 mostly random – 08 reduced movements (almost static, probably turning on) AUUUU : Activity information. UUUU is written in decimal but is meant to refer to a bitmask, with the following convention: – 0000 N/A – 0001 Normal walking – 0002 Looking to smartphone (texting) or looking to floor – 0004 Talking to phone on ear (phone call) – 0008 Facing fists up and down – 0016 Standing still moving up and down – 0032 Moving chairs – 0064 User pushing chair – 0128 User making squats to modify his/her height – 0256 Standing nearby image border (turning on) GXX : Grouping information. XX is written in decimal but is meant to refer to a bitmask, with the following convention: – 00 N/A – 01 mostly not forced – 02 mostly forced to be close CWW : Accesories information. WW is written in decimal but is meant to refer to a bitmask, with the following convention: – 00 No – 01 Some users have hats – 02 All user have hats – 04 Some users have caps – 08 All user have caps SVVVV : Sequence number information. VVVV is the sequence number and is unique across the database. All the variables defined above (XX, YY, UUUU, XX, WW, VVV ) are left zero-padded, so that parsing the filenames is trivial. Filename extensions: The distributed filenames have an extension that identifies their type, as follows: z16: Depth information file. ir16: Infrared information file. gt: Ground truth information. ToF Camera Specifications: The camera used in our recordings is a Kinect 2 for windows device, with the following main character- istics : Depth sensing – 512 x 424 – 30 Hz – FOV: 70 x 60 – One mode: 0.5–4.5 meters 1080p color camera – 30 Hz (15 Hz in low light) Active infrared (IR) capabilities – 512 x 424 – 30 Hz Microphone array (4 microphones) Intrinsic parameters obtained from calibration: fx=367.286994337726; fy=367.286855347968; cx=255.165695200749; cy=211.824600345805; If you make use of this databases and/or its related documentation, you are kindly requested to cite the paper: David Fuentes-Jimenez, Roberto Martin-Lopez, Cristina Losada-Gutierrez, David Casillas-Perez, Javier Macias-Guarasa, Carlos A. Luna, Daniel Pizarro. DPDnet: A Robust People Detector using Deep Learning with an Overhead Depth Camera, Expert Systems with Applications, 2019, 113168, ISSN 0957-4174 https://doi.org/10.1016/j.eswa.2019.113168. (http://www.sciencedirect.com/science/article/pii/S0957417419308851) Carlos A. Luna, Cristina Losada-Gutierrez, David Fuentes-Jimenez, Alvaro Fernandez-Rincon, Manuel Mazo, Javier Macias-Guarasa. Robust People Detection Using Depth Information from an Overhead Time-of-Flight Camera Expert Systems with Applications, Available online 26 November 2016, ISSN 0957-4174 http://dx.doi.org/10.1016/j.eswa.2016.11.019 (http://www.sciencedirect.com/science/article/pii/S0957417416306480)



