遇见数据集

Code and data associated with: Searching the web builds fuller picture of arachnid trade

收藏
Zenodo2022-07-20 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

Data and code used in the paper: Searching the web builds fuller picture of arachnid trade. Throughout the methods we have indicated the stage of analysis each data component was used and the code script connected. We have numbered to code and data supplements to reflect as closely as possible the order in which data generation and summary was undertaken. The following provide additional details linked to each of the data files. Data S1 - Website data: lang = language of the search engine used, ad hoc websites had language described after discovery; engine = the search engine used; page = the page on which the website appeared from the search engine; searchdate = search date in YYYY-mm-dd HH:MM:SS; link = link to the webpage, redacted to protect website identity; reviewdate = date revewied for arachnids being sold and search strategy; sells = whether the website sells arachnids (1 == sells); allow = whether the site explcicilt forbids automated searching (1 == allows, NA when search method was not fully automated, e.g., single page); type = the type of the website (e.g., trade, classified ads); order = whether arachnids where organised in a particular ways; target = a refined target URL to start search; method = the search method chosen, see methods for details; refine = any refinement or filter than could constrain the scope of the website to be searched; spages = the number of pages required to cycle through to cover the entire stock (also separated by ; if multiple cycles where needed or multiple single pages could be easily collected); prelimCheck = whether the website passed initial checks for arachnid selling; notes = any details that might need special attention during searches; webID = code used for subsequent data summary. Data S2 - Raw keyword searches outputs: species keywords. sp = the modern species or genus that a keyword is associated with; page = the number of the page the keyword was detected on; keyw = the exact keyword that was detected; spORgen = whether the keyword was a species binomial or just genus; termsSurrounding = the words surrounding a genus keyword detection (only applies to Data S3); webID = the website ID. Data S3 – Raw keyword searches outputs: genus keywords. sp = the modern species or genus that a keyword is associated with; page = the number of the page the keyword was detected on; keyw = the exact keyword that was detected; spORgen = whether the keyword was a species binomial or just genus; termsSurrounding = the words surrounding a genus keyword detection (multiple detections separated by ;); webID = the website ID. Data S4 - Raw keyword search outputs: temporal sample. sp = the modern species or genus that a keyword is associated with; page = the number of the page the keyword was detected on; keyw = the exact keyword that was detected; spORgen = whether the keyword was a species binomial or just genus; termsSurrounding = the words surrounding a genus keyword detection (multiple detections separated by ;); webID = the website ID; timestamp.parse = the timestamp extracted from the archived web page; year = a simplified timestamp including only the year. Data S5 - LEMIS data used. An arachnid filtered version of <sup>74,75</sup>. Data S6 - CITES trade database data used <sup>76</sup>. Data S7 - CITES appendices data used <sup>77</sup>. Data S8 - IUCN Redlist data used <sup>78</sup>. Data S9 - Compiled final dataset, with data deriving from WSC, Scorpion files, ITIS, WAM and the data collection process. speciesId = a numeric code, one per species; clade = the clade the species belongs to; family = the family the species belongs to; genus = the genus of the species; species = the species epithet; author = the species authority name; year = the species authority year; parentheses = whether parentheses are needed with the authority; distribution = WSC original distribution descriptions; invalid = whether the species is considered valid; source = the species source, either World Spider Catalogue, Scorpion files, ITIS or WAM; accName = the species binomial being used as our accepted name; allNames = the accepted species binomial and all synonyms; allGenera = the accepted genus, and all other genera the species has belonged to at one point; onlineTradeSnap = whether the species was detected via a match to the accName in the snapshot data; onlineTradeSnap_Any = whether the species was detected via any synonym in the snapshot data; onlineTradeSnap_genus = whether the genus was detected via a match to the genus in the snapshot data; onlineTradeSnap_genusAny = whether the genus was detected via any synonym in the snapshot data; onlineTradeTemp = whether the species was detected via a match to the accName in the temporal data; onlineTradeTemp_Any = whether the species was detected via any synonym in the temporal data; onlineTradeTemp_genus = whether the genus was detected via a match to the genus in the temporal data; onlineTradeTemp_genusAny = whether the genus was detected via any synonym in the temporal data; onlineTradeEither = whether the species was detected via a match to the accName in the temporal data or snapshot data; onlineTradeEither_Any = whether the species was detected via any synonym in the temporal data or snapshot data; LEMIStrade = whether the species was detected via a match to the accName in the LEMIS data; LEMIStrade_Any = whether the species was detected via any synonym in the LEMIS data; LEMIStrade_genus = whether the genus was detected via any synonym in the LEMIS data; LEMIStrade_genusAny = whether the genus was detected via any synonym in the LEMIS data; CITEStrade = whether the species was detected via a match to the accName in the CITES trade database data; CITEStrade_Any = whether the species was detected via any synonym in the CITES trade database data; CITEStrade_genus = whether the genus was detected via any synonym in the CITES trade database data; CITEStrade_genusAny = whether the genus was detected via any synonym in the CITES trade database data; CITESapp = the CITES appendix the species is listed under using an exact match to the accName; CITESapp_Any = the CITES appendix the species is listed under using any match to any of the species’ synonyms; redlist = the IUCN Redlist category the species is listed under using an exact match to the accName; redlist_Any = the IUCN Redlist category the species is listed under using any match to any of the species’ synonyms; extactMatchTraded = the species is detected in any of the trade sources via a match to the accName; anyMatchTraded = the species is detected in any of the trade sources via a match to any species’ synonym. Data S10 - Forum listings of “What species are you currently keeping” from an online fora posted between 9th September 2021 and 9th October 2021, to provide an idea of online discussions. Each user with a separate list is provided in a separate tab. Morph_collector is the same as poster1, but the potential cryptic species or morphs are noted separately to make them clearer. Data S11 – Distribution information for spiders. Only two columns used in summaries: accName = the accepted name used throughout summaries; NAME = the country name the spider occurs in. Data S12 - Distribution information for scorpions. species = the accepted name used throughout summaries; NAME = the country name the scorpions occurs in. Code S1 - Search URL Extract.R Code S2 - Retrieve web data.R Code S3 - Temporal Classified Ads.R Code S4 - Keyword Generation.R Code S5 - Keyword Search.R Code S6 - LEMIS filter and summary.R Code S7 - Compiling results.R Code S8 - Summary Figures.R Code S9 - Temporal Figures.R Code S10 - New description figure.R Code S11 - Term exploration.R Code S12 - LEMIS summary and mapping.R

本研究配套的数据与代码均来自论文《利用网络搜索全面刻画蛛形纲动物贸易图景》(Searching the web builds fuller picture of arachnid trade)。在研究方法部分,我们已标注各数据组件所对应的分析阶段及关联的代码脚本。我们对代码与数据补充材料进行了编号,以尽可能贴合数据生成与汇总的实际执行顺序。下文将针对各数据文件逐一补充说明。 数据S1——网站数据: lang:所用搜索引擎的语言;临时采集的网站的语言将在发现后标注 engine:所用搜索引擎 page:该网站在搜索引擎结果页面中的排名位置 searchdate:搜索日期,格式为YYYY-mm-dd HH:MM:SS link:网页链接,为保护网站主体身份已进行脱敏处理 reviewdate:针对网站售卖蛛形纲动物情况及搜索策略进行审核的日期 sells:该网站是否售卖蛛形纲动物(1代表是) allow:该网站是否明确允许自动化搜索(1代表允许;若搜索方式非完全自动化,如仅单页搜索,则标注为NA) type:网站类型(如交易平台、分类广告平台等) order:蛛形纲动物是否在网站中按特定方式分类 target:用于启动搜索的优化目标URL method:所选搜索方法,具体说明详见研究方法部分 refine:可用于限制搜索网站范围的任何优化条件或筛选规则 spages:遍历全部在售库存所需的页面数量;若需多次循环遍历或可轻松收集多个独立单页,则以分号分隔各数值 prelimCheck:该网站是否通过蛛形纲动物售卖情况的初始审核 notes:搜索过程中需特别关注的相关细节 webID:用于后续数据汇总的编码 数据S2——原始关键词搜索结果:物种关键词 sp:与关键词关联的现生物种或属 page:检测到关键词的结果页面编号 keyw:检测到的精确关键词 spORgen:该关键词是物种双名法命名还是仅属名 termsSurrounding:属名关键词检测结果周边的文本(仅适用于数据S3) webID:网站ID 数据S3——原始关键词搜索结果:属关键词 sp:与关键词关联的现生物种或属 page:检测到关键词的结果页面编号 keyw:检测到的精确关键词 spORgen:该关键词是物种双名法命名还是仅属名 termsSurrounding:属名关键词检测结果周边的文本(多个检测结果以分号分隔) webID:网站ID 数据S4——原始关键词搜索结果:时间样本 sp:与关键词关联的现生物种或属 page:检测到关键词的结果页面编号 keyw:检测到的精确关键词 spORgen:该关键词是物种双名法命名还是仅属名 termsSurrounding:属名关键词检测结果周边的文本(多个检测结果以分号分隔) webID:网站ID timestamp.parse:从存档网页中提取的时间戳 year:仅包含年份的简化时间戳 数据S5——所用LEMIS(执法管理信息系统,Law Enforcement Management Information System)数据:经蛛形纲动物筛选的参考文献<sup>74,75</sup>数据集 数据S6——所用濒危野生动植物种国际贸易公约(Convention on International Trade in Endangered Species of Wild Fauna and Flora, CITES)贸易数据库数据<sup>76</sup> 数据S7——所用CITES附录数据<sup>77</sup> 数据S8——所用国际自然保护联盟红色名录(International Union for Conservation of Nature Red List, IUCN Red List)数据<sup>78</sup> 数据S9——最终汇总数据集,其数据来源于世界蜘蛛目录(World Spider Catalogue, WSC)、蝎子文件(Scorpion files)、综合分类信息系统(Integrated Taxonomic Information System, ITIS)以及西澳大利亚博物馆(Western Australian Museum, WAM),并结合本研究的数据收集流程生成。各字段说明如下: speciesId:每个物种对应的唯一数值编码 clade:该物种所属的演化支 family:该物种所属的科 genus:该物种所属的属 species:该物种的种加词 author:该物种的命名者姓名 year:该物种的命名年份 parentheses:命名者信息是否需要添加括号 distribution:世界蜘蛛目录原有的分布描述 invalid:该物种是否被认定为无效物种 source:该物种的数据源,可选值为世界蜘蛛目录、蝎子文件、综合分类信息系统或西澳大利亚博物馆 accName:本研究采用的该物种的接受学名(双名法) allNames:该物种的接受学名及所有异名 allGenera:该物种当前的有效属,以及该物种曾被归入的所有其他属 onlineTradeSnap:该物种是否通过快照数据中与接受学名的匹配被检测到 onlineTradeSnap_Any:该物种是否通过快照数据中与任意异名的匹配被检测到 onlineTradeSnap_genus:该物种所在的属是否通过快照数据中与属名的匹配被检测到 onlineTradeSnap_genusAny:该物种所在的属是否通过快照数据中与任意异名属名的匹配被检测到 onlineTradeTemp:该物种是否通过时间序列数据中与接受学名的匹配被检测到 onlineTradeTemp_Any:该物种是否通过时间序列数据中与任意异名的匹配被检测到 onlineTradeTemp_genus:该物种所在的属是否通过时间序列数据中与属名的匹配被检测到 onlineTradeTemp_genusAny:该物种所在的属是否通过时间序列数据中与任意异名属名的匹配被检测到 onlineTradeEither:该物种是否通过快照数据或时间序列数据中与接受学名的匹配被检测到 onlineTradeEither_Any:该物种是否通过快照数据或时间序列数据中与任意异名的匹配被检测到 LEMIStrade:该物种是否通过LEMIS数据中与接受学名的匹配被检测到 LEMIStrade_Any:该物种是否通过LEMIS数据中与任意异名的匹配被检测到 LEMIStrade_genus:该物种所在的属是否通过LEMIS数据中与属名的匹配被检测到 LEMIStrade_genusAny:该物种所在的属是否通过LEMIS数据中与任意异名属名的匹配被检测到 CITEStrade:该物种是否通过CITES贸易数据库数据中与接受学名的匹配被检测到 CITEStrade_Any:该物种是否通过CITES贸易数据库数据中与任意异名的匹配被检测到 CITEStrade_genus:该物种所在的属是否通过CITES贸易数据库数据中与属名的匹配被检测到 CITEStrade_genusAny:该物种所在的属是否通过CITES贸易数据库数据中与任意异名属名的匹配被检测到 CITESapp:通过与接受学名的精确匹配所确定的该物种所属的CITES附录 CITESapp_Any:通过与该物种任意异名的匹配所确定的该物种所属的CITES附录 redlist:通过与接受学名的精确匹配所确定的该物种所属的IUCN红色名录等级 redlist_Any:通过与该物种任意异名的匹配所确定的该物种所属的IUCN红色名录等级 extactMatchTraded:该物种是否通过与接受学名的匹配在任意贸易数据源中被检测到 anyMatchTraded:该物种是否通过与任意异名的匹配在任意贸易数据源中被检测到 数据S10——2021年9月9日至2021年10月9日期间某在线论坛发布的「你目前饲养的是什么物种」主题帖子,用于反映相关在线讨论情况。每个用户的单独列表将单独存储在一个标签页中。Morph_collector与poster1内容一致,但为便于区分潜在隐存种或形态型,将其单独标注。 数据S11——蜘蛛分布信息。汇总分析中仅使用以下两列数据: accName:汇总分析中采用的接受学名 NAME:该蜘蛛分布的国家名称 数据S12——蝎子分布信息: species:汇总分析中采用的接受学名 NAME:该蝎子分布的国家名称 代码S1——Search URL Extract.R 代码S2——Retrieve web data.R 代码S3——Temporal Classified Ads.R 代码S4——Keyword Generation.R 代码S5——Keyword Search.R 代码S6——LEMIS filter and summary.R 代码S7——Compiling results.R 代码S8——Summary Figures.R 代码S9——Temporal Figures.R 代码S10——New description figure.R 代码S11——Term exploration.R 代码S12——LEMIS summary and mapping.R

提供机构:
Zenodo
创建时间:
2022-04-18
二维码
社区交流群
二维码
科研交流群
商业服务