遇见数据集

UrbanOccupationsOETR_1840s_Ottoman_Bursa_District_TMT_geosample_dataset

收藏
Zenodo2024-08-13 更新2026-05-26 收录
官方服务:

资源简介:

With the UrbanOccupationsOETR, a European Research Council-funded research project hosted at Koç University 2016-2022, we wanted to highlight the importance of rural economic dynamics to explain differences in long-term regional economic development in the late Ottoman Empire. We provide an Excel dataset on the crop-specific agricultural mix and land area of an Ottoman region, Bursa, in the 1840s. This dataset is the result of a new geosampling methodology we devised, representing a key development in the agricultural and overall economic history of Southeast Europe and the Middle East. The 1840s serve as a good period to choose for base years mainly due to three main factors to sample economic data on a regional scale. First, due to Tanzimat reforms (planned and only partially accomplished transformation of the Ottoman central administration in the mid-nineteenth century), the 1840s marked a watershed of bureaucratical information gathering. Especially, the temettuat registers were created as a by-product to realize a drastic change in tax collection. With at least in its first iteration, the unsuccessful abolishment of tax-farming by the Tanzimat decree in 1839, the Ottoman central administration aimed to transform the existing indirect and communal taxation with direct and individual modalities. To accomplish this goal, the administration had to survey the tax base, which was in disguise due to centuries-long tax farming practices. The temettuat registers were conducted in the core regions of the empire with the main exception of the imperial capital, Istanbul. Second, the 1840s correspond to the last period before the beginning of drastic territorial losses, primarily in Southeast Europe, which triggered in size and frequency unprecedented waves of emigration and immigration between the core territories of the empire both in Southeast Europe as well as in Anatolia, which continued until the official demise or the implosion of the empire. Third and lastly, the 1840s serves as a very suitable point to assess the dynamics of pre-industrial and ancienne regime agricultural dynamics due to the lack of modern means of mechanization, irrigation, and fertilization combined with extremely rudimentary transport facilities. The temettuat surveys are invaluable resources for they provide agricultural asset- / crop-type specific agricultural mix information with cultivation area per household. However, extracting their detailed information requires a team and years. To overcome this, we developed a sampling strategy that selected five locations per subdistrict using the Analytical Hierarchy Process (AHP), considering factors of agricultural suitability (85% weight), connectivity to historical roads (within a 500-meter to the closest road or to the Danube, 15% weight, justified by its impact on suitability), and subdistrict population size (chosen villages must represent at least 5% of the subdistrict's total population). Our geosampling methodology of the 1840s tax registers (temettuat) is based on contemporary Ottoman population registers. With this geosampling method, we aim to estimate the regional (district (sancak) and subdistrict (kaza)) level total area of cultivation and shares of the agricultural mix for key products. We are using two mid-nineteenth-century datasets: Ottoman tax (TMT) (temettuat) surveys for agricultural asset / crop type and cultivation area and the population (nüfus) (NFS) registers for population-based sampling. Connectivity is based on a detailed and provenly accurate 1940s German military map of Turkey, Deutsche Heereskarte (DHK). The agricultural suitability raster is an amalgamation of the Land Capability Classification (LCC) encapsulating the variables of soil quality and quantity and the Digital Elevation Model (DEM) based on Shuttle Radar Topography Mission with 30-meter-resolution and comprising elevation, slope, and ruggedness data. In the end, a geosampling initiative was undertaken across six regions in Southeast Europe and Anatolia, namely Ankara, Bursa, Plovdiv, Ruse, Manisa, and Edirne, covering a total of 277 locations with 17,675 households. Our project team entered the economic data from those records into a Microsoft Access database. We employed a specially crafted data entry template to systematically organize the tax survey data into multiple categories. After geosampling locations, our objective extended to deriving estimates for the total cultivated area within each subdistrict and regions. To achieve this goal, it was imperative that the data undergoes coding the cultivation areas into a standardized and comparable land-use scheme. We adopted the Corine Land Cover (CLC) nomenclature from the European Union's Earth Observation Programme (Copernicus), established in 1985 and regularly updated. Our study followed the revised guidelines issued by the European Environment Agency on 10.05.2019. Despite its primary design for contemporary land cover analysis, CLC nomenclature proved well-suited for accurately representing the agricultural tax data and the historical context of the 1845 Ottoman tax surveys. In our analysis, we coded micro-level cultivated land entries associated with individual households, using CLC's highest detail level. Successfully, every cultivated land entry was coded into the third level of detail in CLC, encompassing sub-categories such as 2.1 – “Arable land”, 2.2 – “Permanent crops”, 2.3 – “Pastures”, and 2.4 – “Heterogeneous agricultural areas”—all falling under the overarching category of 2 - Agricultural areas. Additionally, we coded entries related to 3.1 - “Forest” and 3.2 – “Shrub and/or herbaceous vegetation associations”, falling under the primary category of 3 – “Forest and seminatural areas.” Finally, cultivation area expressed in Ottoman measurement units like dönüm (1/9,2 of a hectare) are converted into hectares to ensure consistency and ease of spatiotemporal comparison. We provide the geosample data of the Bursa region, positioned in Western Anatolia, renowned for its historical and economic significance, large and cosmopolitan population, and diverse geophysical characteristics. This data covers all the geosampled households, the individuals residing in them, and their CLC-coded agricultural assets / crops with quantity / cultivation area. The Bursa region comprised 591 geolocated settlements in 12 subdistricts in 1840. Notably, the city of Bursa, serving as the major urban center and regional capital, was intentionally left out of the sample. Additionally, the subdistrict of Pazarköy, with its 14 settlements, was excluded. Despite being initially part of the Bursa region in population registers, it became attached to the northern neighboring Kocaili district in 1845. Consequently, out of the 576 remaining settlements, we geosampled 55 populated places from 11 subdistricts, covering 3547 households, representing 12% of the total households in the region, totaling 30,518. The variables of the tax surveys of the geosampled locations were read, extracted, and entered into the customized Microsoft Access database. In the Bursa region, there are a total of 13,344 entries for agricultural assets coded with CLC across all 55 sampled populated areas. Our dataset includes all these entries and covers 3,325 households (out of the total of 3,547 sampled households) that owned these assets. This allows for a comprehensive analysis of the agricultural mix and land area at a detailed level. The tax survey data was transcribed in Turkish using modern Turkish spelling and punctuation to keep the nuances of the original source. That said, because the original register information is largely presented in a standardized fashion and grouped under detailed variables, the data can easily be translated into other languages and coded into specific coding schemes. The categories and descriptions of the variables of the geosample dataset for the Bursa region are as follows: Category Variable Description GeoCode “GeoCode” UniqueID belonging to a specific geosampled location Location “Longitude” & “Latitude” Geographical coordinates used to specify the precise location of a geosampled location on the Earth's surface Geographic unit of entry “Region” & “SubDistrict” & “Location” Geographic unit of entry, including region (district/sancak); subdistrict (kaza); and geosampled location as they appear in the population registers Unique key/ID “HouseID” Unique and consecutive ID belonging to a specific household, automatically generated by Microsoft Access Register specifics “RegisterNo” Archival code of the population register whose data is being entered “Household” Number of the household (specified by the registers as Menzil, Persian word for house), as appears in the register Unique key/ID "IndivID" Unique ID belonging to a specific individual, automatically generated by Microsoft Access Ethno-religious identity “EthnoReligiousIdentity” Ethnoreligious identity of the individual as given in the register Individual interrelationships “RelationtoHouseholdHead” Shows the individual’s relationship or lack thereof to the first recorded male within the household. If he was the “_firstRegisteredMale,” then he is recorded as such. The type of relationship to the “_firstRegisteredMale”, such as son (oğlu), grandson (hafidi/torunu), tenant (kiracı), or slave (gulamı/kölesi), or lack of it (in the case of a new individual moved to the household) are specified. “FamilyName” The individual's family name, in the rare occasion before the official introduction of family names (not to be confused with title) “Title1” Title(s) appearing before an individual’s name “NameI” The individual's name(s) "Title2” Title(s) appearing after an individual’s name “Conjunction” Arabic patronymic ibn, bin, and veled (equivalent to the “-son” suffix in English, meaning the son of "NameII") and Turkish patronymic oğlu (the opposite, meaning the son of "NameI", appearing rarely), that links "NameI" and "NameII" “TitleII1" Title(s) appearing before the father’s name (if the conjunction is "oğlu", then the title of the son (the individual himself) “NameII” The father's name(s) (if the conjunction is "oğlu", then the name of the son (the individual himself) “TitleII2” Title(s) appearing after the father’s name (if the conjunction is "oğlu", then the title of the son (the individual himself) Occupation “OccupationI” Standardized version of the occupation [e.g.,barber] “OccupationIStatus” Status of the occupation [e.g., apprentice]. If an occupation did not have a status, then “statüsüz” (no status) was entered. “OccupationII” & “OccupationIIStatus” Versions of the descriptors above if the individual is employed in multiple occupations Agricultural Asset / Crop "AgrID" UniqueID belonging to a specific agricultural asset / crop belonging to an individual "Agriculture" Type of the agricultural asset / crop "CLCAgricultureCode" CLC-code of the agricultural asset / crop "CategoryUnit" Is applicable when the quantity of an agricultural asset or crop is specified using specific terms ("aded", "res", "eşcar", "sak" [usually for individual trees]), and when the area of an agricultural asset or crop is described in vague terms ("bab", "kıta" [usually for fields, gardens, and vineyards]) "Unit" Quantity of the "CategoryUnit" "CategoryArea" The land area type of the agricultural asset / crop in Ottoman measurement units, like "Dönüm" "Area" Quantity of the "CategoryArea" "Total area of cultivated land (CategoryArea) in hectares (if applicable)" Quantity of the "CategoryArea" converted into hectares

本研究依托由欧洲研究委员会资助、设于科奇大学(Koç University)的UrbanOccupationsOETR项目(2016-2022年),旨在阐明乡村经济动态对于解释晚期奥斯曼帝国长期区域经济发展差异的重要性。本研究提供一套Excel格式数据集,涵盖19世纪40年代奥斯曼帝国布尔萨(Bursa)地区的作物特异性农业结构与土地面积。该数据集系本研究设计的全新地理抽样方法的成果,是东南欧与中东农业史及整体经济史研究的重要进展。 选择19世纪40年代作为基期以开展区域尺度经济数据抽样,主要基于三大核心因素。其一,坦齐马特改革(Tanzimat reforms)即19世纪中期奥斯曼中央行政的计划性且仅部分完成的转型,标志着官僚信息收集的分水岭。特梅杜阿特登记册(temettuat registers)正是为实现税收征管的重大变革而产生的副产品。1839年坦齐马特法令首次尝试废除包税制但宣告失败,奥斯曼中央政府旨在将原有的间接集体税制改造为直接个人税制。为达成该目标,行政当局需清查因数百年包税实践而模糊不清的税基。特梅杜阿特登记册在帝国核心区域推行,唯独未覆盖帝国首都伊斯坦布尔。其二,19世纪40年代是帝国首次大规模领土丧失(主要发生在东南欧)之前的最后时期,该领土丧失引发了规模与频率均前所未有的移民潮,涉及帝国东南欧与安纳托利亚核心领土之间的人口迁移,该趋势一直持续至帝国正式消亡或崩溃。其三,19世纪40年代是评估前工业与旧制度(ancien régime)农业动态的极佳时间节点,彼时既无现代机械化、灌溉与施肥手段,交通设施也极度原始。 特梅杜阿特调查具有极高价值,因其提供了按农业资产/作物类型划分的农业结构信息,以及每户的耕种面积。但提取其中的详细信息需要团队投入数年时间。为此,本研究开发了一套抽样策略:采用层次分析法(Analytical Hierarchy Process, AHP)为每个分区选取5个地点,考量的因素包括农业适宜性(权重85%)、与历史道路的连通性(距离最近道路或多瑙河500米范围内,权重15%,该权重设置的合理性在于其对适宜性的影响),以及分区人口规模(选中村庄需至少占分区总人口的5%)。 本研究针对19世纪40年代税务登记册(特梅杜阿特)的地理抽样方法,以同时期奥斯曼人口登记册为基础。本研究的目标是估算区域(县(sancak)与分区(kaza))层面的总耕种面积与主要农产品的农业结构占比。本研究使用两套19世纪中期数据集:一是针对农业资产/作物类型与耕种面积的奥斯曼税务(TMT, Ottoman tax (TMT))调查,二是用于人口抽样的人口(nüfus, NFS)登记册。连通性数据基于一套详细且经证实准确的1940年代德国军用地图(Deutsche Heereskarte, DHK)。农业适宜性栅格由土地能力分类(Land Capability Classification, LCC)与基于航天飞机雷达地形测绘任务(Shuttle Radar Topography Mission)的数字高程模型(Digital Elevation Model, DEM)合并而成,其中DEM分辨率为30米,包含海拔、坡度与崎岖度数据。 最终,本研究在东南欧与安纳托利亚的六个区域开展了地理抽样工作,分别为安卡拉、布尔萨、普罗夫迪夫、鲁塞、马尼萨与埃迪尔内,总计覆盖277个地点与17675户家庭。本研究团队将这些记录中的经济数据录入微软Access(Microsoft Access)数据库,并采用专门定制的数据录入模板,将税务调查数据系统地归类至多个类别中。 完成地理抽样地点选取后,本研究的目标进一步拓展至估算每个分区与区域的总耕种面积。为实现该目标,必须将耕种面积编码至一套标准化且可比较的土地利用方案中。本研究采用欧盟地球观测计划哥白尼计划(Copernicus)推出的CORINE土地覆盖(Corine Land Cover, CLC)命名法,该命名法于1985年建立并定期更新,且本研究遵循了欧洲环境署2019年5月10日发布的修订指南。尽管CLC命名法最初是为当代土地覆盖分析设计的,但实践证明其非常适合准确呈现1845年奥斯曼税务调查的农业税务数据与历史背景。 在本研究的分析中,我们采用CLC的最高细节层级对与单个家庭相关的微观耕地条目进行编码。最终,所有耕地条目均被编码至CLC的第三细节层级,涵盖以下子类别:2.1——“耕地(Arable land)”、2.2——“多年生作物(Permanent crops)”、2.3——“牧场(Pastures)”、2.4——“混合型农业区域(Heterogeneous agricultural areas)”,上述类别均归属于2大类“农业区域”。此外,我们还对归属于3大类“森林与半自然区域”的条目进行了编码,包括3.1——“森林(Forest)”与3.2——“灌丛和/或草本植被群落(Shrub and/or herbaceous vegetation associations)”。 最后,本研究将以奥斯曼单位如德南(dönüm,1/9.2公顷)表示的耕种面积转换为公顷,以确保数据一致性并便于时空比较。 本研究提供位于安纳托利亚西部的布尔萨地区地理抽样数据,该地区以其历史与经济重要性、庞大且多元的人口,以及多样的地球物理特征而闻名。本数据集涵盖所有地理抽样家庭、居住于其中的个体,以及他们经CLC编码的农业资产/作物与对应的数量/耕种面积。 1840年,布尔萨地区下辖12个分区,共计591个已地理定位的定居点。值得注意的是,作为主要城市中心与区域首府的布尔萨市未被纳入抽样范围。此外,拥有14个定居点的帕扎尔科伊(Pazarköy)分区也被排除在外,尽管其最初在人口登记册中隶属于布尔萨地区,但在1845年被划归至北部相邻的科恰伊利(Kocaili)区。因此,在剩余的576个定居点中,我们从11个分区中抽样了55个居民点,覆盖3547户家庭,占该地区总住户30518户的12%。抽样地点的税务调查数据已被读取、提取并录入定制的微软Access数据库。 在布尔萨地区的55个抽样居民点中,共计存在13344条经CLC编码的农业资产条目。本数据集包含所有这些条目,覆盖了3325户拥有这些资产的家庭(抽样家庭总数为3547户),这使得我们能够开展精细化的农业结构与土地面积分析。 本研究采用现代土耳其语拼写与标点转录税务调查数据,以保留原始资料的细节特征。需要说明的是,由于原始登记信息大多采用标准化格式并按详细变量分组,该数据可轻松翻译成其他语言,并编码至特定的编码方案中。 布尔萨地区地理抽样数据集的变量分类与描述如下: 分类 变量 描述 地理编码(GeoCode) “GeoCode” 特定地理抽样地点的唯一标识 位置 “Longitude” & “Latitude” 用于指定地理抽样地点在地球表面精确位置的地理坐标 录入地理单元 “Region” & “SubDistrict” & “Location” 录入的地理单元,包括区域(县/桑贾克(sancak))、分区(卡扎(kaza)),以及人口登记册中记载的地理抽样地点 唯一键/标识 “HouseID” 特定家庭的唯一连续标识,由微软Access自动生成 登记详情 “RegisterNo” 录入数据所依据的人口登记册的档案编码 “Household” 登记册中记载的家庭序号(登记册中将家庭称为Menzil,即波斯语中“房屋”之意) 唯一键/标识 “IndivID” 特定个体的唯一标识,由微软Access自动生成 族裔宗教身份 “EthnoReligiousIdentity” 登记册中记载的个体族裔宗教身份 个体亲属关系 “RelationtoHouseholdHead” 显示个体与家庭中首位登记男性的关系或无关联。若个体为“_firstRegisteredMale”,则将其记录为该值。与“首位登记男性”的关系类型包括:儿子(oğlu)、孙子(hafidi/torunu)、佃户(kiracı)、奴隶(gulamı/kölesi),或无关联(如新迁入家庭的个体)。 “FamilyName” 个体的姓氏,在官方正式引入姓氏之前极为罕见(注意勿与头衔混淆) “Title1” 出现在个体姓名前的头衔 “NameI” 个体的名字 “Title2” 出现在个体姓名后的头衔 “Conjunction” 阿拉伯语父名前缀ibn、bin、veled(对应英语中的“-son”后缀,意为“……之子”),以及土耳其语父名后缀oğlu(意为“……之子”,较为罕见),用于连接“NameI”与“NameII” “TitleII1” 出现在父亲姓名前的头衔(若连接词为“oğlu”,则为个体本人的头衔) “NameII” 父亲的名字(若连接词为“oğlu”,则为个体本人的名字) “TitleII2” 出现在父亲姓名后的头衔(若连接词为“oğlu”,则为个体本人的头衔) 职业 “OccupationI” 职业的标准化表述,例如“理发师” “OccupationIStatus” 职业的身份,例如“学徒”。若职业无特定身份,则填入“statüsüz(无身份)” “OccupationII & OccupationIIStatus” 若个体从事多重职业,则使用上述描述的对应版本 农业资产/作物 “AgrID” 特定个体拥有的农业资产/作物的唯一标识 “Agriculture” 农业资产/作物的类型 “CLCAgricultureCode” 农业资产/作物的CLC编码 “CategoryUnit” 当农业资产或作物的数量以特定术语(aded、res、eşcar、sak,通常用于单个树木)表示,或农业资产或作物的面积以模糊术语(bab、kıta,通常用于田地、花园与葡萄园)表示时,适用该字段 “Unit” “类别单位”对应的数量 “CategoryArea” 以奥斯曼计量单位(如德南(dönüm))表示的农业资产/作物的土地面积类型 “Area” “类别面积”对应的数量 “Total area of cultivated land (CategoryArea) in hectares (if applicable)” 已转换为公顷的“类别面积”对应的总耕地面积

提供机构:
Zenodo
创建时间:
2024-07-01
二维码
社区交流群
二维码
科研交流群
商业服务