UrbanOccupationsOETR_temettuat_cultivation_CLC_6_region_rural_geosample
收藏资源简介:
1. Technical Description, Scope, and Methodological Framework This record provides a six-region Ottoman geosampling dataset for the study of regional land use and agricultural production in the mid-nineteenth-century Ottoman Empire. It was produced in connection with UrbanOccupationsOETR, a European Research Council-funded research project hosted at Koç University between 2016 and 2022, which foregrounded the importance of rural economic dynamics for explaining long-term regional economic development in the late Ottoman Empire. The present dataset extends that agenda by converting selected household-level rural fiscal records into structured Excel data on agricultural mix, cultivated land area, agricultural assets, and local production patterns in six Ottoman regions: Ankara, Bursa, Edirne, Manisa, Plovdiv, and Ruse. The region names Plovdiv and Ruse are used in their modern English-language forms for clarity and consistency, corresponding to the Ottoman regional frameworks centered on Filibe and Rusçuk, respectively. For these regions, Ottoman-period place names are also provided alongside their modern Bulgarian equivalents in parentheses where this facilitates identification, as in Hezargrad (Razgrad) and İstanimaka (Asenovgrad). The main empirical source is the 1845 Ottoman tax surveys, commonly known as temettuat registers. These surveys preserve household-level, handwritten, non-tabulated fiscal records dispersed across thousands of populated places. The dataset makes these records usable for regional and subregional agricultural history by linking extracted household-level entries to individuals where applicable, households, populated places, subdistricts, regions, coordinates, measurement units, convertible areas, and Corine Land Cover (CLC) codes. In total, the extracted sample includes 17,848 households. Of these, 15,745 households have agricultural entries, while 2,103 households recorded no cultivated land. The Data sheet contains 46,291 spreadsheet rows in total, including rows for non-cultivating households. Of these, 44,188 rows are agricultural entries. These entries record agricultural assets, cultivated land, permanent crops, pasture, forest, livestock-related categories where applicable, crop or asset names, measurement units, convertible areas, and CLC codes. The dataset is therefore both an empirical resource and a methodological contribution to the agricultural, land-use, and regional economic history of Southeast Europe and the Middle East in the 1840s. The key methodological contribution is not simply sampling. A basic random sample of populated places would reduce archival labor, but it would not necessarily capture the environmental and accessibility differences that shaped pre-industrial agriculture. The geosampling strategy is therefore both suitability-stratified and population-aware. It selects populated places across different agricultural-suitability and accessibility conditions, while using population registers to preserve subdistrict-level household representation. The geosampling procedure is explained in detail in the article associated with this dataset. In simple terms, the method begins with the 1840s Ottoman population registers. These registers identify the populated places in each subdistrict, their household numbers, and their administrative affiliations. They define the sampling universe, provide the household denominators, and make it possible to measure how much of each subdistrict is represented in the sample. The tax surveys then provide the household-level evidence on land, crops, livestock, and other agricultural assets. GIS layers are used to decide which populated places should be sampled. Each populated place is evaluated through the agricultural and accessibility conditions around it. Agricultural suitability is the main criterion and receives 85 percent of the AHP weight. It reflects factors such as soil quality, land capability, elevation, slope, ruggedness, and access to cultivable land. Connectivity receives 15 percent of the AHP weight. It is used as a secondary measure of relative accessibility to wheeled roads, and, in the Danubian case, to inland-water transport. The purpose of the AHP classification is not to select only the largest, richest, or most productive places. Its purpose is to divide the populated places in each subdistrict into different suitability classes and then select places from across those classes. Where possible, the sample includes one populated place from each of the five AHP-derived suitability classes. In larger subdistricts, it may include two populated places from each class. The selected places are also expected to represent approximately at least 5 percent of the households recorded in the population registers for that subdistrict. This household coverage rate is later used to scale the extracted tax-survey data from the sample to the subdistrict level. Table 1. Main components of the geosampling method Method component Description Population representation Derived from the 1840s Ottoman population registers. These registers define the populated-place universe, subdistrict membership, household denominators, and sample coverage. They are also used later for scaling sampled tax-survey data to the subdistrict level. Agricultural suitability Derived from suitability rasters combining soil quality, Land Capability Classification, DEM, SRTM-30, elevation, slope, and ruggedness. This is the main selection criterion and receives 85 percent of the AHP weight. Connectivity Derived from georeferenced historical road networks, especially the Deutsche Heereskarte and the Generalkarte. This is used as a secondary proxy for relative accessibility and receives 15 percent of the AHP weight. AHP classification Agricultural suitability and connectivity are combined to classify the 90-minute walking-distance polygon around each populated place. These classifications help select populated places from different agricultural-suitability and accessibility conditions within each subdistrict. Sample selection rule Where possible, the sample includes populated places from different AHP-derived suitability classes within each subdistrict. The selected places should also represent approximately at least 5 percent of the households recorded in the population registers for that subdistrict. Extraction rule After a populated place is selected, the complete corresponding tax-survey data for that place are extracted. The sample is therefore selective at the populated-place level, but complete at the household level within each selected place. Scaling principle For estimation, sampled subdistrict totals can be multiplied by the inverse of the household coverage rate of the sample in that subdistrict. This allows household-level tax-survey data from sampled populated places to be scaled to subdistrict estimates. Once a populated place is selected, its corresponding tax-survey data are extracted in full into a Microsoft Access database specifically tailored to the needs of the project. The method is therefore selective at the populated-place level, but complete at the household level within each selected place. This preserves the internal structure of the sampled rural community while still making regional comparison possible. After extraction, agricultural entries are coded according to Corine Land Cover categories. This allows Ottoman fiscal descriptions of cultivated land, permanent crops, pasture, forest, and heterogeneous agricultural uses to be analyzed through a comparative land-use vocabulary. The coding preserves the household-level structure of the original tax surveys while enabling aggregation at the populated-place, subdistrict, regional, and, where appropriate, cross-regional scales. These CLC labels are used as modern comparative land-use categories, not as Ottoman emic classifications. For example, category 2.1.1, “Non-irrigated arable land,” is not a grain-specific category in the CLC system. In the Ottoman tax-survey context of this dataset, however, it is used operationally as a proxy for rainfed grain cultivation because the relevant entries overwhelmingly describe non-irrigated arable grain land rather than permanent crops, irrigated fields, rice fields, pasture, or tree crops. Measurement units require special care. The temettuat surveys were part of a broader Ottoman effort to standardize fiscal administration, but standardization was more successful for recorded produce than for cultivated area. Produce subject to tithe assessment was expressed through the Istanbul kile and related units, making grain quantities convertible into metric weight, with one kile corresponding to 25.6 kilograms. Cultivated area was less consistently standardized. The Ottoman-Istanbul dönüm, equivalent to 920 square meters, and its quarter unit, the evlek, can be converted into modern area units, but other entries use localized or ambiguous units. These non-standard or ambiguous units include, for instance, kıta, meaning a plot or parcel rather than a fixed surface measure; seed-, produce-, or load-denominated descriptions such as kile and araba; tree-crop entries recorded as counts of trees; and fractional or share-based notations. Yabanabad in Ankara is an important example of this last problem, since some land there appears through fractional or share-based notations, probably referring to jointly held fields rather than directly measurable surface area. Such entries are preserved in the dataset but excluded from modern area conversions when they cannot be converted into square meters or hectares through a fixed and reliable rule. The main unit of observation in the Data sheet is the agricultural asset or crop entry. Each entry is linked to a household through HouseID, to an individual through IndividualID where applicable, to a populated place through GeoCode and Location, and to broader administrative units through Subdistrict and Region. A single household can therefore appear in multiple rows if it recorded several agricultural assets or crop types. Researchers should choose their analytical denominator carefully. For asset-level analysis, the row is the appropriate unit. For household-level analysis, rows should be aggregated by HouseID. For populated-place or subdistrict analysis, users should aggregate entries by GeoCode, Location, Subdistrict, and Region. Geographic observations are organized through documented administrative membership rather than reconstructed nineteenth-century boundary polygons. This is deliberate. Precise historical boundary polygons cannot always be reconstructed without speculation. By linking households and agricultural entries to geolocated populated places and documented administrative units, the dataset can be reaggregated at chosen spatial scales and compared with later administrative, cadastral, demographic, agricultural, and land-use datasets. Table 2. Regional coverage of the extracted dataset Region Populated places Subdistricts Households in extracted sample Rows in Data sheet Ankara 40 8 1,236 3,271 Bursa 55 11 3,547 11,349 Edirne 55 9 3,095 7,627 Manisa 82 17 6,959 16,338 Plovdiv 20 4 1,920 4,934 Ruse 25 4 1,091 2,772 Total 277 53 17,848 46,291 Note: The Data sheet contains 46,291 spreadsheet rows in total, including rows for non-cultivating households. Of these, 44,188 rows are agricultural entries. The dataset was designed to support comparative and longitudinal agricultural history for regions of the Ottoman Empire and its successor states. Rather than treating the tax surveys as isolated local case-study material, the project converts selected non-tabulated household records into spatially structured evidence. This makes it possible to analyze cultivation patterns, agricultural mix, landholding, estate concentration, rural inequality, and local production modalities in relation to environmental conditions, accessibility, and administrative membership. 2. Data Categories, Variables, and Coding Tables The following tables describe the main analytical fields, CLC categories, and measurement-unit categories used in the workbook. 2.1 Primary Variables in the Data Sheet Variable Level/type Definition and use HouseID Household identifier Unique identifier for a household in the extracted dataset. Multiple agricultural entries may share the same HouseID. Treat as an identifier, not as a numeric measurement. GeoCode Populated-place identifier Unique code assigned to a geosampled populated place. Used to link households and agricultural entries to a location. Latitude_Num Coordinate Latitude of the populated place as stored in decimal form. Longitude_Num Coordinate Longitude of the populated place as stored in decimal form. Region Regional unit Name of the sampled region or district-level unit, corresponding broadly to the Ottoman sancak-level regional frame used in the article. SubDistrict Subregional unit Name of the sampled subregional unit used for sampling, representation, and scaling. In most cases this corresponds to the Ottoman kaza-level frame. Only İstanimaka (Asenovgrad) and Tutrakan belonged to different administrative hierarchies across the population-register and tax-survey contexts; they are retained in the SubDistrict field for consistency across the six-region dataset. Location Populated place Name of the geosampled populated place, such as village or a uqarter, as transcribed and standardized in the project data. RegisterNo Archival/source reference Number or code of the Ottoman register from which the entry was extracted. When the value appears only as a number, it refers to the ML.VRD.TMT.d. series of the Ottoman Archives, unless otherwise stated. When the value begins with “KK,” it refers to the Kamil Kepeci classification/fond. HaneNo Original household number Household number as recorded in the Ottoman register, corresponding to hane/menzil numbering. AgrID Agricultural entry identifier Unique identifier for a specific agricultural asset or crop entry. IndividualID Individual identifier Identifier linking an agricultural entry to an individual where the original tax-survey structure assigns the asset to a person, usually a household head or economically recorded individual. Agriculture Original agricultural label Transcribed agricultural asset, crop, land, tree, pasture, or related category from the tax survey, rendered in modern Turkish spelling while preserving source meaning. CLCAgricultureCode CLC code Corine Land Cover code assigned to the agricultural entry. Used to standardize heterogeneous Ottoman fiscal categories into a comparative land-use vocabulary. CtgUnit Count/category unit Non-area unit or category used for quantities such as counts of trees, hives, structures, parcels, or ambiguous units. Examples include Aded, Bab, Eşcar, Kıta, res, sak, _none, and __undeciphered. Unit Quantity of CtgUnit Numeric quantity associated with CtgUnit when such a count/category unit is recorded. CtgArea Area or quantity unit Original area or land-quantity unit recorded in the source. Common values include Dönüm and Evlek; non-area or ambiguous values include Kile, Kıyye/kıyye, Şinik, çeki, sehm, Araba, _none, and _undeciphered. Area Quantity of CtgArea Numeric quantity associated with the unit recorded in CtgArea. Convertible_to_Ottoman_Dönüm True/False conversion flag True/False field indicating whether the recorded unit can be reliably converted to Ottoman-Istanbul dönüm under the project’s conversion rules. True = convertible; False = not convertible. Ottoman_Dönüm Normalized Ottoman area Converted area in Ottoman-Istanbul dönüm where reliable conversion is possible. One Ottoman-Istanbul dönüm is equivalent to 920 square meters. Area_in_square_meters Metric conversion field Metric area derived from Ottoman_Dönüm using the conversion 1 Ottoman-Istanbul dönüm = 920 square meters. This field should be used only where Convertible_to_Ottoman_Dönüm = True. 2.2 CLC Codes Observed in the Workbook CLC code CLC label Observed count Interpretive note 2.1.1 Non-irrigated arable land 22,052 Operational proxy for rainfed grain cultivation in this Ottoman tax-survey context; not a universal grain-specific CLC category. 2.1.2 Permanently irrigated land 1,874 Irrigated arable cultivation where recorded. 2.1.3 Rice fields 95 Rice cultivation. 2.2.1 Vineyards 9,643 Vineyard entries. 2.2.2 Fruit trees and berry plantations 4,293 Fruit/tree-crop categories not classified as olive groves. 2.2.3 Olive groves 3,412 Olive-related entries. 2.3.1 Pastures 1,268 Pasture and meadow-related entries. 2.4.3 Land principally occupied by agriculture, with significant natural vegetation 18 Mixed agricultural/natural vegetation cases. 3.1.3 Mixed forest 1 Forest/seminatural category. 3.2.1 Natural grasslands 1 Seminatural vegetation category. 3.2.2 Moors and heathland 8 Seminatural vegetation category. 3.2.3 Sclerophyllous vegetation 173 Seminatural vegetation category. 3.2.4 Transitional woodland-shrub 570 Seminatural vegetation category. 4.1.1 Inland marshes 1 Wetland category. NA / blank No CLC code, non-cultivating row, or uncoded entry Varies by analytical filter Not a CLC land-use category. Use with caution, because blank or NA values may reflect different row types or coding situations. 2.3 Measurement Unit and Category Values Observed in the Workbook The following table lists all unit and category values observed in the CtgArea and CtgUnit fields of the workbook, including non-area terms and internal data flags. Counts are observed nonblank values in the workbook. They should be used as documentation of the dataset, not as direct measures of land area. Unit/category Field Observed count Type Conversion / interpretation rule Dönüm CtgArea 33,516 Area Convertible. Ottoman-Istanbul dönüm = 920 square meters = 0.092 hectares. Evlek CtgArea 2,481 Area Convertible. Quarter unit of the dönüm; convert through Ottoman dönüm before square meters or hectares. Kile CtgArea 992 Dry-capacity / seed or produce quantity Not a surface-area unit. For grains, one Istanbul kile corresponds to 25.6 kilograms; kile-denominated land descriptions should not be converted to hectares without a local rule. Kıyye / kıyye CtgArea 227 Weight quantity Not a surface-area unit. Preserved but excluded from area conversion. Şinik CtgArea 70 Dry-capacity / seed quantity Not a surface-area unit. Preserved but excluded from area conversion. çeki CtgArea 120 Weight or load quantity Not a surface-area unit. Preserved but excluded from area conversion. Araba CtgArea 36 Cartload / load quantity Cartload-based description. Preserved but excluded from surface-area conversion. sehm CtgArea 42 Share-based notation Probably refers to shares in jointly held fields in some locations. Preserved but excluded from area conversion unless local evidence permits conversion. _none CtgArea 411 Internal blank/none flag Indicates no recorded CtgArea value in the workbook field. Do not treat as zero area without checking the row context. _undeciphered CtgArea 51 Internal undeciphered flag Source unit could not be confidently read. Preserved but excluded from area conversion. Aded CtgUnit 3,800 Count/category unit Count unit, often used for discrete assets. Not automatically convertible into surface area. Kıta CtgUnit 953 Parcel/plot term Means a plot or parcel rather than a fixed surface measure. Not convertible into square meters or hectares through a fixed rule. Bab CtgUnit 92 Count/category unit Category/count term in the source data. Not automatically convertible into hectares. Eşcar CtgUnit 120 Tree count/category Tree-related count/category term. Some tree-crop entries are recorded as counts of trees rather than cultivated surface area. res CtgUnit 148 Count/category unit Count/category term as recorded in the data. Not a surface-area measure. sak CtgUnit 189 Count/category unit Count/category term as recorded in the data. Not a surface-area measure. _none CtgUnit 218 Internal blank/none flag Indicates no recorded CtgUnit value in the workbook field. Interpret according to row context. __undeciphered CtgUnit 2 Internal undeciphered flag Source category unit could not be confidently read. Preserved but not converted. For area-based analysis, users should filter to entries with reliable area units and positive area values. In practice, this means using Dönüm and Evlek entries that have been converted to Ottoman_Dönüm and Area_in_square_meters. Kıta, araba, kile-denominated descriptions, tree counts, fractional shares, and other non-area entries should not be converted into surface area unless a separate local conversion rule is independently established.



