meta-analysis.cz: harmonised estimate-level data from meta-analyses in economics
收藏资源简介:
Estimate-level data from meta-analyses in economics and the social sciences, with the study and design characteristics hand-coded for each original paper. What this deposit contains. The harmonised estimate-level table pooling 40 economics meta-analysis literatures, the machine-readable dataset index, a column-level codebook for each of the 44 source datasets, and the documentation. It does NOT contain the per-dataset format conversions, the papers themselves, or their replication packages; those remain at https://meta-analysis.cz. What the numbers mean. 44 datasets contain 65,349 rows in the converted source files. After each paper's own analysis selection, and where necessary a reshape of wide data into estimate-level form, those become 53,960 estimates. The harmonised table pools 48,355 of them across 40 literatures. Each dataset also ships a column-level codebook. Rows are not always independent estimates. The harmonised table holds one row per harmonised observation. Two literatures are impulse responses: price_puzzle contributes 1,415 rows across seven horizons, and house_prices contributes 1,555 rows at one row per impulse-response horizon, so rows within a single response function are not independent of each other. Use the horizon column before treating rows as independent observations. Review status: every literature has been checked against its source. All 40 pooled literatures are verified, with the evidence quoted per dataset: 20 reviewed against the paper's own replication code or published results by hand, and 20 confirmed mechanically by reading each paper's code and comparing the variables it regresses with the columns published here. None rests on arithmetic column-matching alone. gasoline_price ships no replication code at all and was therefore verified against its paper's published results directly: the abstract reports a corrected long-run elasticity of -0.31 and short-run of -0.09, with published averages 'exaggerated twofold', and the shipped data gives means of -0.691 and -0.227, reproducing that twofold exaggeration and the long-run/short-run ordering. Every dataset publishes an audit_status. Changes since 0.9.0-beta. This release corrects defects found by an independent audit of the beta files. price_puzzle repeated every estimate seven times, because its source file is already long on the impulse-response horizon while its response columns are wide and the reshape did not deduplicate; it now carries 1,415 rows across all seven horizons, including the trough and peak responses that the beta discarded entirely. remittances published the raw regression coefficients its source file carries, over five different dependent variables; it now publishes the partial correlations its paper actually analyses, and its sample-size column, which was a row counter rather than a sample size, is set to null. pub_year held a standardised model regressor rather than a publication year in five literatures, because the column was matched by name; year columns are now validated by value. trust is newly included, contributing the 284 estimates that size does not already carry. t_stat and precision are now derived in double precision throughout. Because of the deduplication the table is smaller than the beta, at 48,355 rows against 54,076, while covering more literatures and more horizons. Read the weight concentration before benchmarking. Each dataset now publishes max_precision_weight_share, the share of total inverse-variance weight held by its single most precise estimate. In several literatures one estimate carries more than half, and in three it carries essentially all of it, because the original paper rounded a standard error to a very small value. Any precision-weighted estimator on those literatures returns approximately that one observation, so a comparison of estimators there measures very little. Check this field before drawing conclusions. Which column holds the effect and which its standard error was resolved arithmetically rather than by name: the correct pair is the one for which the effect divided by its standard error reproduces the t-statistic the dataset already reports. Where that was not decisive, the mapping was taken from the paper's own published replication code, and the evidence is recorded per dataset. Every row of the harmonised table carries the source file and the exact column names it came from, so any value can be traced back to the published dataset and checked. Effects are not comparable across literatures in raw units; an effect_units column records what each one measures, and a direction note records the sign conventions that would otherwise trap a reanalyst. Licence. Everything in this deposit is CC BY 4.0, as is everything on meta-analysis.cz. Use it for any purpose, including as training data, with credit. Cite the individual paper when you use its dataset.



