遇见数据集

Market Calibration Dataset: De-Vigged Betting Probabilities vs Real Outcomes, 36,476 Observations

收藏
Zenodo2026-08-02 更新2026-08-13 收录
官方服务:

资源简介:

# Market Calibration Dataset — de-vigged betting probabilities vs real outcomes **36,476 outcome observations (both sides of 18,238 settled two-outcome betting markets)** across five sports: each row is a vig-free market-implied probability and whether the outcome happened. A modern, multi-sport dataset for studying market calibration and the favourite–longshot bias — a literature that still leans heavily on decades-old horse-racing data. | | | |---|---| | **Rows** | 36,476 (both sides of every market — the dataset nets to 0.500/0.500 by construction) | | **Markets** | 18,238 settled two-outcome markets | | **Sports** | 5 (soccer, baseball, basketball, ice hockey, Australian rules) | | **Period** | November 2024 – July 2026 (month granularity) | | **Licence** | CC BY 4.0 | ## Columns | Column | Meaning | |---|---| | `event_month` | Month the event was played (YYYY-MM) | | `sport` | Sport key | | `market_family` | `game line` (match-level markets) or `player prop` | | `market_implied_probability` | The outcome's **de-vigged** market-implied probability (vig removed across the market's outcomes), rounded to 3 dp | | `outcome` | 1 if the outcome occurred, 0 otherwise | ## Construction — read before analysing - **Both sides of every market are included.** For every observed outcome at probability *p*, its complement appears at *1−p* with the opposite result. The dataset therefore nets to exactly 0.500/0.500 overall, and calibration deviations appear as an antisymmetric curve. This is the standard construction for bias analysis and means the dataset carries **no information about any bettor's or model's selections** — only about the market itself. - **Probabilities are de-vigged**, not raw. Raw implied probabilities contain the bookmaker's margin; these have had it removed across the market's outcomes. Note that de-vig method choice affects how margin is allocated across the probability range — a caveat relevant to interpreting tail behaviour, and itself a research-worthy question the data supports. - **Sampling:** markets are those covered by an analytics pipeline across five sports; within a market, inclusion of both sides is complete by construction. Rows carry no event, team, player, bookmaker or exact-date identifiers and cannot be mapped to any individual market or price. ## Headline findings - **Betting markets are impressively well calibrated**: across every probability decile, real outcome frequencies track de-vigged implied probabilities within about ±2 percentage points. - **A small, systematic tilt survives vig removal**: outcomes priced below ~50% slightly *overperform* their vig-free probabilities (+1 to +2 points) while favourites slightly underperform — consistent with bookmakers loading their margin most heavily onto longshot prices. In raw (vigged) prices this manifests as the classic favourite–longshot bias. - Player-prop markets show a larger tilt than match-level markets, peaking around the 30–40% / 60–70% probability bands (±5 points). ## Licence and citation CC BY 4.0 — free to use, including commercially, with attribution. > Bet Better (2026). *Market Calibration Dataset: de-vigged betting probabilities vs outcomes, > 36,476 observations.* https://betbetter.world/studies/market-calibration Maintained by [Bet Better](https://betbetter.world), which also publishes the [Bookmaker Margin Panel](https://doi.org/10.5281/zenodo.21755380), a [free sports model API](https://betbetter.world/api/) and an [open AFL dataset](https://doi.org/10.5281/zenodo.21612783). > No odds, prices, bookmaker identities or wagering data are included. 18+. Most people lose > money gambling. If gambling is causing you harm: gamblinghelponline.org.au.

提供机构:
Bet Better
创建时间:
2026-08-02
二维码
社区交流群
二维码
科研交流群
商业服务