遇见数据集

Domain name availability benchmark

收藏
Zenodo2026-08-07 更新2026-08-13 收录
官方服务:

资源简介:

This dataset records whether each of 50,060 candidate domain names was registered or available at the time it was queried, measured directly against registry RDAP services: 5,006 labels crossed with 10 TLDs (.com, .net, .org, .io, .co, .ai, .app, .dev, .xyz, .tech), operated by 7 registry operators. Checks were performed between 2026-08-05T21:13Z and 2026-08-05T23:49Z, with a follow-up pass on 2026-08-06 re-querying 2,079 rate-limited names at 2-second pacing; all resolved, leaving no indeterminate verdicts. Availability is a snapshot, not a durable property: each check carries its own timestamp, every verdict describes registry state at that recorded moment, and names may have been registered or released since. Corpus. Labels come from a frozen benchmark corpus (version 1, frozen 2026-07-31 — the freeze date of the name list, not the measurement date), stratified into four groups: dictionary-words, compounds, invented, and patterns. Every label is checked against every TLD, a fully crossed design, so availability is directly comparable across TLDs rather than confounded by a different sample per TLD. Method. Registry RDAP endpoints were resolved via the IANA bootstrap registry (RFC 9224), with a hand-verified supplement for .io and .co; the source used for each TLD is recorded in the data. Verdicts are operational, not guarantees of registrability: an RDAP response of HTTP 404 is recorded as available, 200 as registered, and anything else as indeterminate rather than forced into either class. Each check records its own timestamp, attempt count, and RDAP HTTP status. For registered names, registration date, expiry date, last-changed date, and EPP domain status (RFC 8056) are additionally included — but only where the registry publishes them; absence of these fields means the registry did not expose them, not that they do not exist. Before any measurement, a control probe queries one known-registered domain per TLD, to catch an endpoint that responds but responds wrongly. Known bias. The label list is drawn from a dictionary, not a frequency list. The dictionary-words and compounds strata therefore over-represent archaic and obscure vocabulary, which is more available than common words. Read those two strata as an optimistic ceiling on availability. Do not label them "common English words". Availability rates should also be read per stratum and per TLD: three of the four strata are constructed rather than sampled from names in demand, so a figure pooled across the corpus reflects its composition rather than the domain market. Files. An NDJSON with one record per check — the primary data. 2,079 records carry a second source entry from the follow-up pass; the original rate-limited observation is retained rather than replaced. A JSON aggregate derived from it, fully reproducible from the NDJSON alone. Results, method and known bias are published at https://domain.yoga/data/availability.

提供机构:
Zenodo
创建时间:
2026-08-07
二维码
社区交流群
二维码
科研交流群
商业服务