tgupj-tiny-router-data
收藏资源简介:
# tiny-router dataset Synthetic data for training and evaluating `tiny-router`, a compact multi-head classifier for short routing decisions. Each example contains a `current_text` field, optional interaction context, and four labels: - `relation_to_previous`: `new`, `follow_up`, `correction`, `confirmation`, `cancellation`, `closure` - `actionability`: `none`, `review`, `act` - `retention`: `ephemeral`, `useful`, `remember` - `urgency`: `low`, `medium`, `high` ## Files - `raw/synthetic.jsonl`: 2,907 synthetic examples before deduplication - `raw/synthetic.deduped.jsonl`: 2,892 deduplicated examples - `synthetic/train.jsonl`: 2,279 examples - `synthetic/validation.jsonl`: 276 examples - `synthetic/test.jsonl`: 337 examples The train, validation, and test splits are derived from the deduplicated file. About 37% of examples have no interaction context and only include `current_text` plus labels. ## Schema ```json { "current_text": "Actually next Monday", "interaction": { "previous_text": "Set a reminder for Friday", "previous_action": "created_reminder", "previous_outcome": "success", "recency_seconds": 45 }, "labels": { "relation_to_previous": "correction", "actionability": "act", "retention": "useful", "urgency": "medium" } } ``` `interaction` is optional and may be `null` or omitted. ## Intended use This dataset is meant for lightweight research and prototyping around message routing, update handling, memory policy, and prioritization for short text inputs. ## Limitations - The data is synthetic, not collected from production traffic. - Labels reflect the routing schema used in this repository and may not transfer cleanly to other products. - It should not be used on its own for high-stakes or fully autonomous decisions.



