This dataset was created using [LeRobot](https://github.com/huggingface/lerobot). ## Dataset Description - **Homepage:** https://tonyzhaozh.github.io/aloha/ - **Paper:** https://arxiv.org/abs/230
This dataset consists of 10.17 hours of annotated male voices in Shanghai dialect that is applicable for Text-to-Speech Synthesis, where 8,499 utterances collected from a 21-year-old man were containe
# FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon ## Data Statistics | Domain (#tokens/#samples) | Iteration 1