tascib/turkish-llm-dataset
收藏资源简介:
--- license: cc-by-sa-4.0 language: - tr tags: - pretraining - turkish - nlp - language-modeling size_categories: - 100M<n<1B --- # Turkish Pretraining Corpus ## Dataset Description This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models. This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı University. ## Data Sources This dataset was constructed from the following sources: - BellaTurca https://huggingface.co/datasets/turkish-nlp-suite/BellaTurca - Cosmos-Turkish-Corpus-v1.0 https://huggingface.co/datasets/ytu-ce-cosmos/Cosmos-Turkish-Corpus-v1.0 - FineWeb-2 Turkish Categorized https://huggingface.co/datasets/altaidevorg/fineweb-2-turkish-categorized - FineWeb-2 (upstream dataset) https://huggingface.co/datasets/HuggingFaceFW/fineweb-2 ## Intended Use This dataset is primarily intended for the following purposes: - Turkish language model pretraining - Continual pretraining - Turkish NLP research - Academic and experimental use This dataset is more suitable for raw text pretraining than for supervised fine-tuning (SFT), as it does not consist of instruction-response pairs. ## Preprocessing - Merging the BellaTurca, Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized corpora - Cleaning inconsistent or malformed samples - Removing duplicate records - Applying additional filtering and normalization where necessary ## License This dataset is released under the **CC BY-SA 4.0** license for the original composition, preprocessing, and documentation created by this project. Notes: - The source datasets remain subject to their original license terms. - FineWeb-2 and FineWeb-2 Turkish Categorized are subject to the **ODC-By 1.0** license and require proper attribution. - Proper attribution is required for all included sources. - Please respect the upstream license terms when using, redistributing, or modifying this dataset. ## Limitations - May contain noise due to automatic collection and merging - May inherit biases from the source datasets - Additional cleaning and validation may be required depending on the use case ## Acknowledgements We would like to thank the creators of the following datasets: - BellaTurca contributors - Cosmos AI Research Group - HuggingFaceFW / FineWeb-2 contributors - altaidevorg / FineWeb-2 Turkish Categorized contributors ## Disclaimer No responsibility or liability is accepted for the use of this dataset.




