遇见数据集

infinite-dataset-hub/AnagramMatrix

收藏
Hugging Face2024-08-29 更新2025-04-12 收录
官方服务:

资源简介:

--- license: mit tags: - infinite-dataset-hub - synthetic --- # AnagramMatrix tags: classification, text, ML-anagram _Note: This is an AI-generated dataset so its content may be inaccurate or false_ **Dataset Description:** The 'AnagramMatrix' dataset comprises pairs of strings, where each pair consists of two words that are anagrams of each other. Each row in the dataset contains a label indicating whether the words are anagrams ('Anagram') or not ('NotAnagram'). The dataset is designed for a classification task in machine learning, where the goal is to train a model to predict whether a given pair of words are anagrams. The dataset includes a variety of word lengths and complexity to provide a balanced challenge for ML algorithms. **CSV Content Preview:** ```csv word1,word2,label listen,silent,Anagram dormitory,dirty room,Anagram conversation,voices rant on,Anagram cinema,iceman,Anagram race,care,Anagram joker,kore,Anagram ``` The dataset structure is simple and ideal for supervised learning models, where the feature space consists of the two input words and the target is the label indicating an anagram relationship. This dataset could be used for various natural language processing tasks, including anagram detection and classification. **Source of the data:** The dataset was generated using the [Infinite Dataset Hub](https://huggingface.co/spaces/infinite-dataset-hub/infinite-dataset-hub) and microsoft/Phi-3-mini-4k-instruct using the query 'anagrams': - **Dataset Generation Page**: https://huggingface.co/spaces/infinite-dataset-hub/infinite-dataset-hub?q=anagrams&dataset=AnagramMatrix&tags=classification,+text,+ML-anagram - **Model**: https://huggingface.co/microsoft/Phi-3-mini-4k-instruct - **More Datasets**: https://huggingface.co/datasets?other=infinite-dataset-hub

license: MIT协议 tags: - 无限数据集中心(Infinite Dataset Hub) - 合成数据集 # AnagramMatrix数据集 tags: 分类任务、文本数据、ML-anagram * 注意:本数据集由人工智能生成,其内容可能存在不准确或虚假信息。 ## 数据集描述 AnagramMatrix数据集由字符串对组成,每对均为互为变位词(anagram)的两个单词。数据集中每一行均包含一个标签,用以标注两个单词是否为变位词,标签取值为"Anagram"(是变位词)或"NotAnagram"(非变位词)。本数据集专为机器学习(ML)分类任务设计,目标为训练模型以预测给定单词对是否互为变位词。数据集涵盖多种长度与复杂度的单词,可为机器学习算法提供难度均衡的测试场景。 ## CSV内容预览 csv word1,word2,label listen,silent,Anagram dormitory,dirty room,Anagram conversation,voices rant on,Anagram cinema,iceman,Anagram race,care,Anagram joker,kore,Anagram 本数据集结构简洁,非常适用于监督学习模型:其特征空间由两个输入单词组成,目标标签则用于标注单词间的变位词关系。本数据集可应用于多项自然语言处理(NLP)任务,例如变位词检测与分类任务。 ## 数据集来源 本数据集通过无限数据集中心(Infinite Dataset Hub)与microsoft/Phi-3-mini-4k-instruct模型,以查询词"anagrams"生成: - **数据集生成页面**:https://huggingface.co/spaces/infinite-dataset-hub/infinite-dataset-hub?q=anagrams&dataset=AnagramMatrix&tags=classification,+text,+ML-anagram - **模型**:https://huggingface.co/microsoft/Phi-3-mini-4k-instruct - **更多数据集**:https://huggingface.co/datasets?other=infinite-dataset-hub

提供机构:
infinite-dataset-hub
二维码
社区交流群
二维码
科研交流群
商业服务