Amir13/ncbi-persian
收藏资源简介:
--- annotations_creators: - expert-generated language: - fa language_creators: - machine-generated license: - other multilinguality: - monolingual pretty_name: ncbi-persian size_categories: - 1K<n<10K source_datasets: - extended|ncbi_disease tags: - named entity recognition task_categories: - token-classification task_ids: - named-entity-recognition train-eval-index: - col_mapping: ner_tags: target tokens: text config: ncbi_disease metrics: - name: Accuracy type: accuracy - args: average: macro name: F1 macro type: f1 - args: average: micro name: F1 micro type: f1 - args: average: weighted name: F1 weighted type: f1 - args: average: macro name: Precision macro type: precision - args: average: micro name: Precision micro type: precision - args: average: weighted name: Precision weighted type: precision - args: average: macro name: Recall macro type: recall - args: average: micro name: Recall micro type: recall - args: average: weighted name: Recall weighted type: recall splits: eval_split: test train_split: train task: token-classification --- # Dataset Card for Dataset Name ## Dataset Description - **Homepage:** - **Repository:** - **Paper:** - **Leaderboard:** - **Point of Contact:** ### Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using [this raw template](https://github.com/huggingface/huggingface_hub/blob/main/src/huggingface_hub/templates/datasetcard_template.md?plain=1). ### Supported Tasks and Leaderboards [More Information Needed] ### Languages [More Information Needed] ## Dataset Structure ### Data Instances [More Information Needed] ### Data Fields [More Information Needed] ### Data Splits [More Information Needed] ## Dataset Creation ### Curation Rationale [More Information Needed] ### Source Data #### Initial Data Collection and Normalization [More Information Needed] #### Who are the source language producers? [More Information Needed] ### Annotations #### Annotation process [More Information Needed] #### Who are the annotators? [More Information Needed] ### Personal and Sensitive Information [More Information Needed] ## Considerations for Using the Data ### Social Impact of Dataset [More Information Needed] ### Discussion of Biases [More Information Needed] ### Other Known Limitations [More Information Needed] ## Additional Information ### Dataset Curators [More Information Needed] ### Licensing Information [More Information Needed] ### Citation Information If you used the datasets and models in this repository, please cite it. ```bibtex @misc{https://doi.org/10.48550/arxiv.2302.09611, doi = {10.48550/ARXIV.2302.09611}, url = {https://arxiv.org/abs/2302.09611}, author = {Sartipi, Amir and Fatemi, Afsaneh}, keywords = {Computation and Language (cs.CL), Artificial Intelligence (cs.AI), FOS: Computer and information sciences, FOS: Computer and information sciences}, title = {Exploring the Potential of Machine Translation for Generating Named Entity Datasets: A Case Study between Persian and English}, publisher = {arXiv}, year = {2023}, copyright = {arXiv.org perpetual, non-exclusive license} } ``` ### Contributions [More Information Needed]
数据集概述
- 名称: ncbi-persian
- 语言: 波斯语 (fa)
- 语言生成方式: 机器生成
- 许可证: 其他
- 多语言性: 单语种
- 大小: 1K<n<10K
- 来源数据集: 扩展自 ncbi_disease
- 标签: 命名实体识别
- 任务类别: 令牌分类
- 任务ID: 命名实体识别
训练与评估指标
- 配置: ncbi_disease
- 指标:
- 准确率 (Accuracy)
- F1分数:
- 宏平均 (F1 macro)
- 微平均 (F1 micro)
- 加权平均 (F1 weighted)
- 精确度:
- 宏平均 (Precision macro)
- 微平均 (Precision micro)
- 加权平均 (Precision weighted)
- 召回率:
- 宏平均 (Recall macro)
- 微平均 (Recall micro)
- 加权平均 (Recall weighted)
- 数据分割:
- 训练集: train
- 评估集: test
- 任务: 令牌分类



