遇见数据集

funpedia

收藏
OpenML2025-02-24 更新2025-12-20 收录
官方服务:

资源简介:

Funpedia- Funpedia (Miller et al., 2017) contains rephrased Wikipedia sentences in a more conversational way. The curators retained only biography related sentences and annotate similar to Wikipedia, to give ABOUT labels. The Multi-Dimensional Gender Bias Classification dataset is based on a general framework that decomposes gender bias in text along several pragmatic and semantic dimensions: bias from the gender of the person being spoken about, bias from the gender of the person being spoken to, and bias from the gender of the speaker. It contains seven large scale datasets automatically annotated for gender information (there are eight in the original project but the Wikipedia set is not included in the HuggingFace distribution), one crowdsourced evaluation benchmark of utterance-level gender rewrites, a list of gendered names, and a list of gendered words in English. text-classification-other-gender-bias: The dataset can be used to train a model for classification of various kinds of gender bias. The model performance is evaluated based on the accuracy of the predicted labels as compared to the given labels in the dataset. Dinan et al's (2020) Transformer model achieved an average of 67.13 accuracy in binary gender prediction across the ABOUT, TO, and AS tasks. This is the dataset 'funpedia', it description is as follows: text: the text to be classified. gender(target): a classification label, with possible values including gender-neutral (0), female (1), male (2), indicating the gender of the person being talked about. persona: a string describing the persona assigned to the user when talking about the entity. title: a string naming the entity the text is about. paper_url = "https://arxiv.org/pdf/1509.01626" original_data_url = "https://huggingface.co/datasets/facebook/md_gender_bias/tree/10c34c50ef78b4a42f6d4eeac80a0ef2d190cd07/funpedia"

创建时间:
2025-02-24
二维码
社区交流群
二维码
科研交流群
商业服务