遇见数据集

poojaruhal/RP-commenting-practices-social-media: RP-commenting-practices-social-media: RP_TOSEM_2020 v.1.0.1 Second release of of the replication Package for the paper "What do Developers Discuss about Code Comment Conventions on Social Media"

收藏
Zenodo2020-08-16 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

RP-commenting-practices-social-media Replication Package for the paper "What do Developers Discuss about Code Comment Conventions on Social Media?" Structure <pre><code>Paper-presenation.pdf Makar_tool/ Data/ stackoverfow_questions_with_answers_by_tags.csv stackoverfow_tags_metrics.csv apache_mailing_list.csv mailing_lists_ASF_@dev_@users_1.csv mailing_lists_ASF_@dev_@users_2.csv quora.csv sample_stackoverfow_questions_with_answers_by_tags.csv Schemas/ apache_mailing_lists.json quora.json stackoverfow_questions_answers_by_tag.json stackoverfow_tag_count.json stackoverfow_tag_metrics.json RQ1/ LDA_input/ stackoverfow_raw_dataset.csv LDA_output/ Mallet/ output_csv/ docs-in-topics.csv topic-words.csv topics-in-docs.csv topics-metadata.csv output_html/ all_topics.html Docs/ Topics/ RQ2/ datasource_rawdata/ mailing_lists_selection_criteria.csv quora.csv stackoverflow.csv manual_analysis_output/ stackoverflow_quora_taxonomy.xlsx </code></pre> Contents of the Replication Package <strong>Paper-presenation.pdf</strong> presents the highlights of the work in a presenation. <strong>Makar_tool/</strong> contains the data processed using the tool for the study <strong>Data/</strong> <code>stackoverfow_questions_with_answers_by_tags.csv</code> - all stackoverflow questions used in the study as stored in Makar <code>stackoverfow_tags_metrics.csv</code> - all data containing the calculations done for stackoverflow tag selection <code>apache_mailing_list.csv</code> - statistically significant sample of <code>mailing_lists_ASF_@dev_@users_1.csv</code> and <code>mailing_lists_ASF_@dev_@users_2.csv</code> used in the study <code>mailing_lists_ASF_@dev_@users_1.csv</code> - mailing list data used in the study as stored in Makar (part 1) <code>mailing_lists_ASF_@dev_@users_2.csv</code> - mailing list data used in the study as stored in Makar (part 2) <code>quora.csv</code> - all quora questions used in the study as stored in Makar <code>sample_stackoverfow_questions_with_answers_by_tags</code> - statistically significant sample of <code>stackoverfow_questions_with_answers_by_tags.csv</code> used in the study <strong>Schemas/</strong> <code>apache_mailing_lists.json</code> - data schema used in Makar to store mailing list data <code>quora.json</code> - data schema used in Makar to store quora data <code>stackoverfow_questions_answers_by_tag.json</code> - data schema used in Makar to store stackoverflow questions data <code>stackoverfow_tag_count.json</code> - data schema used in Makar to lookup number of questions per tag available in stackoverflow <code>stackoverfow_tag_metrics.json</code> - data schema used in Makar to stackoverflow tag metrics data <strong>RQ1/</strong> - contains the data used to answer RQ1 <strong>LDA_input/</strong> - input data used for LDA analysis <code>stackoverfow_raw_dataset.csv</code> - stackoverflow questions used to perform LDA analysis <strong>LDA_output/</strong> <strong>Mallet/</strong> - contains the LDA output generated by MALLET tool <strong>output_csv/</strong> <code>docs-in-topics.csv</code> - documents per topic <code>topic-words.csv</code> - most relevant topic words <code>topics-in-docs.csv</code> - topic probability per document <code>topics-metadata.csv</code> - metadata per document and topic probability <strong>output_html/</strong> - Browsable results of mallet output <code>all_topics.html</code> <code>Docs/</code> <code>Topics/</code> <strong>RQ2/</strong> - contains the data used to answer RQ2 <strong>datasource_rawdata/</strong> - contains the raw data for each source <code>mailing_lists_selection_criteria.csv</code> - criteria used to select mailing_lists. <code>quora.csv</code> - contains the processed dataset (like removing HTML tags). To know more about the preprocessing steps, please refer to the reproducibility section in the paper. The data is preprocessed using Makar tool. <code>stackoverflow.csv</code> - contains the processed stackoverflow dataset. To know more about the preprocessing steps, please refer to the reproducibility section in the paper. The data is preprocessed using Makar tool. <strong>manual_analysis_output/</strong> <code>stackoverflow_quora_taxonomy.xlsx</code> - contains the classified dataset of stackoverflow and quora and description of taxonomy. <code>Taxonomy</code> - contains the description of the first dimension and second dimension categories. Second dimension categories are further divided into levels, separated by <code>|</code> symbol. <code>stackoverflow-posts</code> - the questions are labelled relevant or irrelevant and categorized into the first dimension and second dimension categories. <code>quota-posts</code> - the questions are labelled relevant or irrelevant and categorized into the first dimension and second dimension categories.

提供机构:
Zenodo
创建时间:
2020-08-16
二维码
社区交流群
二维码
科研交流群
商业服务