遇见数据集

<b>Lexical Demands and Features of English Textbooks for Vietnamese 10th Graders: An In-depth Comparison of Listening Sections</b>

收藏
Figshare2024-01-07 更新2026-04-08 收录
官方服务:

资源简介:

<b>Research Design</b>In this research, the quantitative method was employed to measure the lexical coverage and lexical richness of the English textbook series. The AntwordProfiler version 2.1.0 (Anthony, 2023) accompanied with BNC/COCA lists created by Nation (2017) was run. Then, it calculated a sort of lexical diversity indices with Tool for the Automatic Analysis of Lexical Diversity (TAALED) (Kyle, Crossley &amp; Jarvis, 2021). The study also employed Jamovi version 2.3.28 to examine whether these textbooks significantly varied through pairwise comparisons.<b>Data Collection</b>In the period of 2022-2023, the MoET released nine new English textbook series for grade 10 students. Unfortunately, <i>Macmillan Move On</i> was excluded from the data source of this study due to its limited popularity. As mentioned earlier, the key data was the listening sections in those English textbooks. The corpus included a total of 54,566 tokens. The size of units and passages were different from each other.<b>Data Preparation</b>The listening transcripts were mostly taken from teacher’s books of these series. Some of them were collected from the subtitles of the audios. Then, a data extraction tool from Microsoft named PowerToys was applied to convert raw texts from the data sources into text files. However, there were some spelling errors due to the unoptimized function of the tool. It could be seen that the characters “m” and “rn” were quite similar, which caused some misinterpretations in the scanning process of the tool. Besides, some uppercase characters such as “I” and “O” were converted as the numbers “1” and “0”. Thus, the researcher retyped the errors manually, removed redundant words, and re-added missing words to the corpus.<b>Data Analysis</b>The AntWordProfiler (Anthony, 2023) was utilised to process all of the text files. The program used the BNC/COCA list (Nation, 2017) to classify words into their frequency levels, count their occurrences, and calculate their sums and percentages. Tokens were arranged in 25 levels and four supplementary lists, moreover, outlying tokens would be put in “Not in the list”. Then, those not-in-the-list words would be returned to their correct positions in the four supplementary lists. The updated data would be processed to report the coverage of each full textbook, and each unit in each textbook was further analysed independently.<i>Research question 2</i> could be answered based on the results of each unit in terms of length, sophistication and diversity. Length was counted by the total tokens of each unit. At the same time, lexical sophistication was calculated according to the sum of word families from the 4,000 to 25,000 levels, and diversity was computed on TAALED using the three lexical diversity indices: MATTR50, HD-D, and MTLD Original. In addition, each textbook was coded according to alphabetic characters, and units in the same textbook series were coded with the same letter. Finally, the data were run by Jamovi to make pairwise comparisons relating to length, sophistication and diversity.

提供机构:
Mai, Nhi
创建时间:
2024-01-07
二维码
社区交流群
二维码
科研交流群
商业服务