遇见数据集

医学影像教学病例题库数据

收藏
浙江省数据知识产权登记平台2025-01-08 更新2025-01-09 收录
官方服务:

资源简介:

在校医学生、规培医生、在职影像科医生均可以通过医学影像教学病例题库系统实现在线学习和管理。对于在校医学院学生,解决学校实践困难,学习资料长时间没有更新升级问题。对于规培医生,解决带教医生没有足够时间和精力去整理学习内容,同时帮助规培医生在短时间内集中学习重点知识,从而顺利完成放射科室轮转任务。对于在职影像科医生,拓展了医生知识面,降低医生查找和管理学习资料时间成本。1.数据采集:由医生在线提交题目信息,包括题目标题、题目国际码、科室名称、题型、创建人、创建时间、所属单位等原始数据字段,题目国际码、题目ID等由系统自动生成,并对相关字段进行脱敏处理。 2.算法规则: 1)使用UIE算法和基于同义词的正则表达式查询,计算出同类型器官名称和疾病名称数量,一方面排除重复数据,另一方面根据同类型疾病数量比例判定题目级别,占比大于50%的分为初级题目,占比20%-50%为中级题目,小于20%的分为高级题目。 2)通过预训练的隐马尔可夫链结合CRF算法从病例分析中自动提取生成器官名称和疾病名称。首先,将原始病例分析内容转成序列Z={z0, z1, …, zm}隐马尔可夫链计算总词表S;设定字的概率,假定前一个字z(t-1)情况下,当前字zt分别为词首字、词中字、词尾字的概率P(zt | z(t-1), S),并选取概率最高的类别作为当前字zt的类别;按照词首字+多个词中子+词尾字,实现病例分析内容分词,得到分词序列W={w0, w1, …, wl};再通过CRF算法识别2类实体:器官/部位、疾病/征象,针对每一个词wk,CRF在已知w(k-1)类别时输出wk为上述2类实体的概率P(wk | w(k-1)),从而得到每一个词的类别;最后,基于影像解剖学和诊断学知识图谱,利用Jaro-Winkler 距离将器官/部位与所有规范化器官名词计算相似度,并选择相似度最高的器官名词作为器官名称。同理,将疾病/征象与所有规范化疾病名词计算相似度得到疾病名词进行实体对齐,疾病名称的提取。

All medical students in school, resident physicians, and practicing radiologists can conduct online learning and management via the medical imaging teaching case question bank system. For medical college students, this system addresses the issues of insufficient clinical practice opportunities and outdated learning materials that have not been updated for a long time. For resident physicians, it solves the problem that preceptors do not have enough time and energy to organize learning content, and helps them focus on key knowledge in a short period to successfully complete their radiology department rotation tasks. For practicing radiologists, it expands their knowledge scope and reduces the time and cost spent on searching and managing learning materials. 1. Data Collection: Doctors submit question information online, including original data fields such as question title, question international code, department name, question type, creator, creation time, and affiliated institution. The question international code, question ID, etc. are automatically generated by the system, and relevant fields are desensitized. 2. Algorithm Rules: 1) Use the UIE algorithm and synonym-based regular expression queries to calculate the number of organ names and disease names of the same type. On one hand, this eliminates duplicate data; on the other hand, it determines the question level based on the proportion of same-type diseases. Questions with a proportion greater than 50% are classified as primary questions, those with a proportion between 20% and 50% are intermediate questions, and those with a proportion less than 20% are advanced questions. 2) Automatically extract and generate organ names and disease names from case analysis using a pre-trained hidden Markov chain combined with the CRF algorithm. First, convert the original case analysis content into a sequence Z={z₀, z₁, …, zₘ}, and use the hidden Markov chain to calculate the total vocabulary S. Set the word probability: assuming that given the previous character z_(t-1), the probability P(zₜ | z_(t-1), S) that the current character zₜ is the word initial character, word middle character, or word end character respectively, and select the category with the highest probability as the category of the current character zₜ. Perform word segmentation on the case analysis content according to the pattern of word initial character + multiple word middle characters + word end character, to obtain the word segmentation sequence W={w₀, w₁, …, wₗ}. Then, use the CRF algorithm to identify two types of entities: organs/parts and diseases/signs. For each word wₖ, the CRF outputs the probability P(wₖ | w_(k-1)) that wₖ belongs to the two types of entities given the category of w_(k-1), thereby obtaining the category of each word. Finally, based on the knowledge graph of imaging anatomy and diagnostics, use the Jaro-Winkler distance to calculate the similarity between organs/parts and all standardized organ nouns, and select the standardized organ noun with the highest similarity as the final organ name. Similarly, calculate the similarity between diseases/signs and all standardized disease nouns to obtain standardized disease names for entity alignment and disease name extraction.

创建时间:
2024-12-04
搜集汇总
数据集介绍
医学影像教学病例题库数据 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务