Aesthetic Trends and Semantic Web Adoption of Media Outlets Identified through Automated Archival Data Extraction
收藏资源简介:
This dataset includes a variety of structured data gathered via various Web data extraction techniques which were employed in order to collect current and archival data from almost a thousand news websites that are popular in Greece, for the purpose of monitoring and recording their progress through time. The collected information, that took the form of a website’s source code and an impression of their homepage in different time instances of the last decade, has been used to identify trends concerning Semantic Web integration, DOM structure complexity, number of graphics, color usage and more. In total more than ten thousands impressions (including screenshots and source code) were analyzed which resulted to conclusions regarding the evolution of aesthetics and the adoption of new technologies.
本数据集包含多种结构化数据,这些数据通过各类网页数据抽取技术采集而来,采集对象为希腊近千家主流新闻网站,旨在获取其现时与历史存档数据,以追踪并记录这些网站随时间推移的发展历程。采集到的信息以网站源代码以及近十年不同时间节点下的首页快照形式存在,已被用于识别与语义网(Semantic Web)集成、文档对象模型(Document Object Model,DOM)结构复杂度、图形数量、色彩使用等相关的发展趋势。总计共分析了超过一万份快照(包含截图与源代码),最终得出了关于网站美学演进与新技术采纳情况的相关结论。




