Accessing the Universe via Algorithm: A Baseline Study for Automating Alt-Text Generation with NASA Data and AI
收藏资源简介:
This is a frozen data repository containing the data derived for a paper with the same title. The abstract of this paper is below, in its current form. Small changes in the abstract may occur between this frozen version and the final publication; in that case, the final publication is to be considered the final versio. "Alt-text" is a textual substitute for visual information on a digital platform, accessed primarily via users of screen readers. Effective alt-text is crucial for inclusive and accessible communication; unfortunately, alt-text creation is often inhibited by the time and skill constraints of the humans writing it. This baseline study aimed to develop a reliable method for automating alt-text generation for publicly released science images from NASA's Chandra X-ray Observatory. We used pre-trained vision-language models and image metadata to generate alt-texts; however, we found the results to be unreliable with this method alone. We iterated the process to include fact-checking and evaluation filters, including a two-tier system of "exogenous” and "endogenous” checks to ensure that the alt-text accurately represented the image and image caption. We then evaluated the alt-texts by using an automated "rubric" to score the tool's ability to create "good" alt-text. This multistage approach is consistent with current AI research and demonstrates the role AI can play in creating alt-text for large scientific datasets, although it currently requires use alongside human expertise. In addition to performing the research, we add our real-world results to the Chandra X-ray Observatory public database to improve access to the images. Since the evaluation step yields a distance metric, our strategy enables adaptability to automated corrections or training improvements.



