Faithful and Fair Text Generation from Structured and Visual Data
收藏资源简介:
This thesis investigates text generation from diverse input types, including knowledge graphs, multimodal data, and medical images. Despite substantial advances in pre-trained language and vision–language models, core challenges persist—particularly unfaithful outputs, hallucinations, and fairness concerns. In knowledge graph–to–text generation, models may introduce content absent from the input. In multimodal settings, they often hallucinate details or overlook important visual features. Fairness remains especially critical in high-stakes domains such as medical report generation, where biased outputs can have serious consequences. This research develops approaches to produce more faithful, salient, and fair text while preserving strong overall performance.



