2025
Autores
Vilaça, L; Viana, P;
Publicação
AIQAM 2025 - Proceedings of the 2nd ACM Workshop in AI-powered Question and Answering Systems, Co-Located with MM 2025
Abstract
Automatically generated text has become a popular method for increasing the volume of information used to pre-train audio-visual foundational models. Such data is often leveraged in large volumes of text corpora for pre-training, which can amplify errors caused by incorrectly labelled or inaccurate information. Since it introduces considerable noise, the need for automatic validation strategies for large text collections is becoming particularly important. We present a reference-free method of dialogue evaluation by examining Topic Consistency (TC) within the generated text. Our approach competes with state-of-the-art techniques and provides a more explainable method via questioning. By utilising knowledge databases like ConceptNet, we also explore expanding the topics’ semantic variations for performance improvement. Our contributions include: 1) An objective evaluation methodology for TC; 2) Extensive experiments on four benchmark datasets. Code and data are available on the project page: github.com/lvilaca16/lm-evaluation. © 2025 Copyright held by the owner/author(s)
The access to the final selection minute is only available to applicants.
Please check the confirmation e-mail of your application to obtain the access code.