Please use this identifier to cite or link to this item: http://hdl.handle.net/1942/41621
Title: Topic modelling and text classification models for applications within EFSA
Authors: VANDEVOORT, Brecht 
BEX, Geert Jan 
CREVECOEUR, Jonas 
NEVEN, Frank 
Issue Date: 2023
Publisher: 
Source: EFSA Supporting Publications, 20 (8) (Art N° 8212E)
Abstract: This report presents an overview of topic modelling and classification models in relation to four case studies in the EFSA project OC/EFSA/AMU/2020/02. As adequate document embeddings have a positive influence on the effectiveness of topic modelling as well as text classification, an extensive number of different possibilities for word and document embeddings are discussed. It was found that a multitude of increasingly more complex embeddings are readily available for off-the-shelf use. But as they are trained on large but mostly general text corpora, their utility for domain specific text varies. Fine tuning or creating document embeddings from scratch is only feasible in the presence of enough data and has an associated computational cost. For some domains (like scientific articles), pretrained embeddings are available. For topic modelling, we discuss standard techniques like non-negative matrix factorization and latent Dirichlet allocation as well as more recent methods based on clustering of document embeddings like Top2Vec and BERTopic. For text classification, we consider hierarchical text classification approaches combined with established techniques for text classification via document embeddings. We propose a selection of techniques for each of the case studies justifying their choice and present a plan for evaluation. Finally, we discuss our findings after having implemented and validated the selected techniques.
Keywords: Natural Language Processing;Topic Modelling;Text Classification
Document URI: http://hdl.handle.net/1942/41621
ISSN: 2397-8325
DOI: 10.2903/sp.efsa.2023.EN-8212
Category: A3
Type: Journal Contribution
Appears in Collections:Research publications

Files in This Item:
File Description SizeFormat 
published_version.pdfPublished version3.98 MBAdobe PDFView/Open
Show full item record

Google ScholarTM

Check

Altmetric


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.