Please use this identifier to cite or link to this item: http://hdl.handle.net/1942/41621
Full metadata record
DC FieldValueLanguage
dc.contributor.authorVANDEVOORT, Brecht-
dc.contributor.authorBEX, Geert Jan-
dc.contributor.authorCREVECOEUR, Jonas-
dc.contributor.authorNEVEN, Frank-
dc.date.accessioned2023-10-26T09:25:45Z-
dc.date.available2023-10-26T09:25:45Z-
dc.date.issued2023-
dc.date.submitted2023-10-16T11:22:24Z-
dc.identifier.citationEFSA Supporting Publications, 20 (8) (Art N° 8212E)-
dc.identifier.urihttp://hdl.handle.net/1942/41621-
dc.description.abstractThis report presents an overview of topic modelling and classification models in relation to four case studies in the EFSA project OC/EFSA/AMU/2020/02. As adequate document embeddings have a positive influence on the effectiveness of topic modelling as well as text classification, an extensive number of different possibilities for word and document embeddings are discussed. It was found that a multitude of increasingly more complex embeddings are readily available for off-the-shelf use. But as they are trained on large but mostly general text corpora, their utility for domain specific text varies. Fine tuning or creating document embeddings from scratch is only feasible in the presence of enough data and has an associated computational cost. For some domains (like scientific articles), pretrained embeddings are available. For topic modelling, we discuss standard techniques like non-negative matrix factorization and latent Dirichlet allocation as well as more recent methods based on clustering of document embeddings like Top2Vec and BERTopic. For text classification, we consider hierarchical text classification approaches combined with established techniques for text classification via document embeddings. We propose a selection of techniques for each of the case studies justifying their choice and present a plan for evaluation. Finally, we discuss our findings after having implemented and validated the selected techniques.-
dc.language.isoen-
dc.publisher-
dc.subject.otherNatural Language Processing-
dc.subject.otherTopic Modelling-
dc.subject.otherText Classification-
dc.titleTopic modelling and text classification models for applications within EFSA-
dc.typeJournal Contribution-
dc.identifier.issue8-
dc.identifier.volume20-
local.format.pages112-
local.bibliographicCitation.jcatA3-
local.type.refereedNon-Refereed-
local.type.specifiedArticle-
local.bibliographicCitation.artnr8212E-
dc.identifier.doi10.2903/sp.efsa.2023.EN-8212-
dc.identifier.eissn-
local.provider.typePdf-
local.uhasselt.internationalno-
item.fulltextWith Fulltext-
item.fullcitationVANDEVOORT, Brecht; BEX, Geert Jan; CREVECOEUR, Jonas & NEVEN, Frank (2023) Topic modelling and text classification models for applications within EFSA. In: EFSA Supporting Publications, 20 (8) (Art N° 8212E).-
item.accessRightsOpen Access-
item.contributorVANDEVOORT, Brecht-
item.contributorBEX, Geert Jan-
item.contributorCREVECOEUR, Jonas-
item.contributorNEVEN, Frank-
crisitem.journal.issn2397-8325-
Appears in Collections:Research publications
Files in This Item:
File Description SizeFormat 
published_version.pdfPublished version3.98 MBAdobe PDFView/Open
Show simple item record

Google ScholarTM

Check

Altmetric


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.