Open-access Evaluation of the Annif automatic indexing system in information science journal articles

Objective:  This study evaluates the performance of the Annif automatic indexing tool in assigning descriptors to scientific articles within the field of Information Science.

Methods:  The research adopted an empirical and exploratory approach. It used a corpus composed of 60 articles written in Portuguese and previously indexed manually with the Brazilian Thesaurus of Information Science (TBCI). The corpus was divided into two subsets: 30 articles were used for training and 30 for testing. Annif was configured with the MLLM backend and compared with the terms assigned through intellectual indexing using the metrics of Precision, Recall, F1-score, and Consistency. A qualitative analysis of the generated terms was also conducted.

Results:  The system demonstrated moderate performance, achieving a Precision of 49.52%, Recall of 50.1%, F1-score of 49.23%, and Consistency of 31.63%.

Conclusions:  It is concluded that Annif is a viable tool for supporting semi-automatic indexing workflows, provided that its suggestions are validated by human indexers.

KEYWORDS:
Annif; machine learning; automatic indexing; digital repositories; Brazilian Thesaurus of Information Science

location_on
Universidade Federal de Santa Catarina Campus Universitário Reitor João David Ferreira Lima - Trindade. CEP-88040-900, Telefone: +55 (48) 3721-2237 - Florianópolis - SC - Brazil
E-mail: encontrosbibli@contato.ufsc.br
rss_feed Acompañe los números de esta revista en su lector de RSS
Ir para arriba Notificar error