Taxonomic-Based Digital Humanities Publications Classification
Digital Humanities, Analysis by topics, Taxonomy, Bayesian models, Classification
Publications in scientific journals and in specialized conferences play a key role in expressing the topics of interest to authors and readers in a given field of knowledge. In this sense, the effort to organize scientific production is vital for the advancement of the dissemination of contents produced in an unequivocal, fast and safe way. Considering the current information flood caused by digital tools, the issue of automated classification becomes urgent and must be addressed in every repository or digital platform of scientific publications. Among other aspects, the use of a taxonomy stands out for its ability to add a hierarchical semantic element to the act of classifying or categorizing concepts and specific information that define the domain of a field of knowledge. Particularly in the field of Digital Humanities, the epistemological culture that has been built by its growing community has given rise to international projects that address the issue in an environment with additional challenges due to its strongly interdisciplinary profile. The objective of this dissertation is to use computational tools for analysis by topics of texts to develop an auxiliary method of lexical classification of publications supported by a taxonomy called {\em TaDiRAH -- Taxonomy of Digital Research Activities in the Humanities}. The proposed method can be seen as a combination of the semantic approach of taxonomy with the lexical approach of automated text analysis. Its categories are free and practical. However, it is not uncommon, and even expected by the interdisciplinary profile, that a publication can be classified into different categories of different levels or the same level of the taxonomy, thus creating overlaps. In addition, the number of publications already classified by the scientific community is still relatively small and, above all, extremely unbalanced between the taxonomy categories. These two aspects that characterize the available sample make the task of reliably classifying publications in Digital Humanities particularly difficult. We propose a method that combines Bayesian classification models from the literature with original approaches to deal with overlaps and imbalances between taxonomy categories. Results of computational experiments carried out with a universe of 443 publications showed that the proposed approaches are, in fact, capable of profoundly improving the performance of the classification methods use.