Publications

Publications by LIAAD

2022

Novel features for time series analysis: a complex networks approach

Authors
Silva, VF; Silva, ME; Ribeiro, P; Silva, F;

Publication
DATA MINING AND KNOWLEDGE DISCOVERY

Abstract
Being able to capture the characteristics of a time series with a feature vector is a very important task with a multitude of applications, such as classification, clustering or forecasting. Usually, the features are obtained from linear and nonlinear time series measures, that may present several data related drawbacks. In this work we introduce NetF as an alternative set of features, incorporating several representative topological measures of different complex networks mappings of the time series. Our approach does not require data preprocessing and is applicable regardless of any data characteristics. Exploring our novel feature vector, we are able to connect mapped network features to properties inherent in diversified time series models, showing that NetF can be useful to characterize time data. Furthermore, we also demonstrate the applicability of our methodology in clustering synthetic and benchmark time series sets, comparing its performance with more conventional features, showcasing how NetF can achieve high-accuracy clusters. Our results are very promising, with network features from different mapping methods capturing different properties of the time series, adding a different and rich feature set to the literature.

CloseRead Abstract

2022

Censored Multivariate Linear Regression Model

Authors
Sousa, R; Pereira, I; Silva, ME;

Publication
RECENT DEVELOPMENTS IN STATISTICS AND DATA SCIENCE, SPE2021

Abstract
Often, real-life problems require modelling several response variables together. This work analyses a multivariate linear regression model when the data are censored. Censoring distorts the correlation structure of the underlying variables and increases the bias of the usual estimators. Thus, we propose three methods to deal with multivariate data under left censoring, namely Expectation Maximization (EM), DataAugmentation (DA) and Gibbs Sampler with Data Augmentation (GDA). Results from a simulation study showthat both DA and GDA estimates are consistent for low and moderate correlation. Under high correlation scenarios, EM estimates present a lower bias.

CloseRead Abstract

2022

Statistical education and official statistics - training future data scientists

Authors
Silva, ME; Campos, P;

Publication
Proceedings of the IASE 2021 Satellite Conference

Abstract
EMOS (The European Master in Official Statistics) was set up to strengthen the collaboration within academia and producers of official statistics and help develop professionals able to work with European official data at different levels in the fast-changing production system of the 21st century. In this paper we address the need for training in Official Statistics, particularly in current times, where new skill sets and competencies are necessary. In particular, the needs for new data sources currently used by national statistical systems require the development of new methodologies. For that purpose, we do a matching between National Statistical Offices (NSO) needs and the offer from universities.

CloseRead Abstract

2022

NER in Archival Finding Aids: Extended

Authors
Cunha, LFD; Ramalho, JC;

Publication
MACHINE LEARNING AND KNOWLEDGE EXTRACTION

Abstract
The amount of information preserved in Portuguese archives has increased over the years. These documents represent a national heritage of high importance, as they portray the country's history. Currently, most Portuguese archives have made their finding aids available to the public in digital format, however, these data do not have any annotation, so it is not always easy to analyze their content. In this work, Named Entity Recognition solutions were created that allow the identification and classification of several named entities from the archival finding aids. These named entities translate into crucial information about their context and, with high confidence results, they can be used for several purposes, for example, the creation of smart browsing tools by using entity linking and record linking techniques. In order to achieve high result scores, we annotated several corpora to train our own Machine Learning algorithms in this context domain. We also used different architectures, such as CNNs, LSTMs, and Maximum Entropy models. Finally, all the created datasets and ML models were made available to the public with a developed web platform, NER@DI.

CloseRead Abstract

2022

NER in Archival Finding Aids: Next Level

Authors
Cunha, LFD; Ramalho, JC;

Publication
INFORMATION SYSTEMS AND TECHNOLOGIES, WORLDCIST 2022, VOL 2

Abstract
Currently, there is a vast amount of archival finding aids in Portuguese archives, however, these documents lack structure (are not annotated) making them hard to process and work with. In this way, we intend to extract and classify entities of interest, like geographical locations, people's names, dates, etc. For this, we will use an architecture that has been revolutionizing several NLP tasks, Transformers, presenting several models in order to achieve high results. It is also intended to understand what will be the degree of improvement that this new mechanism will present in comparison with previous architectures. Can Transformer-based models replace the LSTMs in NER? We intend to answer this question along this paper.

CloseRead Abstract

2022

Fine-Tuning BERT Models to Extract Named Entities from Archival Finding Aids

Authors
Costa Cunha, LF; Ramalho, JC;

Publication
Proceedings of the 26th International Conference on Theory and Practice of Digital Libraries - Workshops and Doctoral Consortium, Padua, Italy, September 20, 2022.

Abstract
In recent works, several NER models were developed to extract named entities from Portuguese Archival Finding Aids. In this paper, we are complementing the work already done by presenting a different NER model with a new architecture, Bidirectional Encoding Representation from Transformers (BERT). In order to do so, we used a BERT model that was pre-trained in Portuguese vocabulary and fine-tuned it to our concrete classification problem, NER. In the end, we compared the results obtained with previous architectures. In addition to this model we also developed an annotation tool that uses ML models to speed up the corpora annotation process. © 2022 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0)

CloseRead Abstract