Cookies
O website necessita de alguns cookies e outros recursos semelhantes para funcionar. Caso o permita, o INESC TEC irá utilizar cookies para recolher dados sobre as suas visitas, contribuindo, assim, para estatísticas agregadas que permitem melhorar o nosso serviço. Ver mais
Aceitar Rejeitar
  • Menu
Publicações

Publicações por LIAAD

2025

METAFORE: algorithm selection for decomposition-based forecasting combinations

Autores
Santos, M; de Carvalho, A; Soares, C;

Publicação
INTERNATIONAL JOURNAL OF DATA SCIENCE AND ANALYTICS

Abstract
Time series forecasting is an important tool for planning and decision-making. Considering this, several forecasting algorithms can be used, with results depending on the characteristics of the time series. The recommendation of the most suitable algorithm is a frequent concern. Metalearning has been successfully used to recommend the best algorithm for a time series analysis task. Additionally, it has been shown that decomposition methods can lead to better results. Based on previously published studies, in the experiments carried out, time series components were used. This work proposes and empirically evaluates METAFORE, a new time series forecasting approach that uses seasonal trend decomposition with Loess and metalearning to recommend suitable algorithms for time series forecasting combinations. Experimental results show that METAFORE can obtain a better predictive performance than single models with statistical significance. In the experiments, METAFORE also outperformed models widely used in the state-of-the-art, such as the long short-term memory neural network architectures, in more than 70%\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$70\%$$\end{document} of the time series tested. Finally, the results show that the joint use of metalearning and time series decomposition provides a competitive approach to time series forecasting.

2025

L-GTA: Latent Generative Modeling for Time Series Augmentation

Autores
Roque, L; Soares, C; Cerqueira, V; Torgo, L;

Publicação
CoRR

Abstract

2025

Simulating Biases for Interpretable Fairness in Offline and Online Classifiers

Autores
Inácio, R; Kokkinogenis, Z; Cerqueira, V; Soares, C;

Publicação
CoRR

Abstract
Predictive models often reinforce biases which were originally embedded in their training data, through skewed decisions. In such cases, mitigation methods are critical to ensure that, regardless of the prevailing disparities, model outcomes are adjusted to be fair. To assess this, datasets could be systematically generated with specific biases, to train machine learning classifiers. Then, predictive outcomes could aid in the understanding of this bias embedding process. Hence, an agent-based model (ABM), depicting a loan application process that represents various systemic biases across two demographic groups, was developed to produce synthetic datasets. Then, by applying classifiers trained on them to predict loan outcomes, we can assess how biased data leads to unfairness. This highlights a main contribution of this work: a framework for synthetic dataset generation with controllable bias injection. We also contribute with a novel explainability technique, which shows how mitigations affect the way classifiers leverage data features, via second-order Shapley values. In experiments, both offline and online learning approaches are employed. Mitigations are applied at different stages of the modelling pipeline, such as during pre-processing and in-processing.

2025

Generating Large Semi-Synthetic Graphs of Any Size

Autores
Tuna, R; Soares, C;

Publicação
CoRR

Abstract

2025

Modelradar: aspect-based forecast evaluation

Autores
Cerqueira, V; Roque, L; Soares, C;

Publicação
MACHINE LEARNING

Abstract
Accurate evaluation of forecasting models is essential for ensuring reliable predictions. Current practices for evaluating and comparing forecasting models focus on summarising performance into a single score, using metrics such as SMAPE. While convenient, averaging performance over all samples dilutes relevant information about model behaviour under varying conditions. This limitation is especially problematic for time series forecasting, where multiple layers of averaging-across time steps, horizons, and multiple time series in a dataset-can mask relevant performance variations. We address this limitation by proposing ModelRadar, a framework for evaluating univariate time series forecasting models across multiple aspects, such as stationarity, presence of anomalies, or forecasting horizons. We demonstrate the advantages of this framework by comparing 24 forecasting methods, including classical approaches and different machine learning algorithms. PatchTST, a state-of-the-art transformer-based neural network architecture, performs best overall but its superiority varies with forecasting conditions. For instance, concerning the forecasting horizon, we found that PatchTST (and also other neural networks) only outperforms classical approaches for multi-step ahead forecasting. Another relevant insight is that classical approaches such as ETS or Theta are notably more robust in the presence of anomalies. These and other findings highlight the importance of aspect-based model evaluation for both practitioners and researchers. ModelRadar is available as a Python package.

2025

Reducing algorithm configuration spaces for efficient search

Autores
Freitas, F; Brazdil, P; Soares, C;

Publicação
INTERNATIONAL JOURNAL OF DATA SCIENCE AND ANALYTICS

Abstract
Many current AutoML platforms include a very large space of alternatives (the configuration space). This increases the probability of including the best one for any dataset but makes the task of identifying it for a new dataset more difficult. In this paper, we explore a method that can reduce a large configuration space to a significantly smaller one and so help to reduce the search time for the potentially best algorithm configuration, with limited risk of significant loss of predictive performance. We empirically validate the method with a large set of alternatives based on five ML algorithms with different sets of hyperparameters and one preprocessing method (feature selection). Our results show that it is possible to reduce the given search space by more than one order of magnitude, from a few thousands to a few hundred items. After reduction, the search for the best algorithm configuration is about one order of magnitude faster than on the original space without significant loss in predictive performance.

  • 20
  • 529