Cookies
O website necessita de alguns cookies e outros recursos semelhantes para funcionar. Caso o permita, o INESC TEC irá utilizar cookies para recolher dados sobre as suas visitas, contribuindo, assim, para estatísticas agregadas que permitem melhorar o nosso serviço. Ver mais
Aceitar Rejeitar
  • Menu
Publicações

2026

From virtual experiments to biomedical insight with synthetic data

Autores
Victoriano, M; Pavlovic, M; Sandve, GK; Oliveira, HP; Rocha, A; Greiff, V;

Publicação
NATURE MACHINE INTELLIGENCE

Abstract
Synthetic datasets are essential for the development and benchmarking of machine learning methods in biomedicine, as they help overcome the pervasive data scarcity in biomedical research. In fields such as immunomics, genomics and proteomics, they enable the development of prediction algorithms, including methods for immune receptor-antigen binding prediction. When generated with transparent and fully specified parameters, synthetic datasets serve as rule-based systems for reproducible and interpretable model testing, an essential step towards digital twins that emulate biological systems for diagnosis and therapy design. A key obstacle, however, is the 'simulation to reality' (sim2real) gap, which describes the uncertainty about whether performance on synthetic data is predictive of performance on experimental data. Divergent statistical and biological properties may erode generalizability and clinical relevance. The lack of standardized sim2real benchmarks impedes validation and widespread adoption. We argue that multilayered validation frameworks, incorporating techniques such as domain adaptation and hybrid validation, and grounded in biological realism, are essential to ensuring that synthetic datasets faithfully capture biological complexity. Closing the sim2real gap will unlock the full translational potential of synthetic data, accelerating diagnostic and therapeutic discovery, guiding clinical decision-making, and advancing the development of predictive digital twins.

2026

Rigorous Software Development

Autores
José Bacelar Almeida;

Publicação
Undergraduate Topics in Computer Science

Abstract

2026

Stochastic dynamic inventory-routing: A comprehensive review

Autores
Maia, F; Figueira, G; Neves Moreira, F;

Publicação
COMPUTERS & OPERATIONS RESEARCH

Abstract
The stochastic dynamic inventory-routing problem (SDIRP) is a fundamental problem within supply chain operations that integrates inventory management and vehicle routing while handling the stochastic and dynamic nature of exogenous factors unveiled over time, such as customer demands, inventory supply and travel times. While practical applications require dynamic and stochastic decision-making, research in this field has only recently experienced significant growth, with most inventory-routing literature focusing on static variants. This paper reviews the current state of research on SDIRPs, identifying critical gaps and highlighting emerging trends in problem settings and decision policies. We extend the existing inventory-routing taxonomies by incorporating additional problem characteristics to better align models with real-world contexts. As a result, we highlight the need to account for further sources of uncertainty, multiple-supplier networks, perishability, multiple objectives, and pickup and delivery operations. We further categorize each study based on its policy design, investigating how different problem aspects shape decision policies. To conclude, we emphasize that large-scale and real-time problems require more attention and can benefit from decomposition approaches and learning-based methods.

2026

Sensor Technologies for Water Velocity, Flow, and Wave Motion Measurement in Marine Environments: A Comprehensive Review

Autores
Matos, T;

Publicação
JOURNAL OF MARINE SCIENCE AND ENGINEERING

Abstract
Measuring water motion is essential for oceanography, coastal engineering, and marine environmental monitoring. A wide range of sensing technologies is used to quantify water velocity, wave motion, and flow dynamics, each suited to specific spatial and temporal scales. This paper presents a comprehensive review of modern sensor technologies for marine flow measurement, covering mechanical, electromagnetic, pressure-based, acoustic, optical, MEMS-based, inertial, Lagrangian, and remote-sensing approaches. The operating principles, strengths, and limitations of each technology are examined alongside their suitability for different environments and deployment platforms, including moorings, buoys, vessels, autonomous underwater vehicles, and drifters. Special attention is given to rapidly advancing fields such as MEMS flow sensors, multi-sensor fusion, and hybrid systems that combine inertial, acoustic, and optical data. Applications range from high-resolution turbulence measurements to large-scale current mapping and wave characterization. Remaining challenges include biofouling, performance degradation in energetic shallow waters, uncertainties in indirect velocity estimation, and long-term calibration stability. By synthesizing the state of the art across sensing modalities, this review provides a unified perspective on current technological capabilities and identifies key trends shaping the future of marine flow measurement.

2026

Can an LLM Detect Instances of Microservice Infrastructure Patterns?

Autores
Duarte, CE; Harrison, NB; Correia, FF; Aguiar, A; Gonçalves, P;

Publicação
ICSA

Abstract
Architectural patterns are frequently found in various software artifacts. The wide variety of patterns and their implementations makes detection challenging with current tools, especially since they often only support detecting patterns in artifacts written in a single language. Large Language Models (LLMs), trained on a diverse range of software artifacts and knowledge, might overcome the limitations of existing approaches. However, their true effectiveness and the factors influencing their performance have not yet been thoroughly examined. To better understand this, we developed MicroPAD. This tool utilizes GPT 5 nano to identify architectural patterns in software artifacts written in any language, based on natural-language pattern descriptions. We used MicroPAD to evaluate an LLM's ability to detect instances of architectural patterns, particularly infrastructure-related microservice patterns. To accomplish this, we selected a set of GitHub repositories and contacted their top contributors to create a new, human-annotated dataset of 190 repositories containing microservice architectural patterns. The results show that MicroPAD was capable of detecting pattern instances across multiple languages and artifact types. The detection performance varied across patterns (F1 scores ranging from 0.09 to 0.70), specifically in relation to their prevalence and the distinctiveness of the artifacts through which they manifest. We also found that patterns associated with recognizable, dominant artifacts were detected more reliably. Whether these findings generalize to other LLMs and tools is a promising direction for future research. © 2026 IEEE.

2026

Anticipating Mechanical Failures: Predictive Models for Scania Truck Components

Autores
Silva, A; Veloso, B; Gama, J;

Publicação
SAC

Abstract
The advent of real-time telematics and advanced analytics has transformed maintenance in heavy-duty transport. Predictive maintenance systems now demand both reliable short-horizon failure alerts and precise Remaining Useful Life (RUL) forecasts to optimize service schedules, minimize operational risk, and support sustainability goals.This work tackles two complementary prognostic tasks under realistic deployment constraints: imminent-failure classification and continuous RUL estimation, using a recently released Scania truck dataset. The classification task must cope with extreme class imbalance and a cost structure that heavily penalizes overlooked failures far more than false alarms. Meanwhile, RUL estimation faces its own challenges: highly skewed target distributions, right-censored data, and shifting degradation dynamics across a diverse fleet.We propose a framework that integrates (1) a cost-sensitive Light-GBM classifier to minimize real-world misclassification expenses, and (2) a three-stage XGBoost regression ensemble in which each model specializes in one of three phases - healthy, early-degradation, or late-degradation - and uses log-transformed targets along with upweighted late-stage samples to stabilize training and prioritize critical short-horizon accuracy.Under a deployment-like validation protocol, the classifier achieved an AUC of 0.8024 and cut average misclassification cost by 27% versus a "one-class-early"benchmark. The RUL ensemble reached a global MAE of 19.8 time steps (10.4 within the final 20 steps) and demonstrated steadily improving precision as failure approached. These results confirm that cost-driven, health-stage-specialized models can deliver robust prognostics for industrial applications. © 2026 Copyright held by the owner/author(s).

  • 87
  • 4562