Cookies
O website necessita de alguns cookies e outros recursos semelhantes para funcionar. Caso o permita, o INESC TEC irá utilizar cookies para recolher dados sobre as suas visitas, contribuindo, assim, para estatísticas agregadas que permitem melhorar o nosso serviço. Ver mais
Aceitar Rejeitar
  • Menu
Publicações

Publicações por CTM

2024

Vision-Radio Experimental Infrastructure Architecture Towards 6G

Autores
Teixeira, FB; Ricardo, M; Coelho, A; Oliveira, HP; Viana, P; Paulino, N; Fontes, H; Marques, P; Campos, R; Pessoa, LM;

Publicação
CoRR

Abstract

2024

VEMOCLAP: A video emotion classification web application

Autores
Sulun, S; Viana, P; Davies, MEP;

Publicação
ISM

Abstract
We introduce VEMOCLAP: Video EMOtion Classifier using Pretrained features, the first readily available and open-source web application that analyzes the emotional content of any user-provided video. We improve our previous work, which exploits open-source pretrained models that work on video frames and audio, and then efficiently fuse the resulting pretrained features using multi-head cross-attention. Our approach increases the state-of-the-art classification accuracy on the Ekman-6 video emotion dataset by 4.3% and offers an online application for users to run our model on their own videos or YouTube videos. We invite the readers to try our application at serkansulun.com/app.

2024

Enhancing Indoor Localisation: a Bluetooth Low Energy (BLE) Beacon Placement approach

Autores
Dias, J; Oliper, D; Soares, MR; Viana, P;

Publicação
2024 IEEE 22ND MEDITERRANEAN ELECTROTECHNICAL CONFERENCE, MELECON 2024

Abstract
This paper addresses the critical challenge of optimising beacon placement to support indoor location services and proposes a methodology to optimise the Base Station (BS) coverage keeping or even improving the system precision. The algorithm builds on top of the building schematics and takes into account several aspects that affect the radio link range (materials attenuation, Line of Sight (LOS) conditions, transmitted power and radio sensibility). The outcome is available as a coverage heat map. It is then compared with a standard layout following existing expert guidelines to evaluate the efficacy of the proposed layout.

2024

Movie trailer genre classification using multimodal pretrained features

Autores
Sulun, S; Viana, P; Davies, MEP;

Publicação
EXPERT SYSTEMS WITH APPLICATIONS

Abstract
We introduce a novel method for movie genre classification, capitalizing on a diverse set of readily accessible pretrained models. These models extract high-level features related to visual scenery, objects, characters, text, speech, music, and audio effects. To intelligently fuse these pretrained features, we train small classifier models with low time and memory requirements. Employing the transformer model, our approach utilizes all video and audio frames of movie trailers without performing any temporal pooling, efficiently exploiting the correspondence between all elements, as opposed to the fixed and low number of frames typically used by traditional methods. Our approach fuses features originating from different tasks and modalities, with different dimensionalities, different temporal lengths, and complex dependencies as opposed to current approaches. Our method outperforms state-of-the-art movie genre classification models in terms of precision, recall, and mean average precision (mAP). To foster future research, we make the pretrained features for the entire MovieNet dataset, along with our genre classification code and the trained models, publicly available.

2024

A Machine Learning App for Monitoring Physical Therapy at Home

Autores
Pereira, B; Cunha, B; Viana, P; Lopes, M; Melo, ASC; Sousa, ASP;

Publicação
SENSORS

Abstract
Shoulder rehabilitation is a process that requires physical therapy sessions to recover the mobility of the affected limbs. However, these sessions are often limited by the availability and cost of specialized technicians, as well as the patient's travel to the session locations. This paper presents a novel smartphone-based approach using a pose estimation algorithm to evaluate the quality of the movements and provide feedback, allowing patients to perform autonomous recovery sessions. This paper reviews the state of the art in wearable devices and camera-based systems for human body detection and rehabilitation support and describes the system developed, which uses MediaPipe to extract the coordinates of 33 key points on the patient's body and compares them with reference videos made by professional physiotherapists using cosine similarity and dynamic time warping. This paper also presents a clinical study that uses QTM, an optoelectronic system for motion capture, to validate the methods used by the smartphone application. The results show that there are statistically significant differences between the three methods for different exercises, highlighting the importance of selecting an appropriate method for specific exercises. This paper discusses the implications and limitations of the findings and suggests directions for future research.

2024

Automated identification of building features with deep learning for risk analysis

Autores
Gouveia, F; Silva, V; Lopes, J; Moreira, RS; Torres, JM; Guerreiro, MS;

Publicação
DISCOVER APPLIED SCIENCES

Abstract
Accurate and up-to-date information about the building stock is fundamental to better understand and mitigate the impact caused by catastrophic earthquakes, as seen recently in Turkey, Syria, Morocco and Afghanistan. Planning for such events is necessary to increase the resilience of the building stock and to minimize casualties and economic losses. Although in several parts of the world new constructions follow more strict compliance with modern seismic codes, a large proportion of existing building stock still demands a more detailed and automated vulnerability analysis. Hence, this paper proposes the use of computer vision deep learning models to automatically classify buildings and create large scale (city or region) exposure models. Such approach promotes the use of open databases covering most cities in the world (cf. OpenStreetMap, Google Street View, Bing Maps and satellite imagery), Therefore providing valuable geographical, topological and image data that may cheaply be used to extract valuable information to feed exposure models. Our previous work using deep learning models achieved, in line with the results from other projects, high classification accuracy concerning building materials and number of storeys. This paper extends the approach by: (i) implementing four CNN-based models to perform classification of three sets of different/extended buildings' characteristics; (ii) training and comparing the performance of the four models for each of the sets; (iii) comparing the risk assessment results based on data extracted from the best CNN-based model against the results obtained with traditional ground data. In brief, the best accuracy obtained with the three tested sets of buildings' characteristics is higher than 80%. Moreover, it is shown that the error resulting from using exposure models fed by automatic classification is not only acceptable, but also far outweighs the time and costs of obtaining a manual and specialised classification of building stocks. Finally, we recognize that automatic assessment of certain complex buildings' characteristics compares to similar limitations of traditional assessments performed by specialized civil engineers, typically related with the identification of the number of storeys and the construction material. However, the identified limitations do not show worse results when compared against the use of manual buildings' assessment. Implement an AI/ML framework for automating the collection of buildings' fa & ccedil;ades pictures annotated with several characteristics required by Exposure Models.Collect, process and filter a 4.239 pictures dataset of buildings' fa & ccedil;ades, which was made publicly available.Train, validate and test several Deep Learning models using 3 sets of building characteristics to produce exposure models with accuracies above 80%.Use heatmaps to show which image areas are more activated for a given prediction, thus helping to explain classification results.Compare simulation results using the predicted exposure model and a manually created exposure model, for the same set of buildings.

  • 42
  • 407