2024
Authors
Teixeira, FB; Ricardo, M; Coelho, A; Oliveira, HP; Viana, P; Paulino, N; Fontes, H; Marques, P; Campos, R; Pessoa, LM;
Publication
CoRR
Abstract
2024
Authors
Sulun, S; Viana, P; Davies, MEP;
Publication
ISM
Abstract
We introduce VEMOCLAP: Video EMOtion Classifier using Pretrained features, the first readily available and open-source web application that analyzes the emotional content of any user-provided video. We improve our previous work, which exploits open-source pretrained models that work on video frames and audio, and then efficiently fuse the resulting pretrained features using multi-head cross-attention. Our approach increases the state-of-the-art classification accuracy on the Ekman-6 video emotion dataset by 4.3% and offers an online application for users to run our model on their own videos or YouTube videos. We invite the readers to try our application at serkansulun.com/app.
2024
Authors
Dias, J; Oliper, D; Soares, MR; Viana, P;
Publication
2024 IEEE 22ND MEDITERRANEAN ELECTROTECHNICAL CONFERENCE, MELECON 2024
Abstract
This paper addresses the critical challenge of optimising beacon placement to support indoor location services and proposes a methodology to optimise the Base Station (BS) coverage keeping or even improving the system precision. The algorithm builds on top of the building schematics and takes into account several aspects that affect the radio link range (materials attenuation, Line of Sight (LOS) conditions, transmitted power and radio sensibility). The outcome is available as a coverage heat map. It is then compared with a standard layout following existing expert guidelines to evaluate the efficacy of the proposed layout.
2024
Authors
Sulun, S; Viana, P; Davies, MEP;
Publication
EXPERT SYSTEMS WITH APPLICATIONS
Abstract
We introduce a novel method for movie genre classification, capitalizing on a diverse set of readily accessible pretrained models. These models extract high-level features related to visual scenery, objects, characters, text, speech, music, and audio effects. To intelligently fuse these pretrained features, we train small classifier models with low time and memory requirements. Employing the transformer model, our approach utilizes all video and audio frames of movie trailers without performing any temporal pooling, efficiently exploiting the correspondence between all elements, as opposed to the fixed and low number of frames typically used by traditional methods. Our approach fuses features originating from different tasks and modalities, with different dimensionalities, different temporal lengths, and complex dependencies as opposed to current approaches. Our method outperforms state-of-the-art movie genre classification models in terms of precision, recall, and mean average precision (mAP). To foster future research, we make the pretrained features for the entire MovieNet dataset, along with our genre classification code and the trained models, publicly available.
2024
Authors
Pereira, B; Cunha, B; Viana, P; Lopes, M; Melo, ASC; Sousa, ASP;
Publication
SENSORS
Abstract
Shoulder rehabilitation is a process that requires physical therapy sessions to recover the mobility of the affected limbs. However, these sessions are often limited by the availability and cost of specialized technicians, as well as the patient's travel to the session locations. This paper presents a novel smartphone-based approach using a pose estimation algorithm to evaluate the quality of the movements and provide feedback, allowing patients to perform autonomous recovery sessions. This paper reviews the state of the art in wearable devices and camera-based systems for human body detection and rehabilitation support and describes the system developed, which uses MediaPipe to extract the coordinates of 33 key points on the patient's body and compares them with reference videos made by professional physiotherapists using cosine similarity and dynamic time warping. This paper also presents a clinical study that uses QTM, an optoelectronic system for motion capture, to validate the methods used by the smartphone application. The results show that there are statistically significant differences between the three methods for different exercises, highlighting the importance of selecting an appropriate method for specific exercises. This paper discusses the implications and limitations of the findings and suggests directions for future research.
2024
Authors
Gouveia, F; Silva, V; Lopes, J; Moreira, RS; Torres, JM; Guerreiro, MS;
Publication
DISCOVER APPLIED SCIENCES
Abstract
Accurate and up-to-date information about the building stock is fundamental to better understand and mitigate the impact caused by catastrophic earthquakes, as seen recently in Turkey, Syria, Morocco and Afghanistan. Planning for such events is necessary to increase the resilience of the building stock and to minimize casualties and economic losses. Although in several parts of the world new constructions follow more strict compliance with modern seismic codes, a large proportion of existing building stock still demands a more detailed and automated vulnerability analysis. Hence, this paper proposes the use of computer vision deep learning models to automatically classify buildings and create large scale (city or region) exposure models. Such approach promotes the use of open databases covering most cities in the world (cf. OpenStreetMap, Google Street View, Bing Maps and satellite imagery), Therefore providing valuable geographical, topological and image data that may cheaply be used to extract valuable information to feed exposure models. Our previous work using deep learning models achieved, in line with the results from other projects, high classification accuracy concerning building materials and number of storeys. This paper extends the approach by: (i) implementing four CNN-based models to perform classification of three sets of different/extended buildings' characteristics; (ii) training and comparing the performance of the four models for each of the sets; (iii) comparing the risk assessment results based on data extracted from the best CNN-based model against the results obtained with traditional ground data. In brief, the best accuracy obtained with the three tested sets of buildings' characteristics is higher than 80%. Moreover, it is shown that the error resulting from using exposure models fed by automatic classification is not only acceptable, but also far outweighs the time and costs of obtaining a manual and specialised classification of building stocks. Finally, we recognize that automatic assessment of certain complex buildings' characteristics compares to similar limitations of traditional assessments performed by specialized civil engineers, typically related with the identification of the number of storeys and the construction material. However, the identified limitations do not show worse results when compared against the use of manual buildings' assessment. Implement an AI/ML framework for automating the collection of buildings' fa & ccedil;ades pictures annotated with several characteristics required by Exposure Models.Collect, process and filter a 4.239 pictures dataset of buildings' fa & ccedil;ades, which was made publicly available.Train, validate and test several Deep Learning models using 3 sets of building characteristics to produce exposure models with accuracies above 80%.Use heatmaps to show which image areas are more activated for a given prediction, thus helping to explain classification results.Compare simulation results using the predicted exposure model and a manually created exposure model, for the same set of buildings.
The access to the final selection minute is only available to applicants.
Please check the confirmation e-mail of your application to obtain the access code.