2025
Autores
Vasconcelos, M; Cavique, L;
Publicação
EXPERT SYSTEMS WITH APPLICATIONS
Abstract
Imbalanced datasets present a challenge in machine learning, especially in binary classification scenarios where one class significantly outweighs the other. This imbalance often leads to models favoring the majority class, resulting in inadequate predictions for the minority class, specifically in false negatives. In response to this issue, this work introduces the MinFNR ensemble algorithm, designed to minimize False Negative Rates (FNR) in imbalanced datasets. The new approach strategically combines data-level, algorithmic-level, and hybrid-level approaches to enhance overall predictive capabilities while minimizing computational resources using the Set Covering Problem (SCP) formulation. Through a comprehensive evaluation of diverse datasets, MinFNR consistently outperforms individual algorithms, showing its potential for applications where the cost of false negatives is substantial, such as fraud detection and medical diagnosis. This work also contributes to ongoing efforts to improve the reliability and effectiveness of machine learning algorithms in real imbalanced scenarios.
2025
Autores
Figueiredo, L; Pinheiro, P; Cavique, L; Marques, N;
Publicação
Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Abstract
This study introduces a system that helps non-expert users find information easily without knowing database languages or asking technicians for help. A specific domain is explored, focusing on a subscrip- tion-based sports facility, which serves as an open-source version of a real case study. Utilizing the star schema, the available data in the database is structured to provide accessibility through Portuguese Natural Language queries. Using a Large Language Model (LLM), SQL queries are generated based on the question and the provided star schema. We created a dataset with 115 highly challenging questions drawn from real-world usage scenarios to validate the correctness of the system. Challenges found during testing, like attribute value interpretation, out-of-scope questions, and temporal interval adequacy issues, highlight the insufficiency of the star schema alone in providing the needed context for generating accurate SQL queries by the LLM. Addressing these challenges through enhanced contextual information shows significant improvement in query correctness, with validation results increasing from 57.76% to 88.79%. This study shows the potential and limitations of LLMs in generating SQL queries from Portuguese Natural Language queries. © The Author(s), under exclusive license to Springer Nature Switzerland AG 2025.
2025
Autores
Guerino, LR; Rizzo Vincenzi, AM;
Publicação
SBQS
Abstract
2025
Autores
Monteiro, CEO; Guerino, LR; Fernandes, GF; Pereira, MH; Souza-Zinader, JPd; Braga, RD; Pocivi, VCB; Vincenzi, AMR;
Publicação
Proceedings of the 31st Brazilian Symposium on Multimedia and the Web (WebMedia 2025)
Abstract
2025
Autores
Moreira, I; Adolfo, LB; Melegati, J; Choma, J; Guerra, E; Zaina, L;
Publicação
XP
Abstract
2025
Autores
Pancher, JC; Melegati, J; Guerra, EM;
Publicação
XP
Abstract
The access to the final selection minute is only available to applicants.
Please check the confirmation e-mail of your application to obtain the access code.