2023
Autores
Bernardo, G; Bernardes, G;
Publicação
Pers. Ubiquitous Comput.
Abstract
2023
Autores
Forero, J; Bernardes, G; Mendes, M;
Publicação
AIMC
Abstract
https://aimc2023.pubpub.org/pub/9z68g7d2 Music has been commonly recognized as a means of expressing emotions. In this sense, an intense debate emerges from the need to verbalize musical emotions. This concern seems highly relevant today, considering the exponential growth of natural language processing using deep learning models where it is possible to prompt semantic propositions to generate music automatically. This scoping review aims to analyze and discuss the possibilities of music generation conditioned by emotions. To address this topic, we propose a historical perspective that encompasses the different disciplines and methods contributing to this topic. In detail, we review two main paradigms adopted in automatic music generation: rules-based and machine-learning models. Of note are the deep learning architectures that aim to generate high-fidelity music from textual descriptions. These models raise fundamental questions about the expressivity of music, including whether emotions can be represented with words or expressed through them. We conclude that overcoming the limitation and ambiguity of language to express emotions through music, some of the use of deep learning with natural language has the potential to impact the creative industries by providing powerful tools to prompt and generate new musical works.
2023
Autores
Carvalho, N; Bernardes, G;
Publicação
AIMC
Abstract
https://aimc2023.pubpub.org/pub/latent-spaces-tonal-music Variational Autoencoders (VAEs) have proven to be effective models for producing latent representations of cognitive and semantic value. We assess the degree to which VAEs trained on a prototypical tonal music corpus of 371 Bach's chorales define latent spaces representative of the circle of fifths and the hierarchical relation of each key component pitch as drawn in music cognition. In detail, we compare the latent space of different VAE corpus encodings — Piano roll, MIDI, ABC, Tonnetz, DFT of pitch, and pitch class distributions — in providing a pitch space for key relations that align with cognitive distances. We evaluate the model performance of these encodings using objective metrics to capture accuracy, mean square error (MSE), KL- divergence, and computational cost. The ABC encoding performs the best in reconstructing the original data, while the Pitch DFT seems to capture more information from the latent space. Furthermore, an objective evaluation of 12 major or minor transpositions per piece is adopted to quantify the alignment of 1) intra- and inter-segment distances per key and 2) the key distances to cognitive pitch spaces. Our results show that Pitch DFT VAE latent spaces align best with cognitive spaces and provide a common-tone space where overlapping objects within a key are fuzzy clusters, which impose a well-defined order of structural significance or stability — i.e., a tonal hierarchy. Tonal hierarchies of different keys can be used to measure key distances and the relationships of their in-key components at multiple hierarchies (e.g., notes and chords). The implementation of our VAE and the encodings framework are made available online.
2023
Autores
Forero, J; Mendes, M; Bernardes, G;
Publicação
KUI
Abstract
This study explores the development of intelligent affective virtual environments generated by bimodal emotion recognition techniques and multimodal feedback. A semantic and acoustic analysis predicts emotions conveyed by spoken language, fostering an expressive and transparent control structure. Textual contents and emotional predictions are mapped to virtual environments in real locations as audiovisual feedback. To demonstrate the application of this system, we developed a case study titled "En train d'oublier,"focusing on a train cemetery in Uyuni, Bolivia. The train cemetery holds historical significance as a site where abandoned trains symbolize the passage of time and the interaction between human activities and nature's reclamation. The space is transformed into an immersive and emotionally poetic experience through oral language and affective virtual environments that activate memories, as the system utilizes the transcribed text to synthesize images and modifies the musical output based on the predicted emotional states. The proposed bimodal emotion recognition techniques achieve 94% and 89% accuracy. The audiovisual mapping strategy allows for considering divergence in predictions generating an intended tension between the graphical and the musical representation. Using video and web art techniques, we experimented with the environments generated to create diverses poetic proposals. © 2023 ACM.
2023
Autores
Carvalho, N; Diogo, D; Bernardes, G;
Publicação
THE 10TH INTERNATIONAL CONFERENCE ON DIGITAL LIBRARIES FOR MUSICOLOGY, DLFM 2023
Abstract
We propose a method for computing the similarity of symbolically-encoded Portuguese folk melodies. The main novelty of our method is the use of a preprocessing melodic reduction at multiple hierarchies to filter the surface of folk melodies according to 1) pitch stability, 2) interval salience, 3) beat strength, 4) durational accents, and 5) the linear combination of all former criteria. Based on the salience of each note event per criteria, we create three melodic reductions with three different levels of note retention. We assess the degree to which six folk music similarity measures at multiple reduction hierarchies comply with collected ground truth from experts in Portuguese folk music. The results show that SIAM combined with 75th quantile reduction using the combined or durational accents best models the similarity for a corpus of Portuguese folk melodies by capturing approximately 84-90% of the variance observed in ground truth annotations.
2023
Autores
Lopes, A; Barboza, JR; Bernardes, G;
Publicação
2023 Immersive and 3D Audio: from Architecture to Automotive, I3DA 2023
Abstract
Immersive audio technologies have broadened postproduction strategies for spatial audio, gaining popularity among mainstream audiences. However, there is a lack of defined procedures and critical thinking regarding audio mixing guidelines for surround sound in popular music. In this context, we conducted an empirical study to identify trends concerning instrument position, trajectories, and dynamics from surround mixings. Furthermore, we assess the degree to which they differ from their stereo renderings. Seven award-winning songs in the Grammy category for Best Immersive Album were analyzed, including surround 5.1 and stereo versions. The study found consistent instrument positions in the songs, with rhythmic instruments and bass in the center, lead vocals spread across front channels, and harmonic instruments in wider positions. Solo instruments occupied left, right, and center channels, with dynamics emphasizing lead vocals and solos. Trajectories were rarely used, indicating channel-based thinking. Limited adoption of immersive audio dimensions and reliance on stereo techniques were observed, with no notable differences between the surround and stereo versions. Identified song outliers are discussed and offer avenues for exploration, highlighting the importance of diverse musical expressions in informing immersive audio mixing. © 2023 IEEE.
The access to the final selection minute is only available to applicants.
Please check the confirmation e-mail of your application to obtain the access code.