Bruno Bianchi y Fermín Travi, integrantes del LIAA, publicaron dos nuevos papers en Proceedings of the 15th Workshop of Cognitive Modeling and Computational Linguistics (CMCL)
“Disentanglement and Compositionality of Letter Identity and Letter Position in Variational Auto-Encoder Vision Models”
Bruno Bianchi, Aakash Agrawal, Stanislas Dehaene, Emmanuel Chemla, Yair Lakretz
Abstract
Human readers can accurately count how many letters are in a word (e.g., 7 in “buffalo”), remove a letter from a given position (e.g., “bufflo”) or add a new one. The human brain of readers must have therefore learned to disentangle information related to the position of a letter and its identity. Such sentanglement is necessary for the compositional, unbounded ability of humans to create and parse new strings, with any combination of letters appearing in any positions. Do modern deep neural models also possess this crucial compositional ability? Here, we tested whether neural models that achieve state-of-the-art on disentanglement of features in visual input can also disentangle letter position and letter identity when trained on images of written words. Specifically, we trained beta variational autoencoder (β-VAE) to reconstruct images of letter strings and evaluated their disentanglement performance using CompOrth – a new benchmark that we created for studying compositional learning and zero-shot generalization in visual models for orthography. The benchmark suggests a set of tests, of increasing complexity, to evaluate the degree of disentanglement between orthographic features of written words in deep neural models. Using CompOrth, we conducted a set of experiments to analyze the generalization ability of these models, in particular, to unseen word length and to unseen combinations of letter identities and letter positions. We found that while models effectively disentangle surface features, such as horizontal and vertical ‘retinal’ locations of words within an image, they dramatically fail to disentangle letter position and letter identity and lack any notion of word length. Together, this study demonstrates the shortcomings of state-of-the-art β-VAE models compared to humans and proposes a new challenge and a corresponding benchmark to evaluate neural models.
Link al paper: https://doi.org/10.63317/3btikw97586i
“Topic-context Dependency on Continuous Semantic Reconstruction of Language from fMRI Signals”
Fermín Travi, Agustín Delmagro, Diego Fernández Slezak, Bruno Bianchi, Juan E Kamienkowski
Abstract
Recent advances in neural language decoding have enabled reconstruction of perceived speech from MRI signals using large language models and brain encoders. While these systems achieve impressive semantic fidelity, the extent to which training data biases signal encoding and the resulting decoded output remains unclear. Here, we investigate how topic-specific training constrains language decoding by partitioning 84 auditory stories into two semantic clusters and training brain encoders on topic-specific subsets across three participants. We find that encoders trained exclusively on one topic systematically bias decoded outputs toward their training domain: family-related stories are reclassified as such when reconstructed with family-trained encoders, but not otherwise. Moreover, topic-constrained encoders perform significantly above chance only within their training topic and fall below random baseline when decoding out-of-topic stimuli. These findings underscore the need for more diverse and semantically rich training data, while also raising questions about the breadth of semantic variability that current brain encoders can effectively capture.
Link al paper: https://doi.org/10.63317/233sezzfxz2m