18 juliol, 2025

Autors:
Leivada, Marcus, Günther & Murphy
Títol:
A Sentence is Worth a Thousand Pictures: Can Large Language Models Understand Hum4n L4ngu4ge and the W0rld behind W0rds?Editorial: Philosophical Transactions of the Royal Society A
Data de publicació: 2025
Pàgines: 19 Més informació
Text completModern Artificial Intelligence applications show great potential for language- related tasks that rely on next-word prediction. The current generation of Large Language Models (LLMs) have been linked to claims about human-like linguistic performance and their applications are hailed both as a step towards artificial general intelligence and as a major advance in understanding the cognitive, and even neural basis of human language. To assess these claims, first we analyze the contribution of LLMs as theoretically informative representations of a target cognitive system vs. atheoretical mechanistic tools. Second, we evaluate the models’ ability to see the bigger picture, through top-down feedback from higher levels of processing, which requires grounding in previous expectations and past world experience. We hypothesize that since models lack grounded cognition, they cannot take advantage of these features and instead solely rely on fixed associations between represented words and word vectors. To assess this, we designed and ran a novel ‘leet task’ (l33t t4sk), which requires decoding sentences in which letters are systematically replaced by numbers. The results suggest that humans excel in this task whereas models struggle, confirming our hypothesis. We interpret the results by identifying the key abilities that are still missing from the current state of development of these models, which require solutions that go beyond increased system scaling.
18 juliol, 2025

Autors:
Dentella, Günther & Leivada
Títol:
Language in vivo vs. in silico: Size matters but Larger Language Models still do not comprehend language on a par with humans due to impenetrable semantic referenceEditorial: PLoS ONE
Data de publicació: 17-07-2025
Més informació
Text completUnderstanding the limits of language is a prerequisite for Large Language Models (LLMs) to act as theories of natural language. LLM performance in some language tasks presents both quantitative and qualitative differences from that of humans, however it remains to be determined whether such differences are amenable to model size. This work investigates the critical role of model scaling, determining whether increases in size make up for such differences between humans and models. We test three LLMs from different families (Bard, 137 billion parameters; ChatGPT-3.5, 175 billion; ChatGPT-4, 1.5 trillion) on a grammaticality judgment task featuring anaphora, center embedding, comparatives, and negative polarity. N = 1,200 judgments are collected and scored for accuracy, stability, and improvements in accuracy upon repeated presentation of a prompt. Results of the best performing LLM, ChatGPT-4, are compared to results of n = 80 humans on the same stimuli. We find that humans are overall less accurate than ChatGPT-4 (76% vs. 80% accuracy, respectively), but that this is due to ChatGPT-4 outperforming humans only in one task condition, namely on grammatical sentences. Additionally, ChatGPT-4 wavers more than humans in its answers (12.5% vs. 9.6% likelihood of an oscillating answer, respectively). Thus, while increased model size may lead to better performance, LLMs are still not sensitive to (un)grammaticality the same way as humans are. It seems possible but unlikely that scaling alone can fix this issue. We interpret these results by comparing language learning in vivo and in silico, identifying three critical differences concerning (i) the type of evidence, (ii) the poverty of the stimulus, and (iii) the occurrence of semantic hallucinations due to impenetrable linguistic reference.
18 març, 2026

Autors:
Pantelidou, Nikoleta, Evelina Leivada, Raquel Montero, Paolo Morosi
Títol:
Community size rather than grammatical complexity better predicts Large Language Model accuracy in a novel Wug TestEditorial: PLoS One 21(3)
Col·lecció: PLoS OneData de publicació: 2026
Text complet
Abstract
The linguistic abilities of Large Language Models are a matter of ongoing debate. This study contributes to this discussion by investigating model performance in a morphological generalization task that involves novel words. Using a multilingual adaptation of the Wug Test, six models were tested across four partially unrelated languages (Catalan, English, Greek, and Spanish) and compared with human speakers. The aim is to determine whether model accuracy approximates human competence and whether it is shaped primarily by linguistic complexity or by the size of the linguistic community, which affects the quantity of available training data. Consistent with previous research, the results show that the models are able to generalize morphological processes to unseen words with human-like accuracy. However, accuracy patterns align more closely with community size and data availability than with structural complexity, refining earlier claims in the literature. In particular, languages with larger speaker communities and stronger digital representation, such as Spanish and English, revealed higher accuracy than less-resourced ones like Catalan and Greek. Overall, our findings suggest that model behavior is mainly driven by the richness of linguistic resources rather than by sensitivity to grammatical complexity, reflecting a form of performance that resembles human linguistic competence only superficially.
Citation: Pantelidou N, Leivada E, Montero R, Morosi P (2026) Community size rather than grammatical complexity better predicts Large Language Model accuracy in a novel Wug Test. PLoS One 21(3): e0343164. https://doi.org/10.1371/journal.pone.0343164
Editor: Wei Lun Wong, National University of Malaysia Faculty of Education: Universiti Kebangsaan Malaysia Fakulti Pendidikan, MALAYSIA
Received: October 16, 2025; Accepted: February 2, 2026; Published: March 11, 2026
Copyright: © 2026 Pantelidou et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: All data files are available from the OSF database (https://osf.io/4z5n6/).
Funding: EL acknowledges funding from the Spanish Ministry of Science, Innovation & Universities MCIN/AEI/https://doi.org/10.13039/501100011033) under the research project CNS2023-144415. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Competing interests: The authors have declared that no competing interests exist.
30 abril, 2025

Autors:
Acedo-Matellán & Real-Puigdollers
Títol:
Boundedness in locative prepositions: Evidence from CatalanEditorial: Natural Language & Linguistic Theory
Data de publicació: 02-01-2025
Pàgines: 39 Més informació
Text completThis paper provides evidence from Catalan for the existence of bounded and unbounded locative prepositions, and proposes that boundedness in the adpositional domain is derived similarly to boundedness in the verbal, nominal or adjectival domains. Our contribution is both empirical and theoretical. First, we show that Catalan has two simple locative prepositions, a and en, which form a minimal pair as far as boundedness is concerned and exhibit, correspondingly, different selection patterns: while bounded a only selects DPs with a quantity interpretation, unbounded en can combine with both NPs and DPs, which receive a homogeneous interpretation. Second, we develop a syntactic and semantic theory to account for these facts that relates them to the crosscategorial property of boundedness: a-PPs, but not en-PPs, contain an aspectual projection that imposes the interpretation that the otherwise homogeneous region denoted by the preposition is delimited. Moreover, we show that the difference between the structures licensed by a and en has consequences for the interpretation of quantifiers within PPs. Specifically, we set eyes upon a particular context in which a and en take a universally quantified singular DP as complement and form a minimal pair. We propose that while the bounded preposition a allows for the interpretation of the quantifier tot ‘all’ as a universal quantifier of parts, the unbounded preposition en does not. Instead, with en the quantifier behaves as an adjective of sorts associated to a maximality operator. Our paper contributes to furthering our understanding of boundedness across categories in human language.
17 setembre, 2020

Autors:
Javier Fernández Sánchez & Dennis Ott
Títol:
DislocationsEditorial: Language and Linguistics Compass, Vol.14 issue 9 (John Wiley & Sons Ltd)
Data de publicació: Setembre 2020
Text completDislocation is a kind of construction in which a phrasal constituent (the dislocate) appears at the outer left or right edge of a gap-less clause (its host) that contains a pronominal correlate of the dislocate. Dislocations are widely attested and presumably universally available across languages. The construction raises a number of problems for core assumptions of syntactic theory, in that these assumptions appear to thwart any coherent resolution of the question of how the dislocate relates to the internal structure of its host. This contribution is divided into two parts. In Part 1, we review central empirical properties of dislocation, which, taken together, appear to defy the laws of syntax as commonly assumed. In Part 2, we review key proposals that have emerged over the last decennia to resolve this paradox and restore dislocations to normalcy.