Leivada, Dentella & Günther (2024). Evaluating the language abilities of humans vs. Large Language Models: Three caveats

Autors:

Evelina Leivada, Vittoria Dentella & Fritz Günther

Títol:

Biolinguistics, vol.18

Editorial: PsychOpen
Data de publicació: 19 abril, 2024

Text complet

We identify and analyze three caveats that may arise when analyzing the linguistic abilities of Large Language Models. The problem of unlicensed generalizations refers to the danger of interpreting performance in one task as predictive of the models’ overall capabilities, based on the assumption that because a specific task performance is indicative of certain underlying capabilities in humans, the same association holds for models. The human-like paradox refers to the problem of lacking human comparisons, while at the same time attributing human-like abilities to the models. Last, the problem of double standards refers to the use of tasks and methodologies that either cannot be applied to humans or they are evaluated differently in models vs. humans. While we recognize the impressive linguistic abilities of LLMs, we conclude that specific claims about the models’ human-likeness in the grammatical domain are premature.

Masullo, Casado, Leivada & Sorace (2025). Register variation and linguistic background modulate accuracy in detecting morphosyntactic errors

Autors:

Masullo, Casado, Leivada & Sorace

Títol:

Register variation and linguistic background modulate accuracy in detecting morphosyntactic errors

Editorial: Isogloss. Open Journal of Romance Linguistics
Data de publicació: 30-03-2025
Pàgines: 36

Més informació
Text complet


Linguistic register is defined as a variety of language shaped by different situational settings. Adapting to register is crucial for successful communication and involves the processing of language features related to register variation. Few studies have focused on the impact of linguistic register on language processing. Our research investigates whether register variation affects the detection of linguistic errors. To determine if linguistic background further impacts the way we deal with register, our sample includes monolingual, bilingual, and bidialectal participants. All groups completed an acceptability judgement task in Italian that features Subject-Verb agreement mismatches presented in high and low register. The results reveal a significant impact of linguistic register on accuracy: morphosyntactic errors are better detected in low-register stimuli. Furthermore, different trends characterize the tested groups. While monolinguals show more similar accuracy rates for low- and high-register sentences, the bilingual groups tend to better spot errors in low-register stimuli. Our findings suggest that register plays an important role in the processing of morphosyntactic errors, highlighting the need to consider both its cognitive and social dimensions. Moreover, the variation observed among the tested groups underscores that language processing can be influenced by factors related to the sociolinguistic dimensions of each linguistic community.

Leivada, Marcus, Günther & Murphy (2025). A Sentence is Worth a Thousand Pictures: Can Large Language Models Understand Hum4n L4ngu4ge and the W0rld behind W0rds?

Autors:

Leivada, Marcus, Günther & Murphy

Títol:

A Sentence is Worth a Thousand Pictures: Can Large Language Models Understand Hum4n L4ngu4ge and the W0rld behind W0rds?

Editorial: Philosophical Transactions of the Royal Society A
Data de publicació: 2025
Pàgines: 19

Més informació
Text complet


Modern Artificial Intelligence applications show great potential for language- related tasks that rely on next-word prediction. The current generation of Large Language Models (LLMs) have been linked to claims about human-like linguistic performance and their applications are hailed both as a step towards artificial general intelligence and as a major advance in understanding the cognitive, and even neural basis of human language. To assess these claims, first we analyze the contribution of LLMs as theoretically informative representations of a target cognitive system vs. atheoretical mechanistic tools. Second, we evaluate the models’ ability to see the bigger picture, through top-down feedback from higher levels of processing, which requires grounding in previous expectations and past world experience. We hypothesize that since models lack grounded cognition, they cannot take advantage of these features and instead solely rely on fixed associations between represented words and word vectors. To assess this, we designed and ran a novel ‘leet task’ (l33t t4sk), which requires decoding sentences in which letters are systematically replaced by numbers. The results suggest that humans excel in this task whereas models struggle, confirming our hypothesis. We interpret the results by identifying the key abilities that are still missing from the current state of development of these models, which require solutions that go beyond increased system scaling.

Dentella, Günther & Leivada (2025). Language in vivo vs. in silico: Size matters but Larger Language Models still do not comprehend language on a par with humans due to impenetrable semantic reference

Autors:

Dentella, Günther & Leivada

Títol:

Language in vivo vs. in silico: Size matters but Larger Language Models still do not comprehend language on a par with humans due to impenetrable semantic reference

Editorial: PLoS ONE
Data de publicació: 17-07-2025

Més informació
Text complet


Understanding the limits of language is a prerequisite for Large Language Models (LLMs) to act as theories of natural language. LLM performance in some language tasks presents both quantitative and qualitative differences from that of humans, however it remains to be determined whether such differences are amenable to model size. This work investigates the critical role of model scaling, determining whether increases in size make up for such differences between humans and models. We test three LLMs from different families (Bard, 137 billion parameters; ChatGPT-3.5, 175 billion; ChatGPT-4, 1.5 trillion) on a grammaticality judgment task featuring anaphora, center embedding, comparatives, and negative polarity. N = 1,200 judgments are collected and scored for accuracy, stability, and improvements in accuracy upon repeated presentation of a prompt. Results of the best performing LLM, ChatGPT-4, are compared to results of n = 80 humans on the same stimuli. We find that humans are overall less accurate than ChatGPT-4 (76% vs. 80% accuracy, respectively), but that this is due to ChatGPT-4 outperforming humans only in one task condition, namely on grammatical sentences. Additionally, ChatGPT-4 wavers more than humans in its answers (12.5% vs. 9.6% likelihood of an oscillating answer, respectively). Thus, while increased model size may lead to better performance, LLMs are still not sensitive to (un)grammaticality the same way as humans are. It seems possible but unlikely that scaling alone can fix this issue. We interpret these results by comparing language learning in vivo and in silico, identifying three critical differences concerning (i) the type of evidence, (ii) the poverty of the stimulus, and (iii) the occurrence of semantic hallucinations due to impenetrable linguistic reference.

Pantelidou, N., Leivada E., Montero, R. & Morosi, P. 2026. Community size rather than grammatical complexity better predicts Large Language Model accuracy in a novel Wug Test.

Autors:

Pantelidou, Nikoleta, Evelina Leivada, Raquel Montero, Paolo Morosi

Títol:

Community size rather than grammatical complexity better predicts Large Language Model accuracy in a novel Wug Test

Editorial: PLoS One 21(3)
Col·lecció:
Data de publicació: 2026

Text complet

Abstract

The linguistic abilities of Large Language Models are a matter of ongoing debate. This study contributes to this discussion by investigating model performance in a morphological generalization task that involves novel words. Using a multilingual adaptation of the Wug Test, six models were tested across four partially unrelated languages (Catalan, English, Greek, and Spanish) and compared with human speakers. The aim is to determine whether model accuracy approximates human competence and whether it is shaped primarily by linguistic complexity or by the size of the linguistic community, which affects the quantity of available training data. Consistent with previous research, the results show that the models are able to generalize morphological processes to unseen words with human-like accuracy. However, accuracy patterns align more closely with community size and data availability than with structural complexity, refining earlier claims in the literature. In particular, languages with larger speaker communities and stronger digital representation, such as Spanish and English, revealed higher accuracy than less-resourced ones like Catalan and Greek. Overall, our findings suggest that model behavior is mainly driven by the richness of linguistic resources rather than by sensitivity to grammatical complexity, reflecting a form of performance that resembles human linguistic competence only superficially.

Citation: Pantelidou N, Leivada E, Montero R, Morosi P (2026) Community size rather than grammatical complexity better predicts Large Language Model accuracy in a novel Wug Test. PLoS One 21(3): e0343164. https://doi.org/10.1371/journal.pone.0343164

Editor: Wei Lun Wong, National University of Malaysia Faculty of Education: Universiti Kebangsaan Malaysia Fakulti Pendidikan, MALAYSIA

Received: October 16, 2025; Accepted: February 2, 2026; Published: March 11, 2026

Copyright: © 2026 Pantelidou et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Data Availability: All data files are available from the OSF database (https://osf.io/4z5n6/).

Funding: EL acknowledges funding from the Spanish Ministry of Science, Innovation & Universities MCIN/AEI/https://doi.org/10.13039/501100011033) under the research project CNS2023-144415. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

Competing interests: The authors have declared that no competing interests exist.