Pantelidou, N., Leivada E., Montero, R. & Morosi, P. 2026. Community size rather than grammatical complexity better predicts Large Language Model accuracy in a novel Wug Test.

Autors:

Pantelidou, Nikoleta, Evelina Leivada, Raquel Montero, Paolo Morosi

Títol:

Community size rather than grammatical complexity better predicts Large Language Model accuracy in a novel Wug Test

Editorial: PLoS One 21(3)
Col·lecció:
Data de publicació: 2026

Text complet

Abstract

The linguistic abilities of Large Language Models are a matter of ongoing debate. This study contributes to this discussion by investigating model performance in a morphological generalization task that involves novel words. Using a multilingual adaptation of the Wug Test, six models were tested across four partially unrelated languages (Catalan, English, Greek, and Spanish) and compared with human speakers. The aim is to determine whether model accuracy approximates human competence and whether it is shaped primarily by linguistic complexity or by the size of the linguistic community, which affects the quantity of available training data. Consistent with previous research, the results show that the models are able to generalize morphological processes to unseen words with human-like accuracy. However, accuracy patterns align more closely with community size and data availability than with structural complexity, refining earlier claims in the literature. In particular, languages with larger speaker communities and stronger digital representation, such as Spanish and English, revealed higher accuracy than less-resourced ones like Catalan and Greek. Overall, our findings suggest that model behavior is mainly driven by the richness of linguistic resources rather than by sensitivity to grammatical complexity, reflecting a form of performance that resembles human linguistic competence only superficially.

Citation: Pantelidou N, Leivada E, Montero R, Morosi P (2026) Community size rather than grammatical complexity better predicts Large Language Model accuracy in a novel Wug Test. PLoS One 21(3): e0343164. https://doi.org/10.1371/journal.pone.0343164

Editor: Wei Lun Wong, National University of Malaysia Faculty of Education: Universiti Kebangsaan Malaysia Fakulti Pendidikan, MALAYSIA

Received: October 16, 2025; Accepted: February 2, 2026; Published: March 11, 2026

Copyright: © 2026 Pantelidou et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Data Availability: All data files are available from the OSF database (https://osf.io/4z5n6/).

Funding: EL acknowledges funding from the Spanish Ministry of Science, Innovation & Universities MCIN/AEI/https://doi.org/10.13039/501100011033) under the research project CNS2023-144415. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

Competing interests: The authors have declared that no competing interests exist.

Jardón, Marx & Wittenberg (2025). Is there a perfect state? Experimental evidence from English and Spanish for the perfect-as-state hypothesis

Autors:

Natalia Jardón Pérez, Elena Marx & Eva Wittenberg

Títol:

Is there a perfect state? Experimental evidence from English and Spanish for the perfect-as-state hypothesis

Editorial: Glossa: a journal of general linguistics
Data de publicació: 24 d'octubre de 2025

Més informació

The distinction between past and perfect has been subject to theoretical debate for decades. Some of the most prominent accounts argue that the perfect turns the mental representation of a past event into the mental representation of a state, based on a past event: the perfect-as-state hypothesis. This subtle distinction is notoriously difficult to trace, and has only been argued for using linguistic tests. Here, we provide evidence from two psycholinguistic experiments (total N=960), each in English and in Spanish, that operationalize stativity through event individuation. Our results show that compared to the past, the perfect leads to event construals that have more in common with states, both in English and in Spanish. These data constitute the first documentation of different event construals based on tenses that only differ in the subtlest of semantic distinctions.

Repiso-Puigdelliura (2025). Effects of language dominance in Catalan-Spanish-English trilinguals’ vowel-initial glottal marking: A Principal Components Analysis approach

Autors:

Gemma Repiso-Puigdelliura

Títol:

Effects of language dominance in Catalan-Spanish-English trilinguals' vowel-initial glottal marking: A Principal Components Analysis approach

Editorial: Glossa: a journal of general linguistics
Data de publicació: 8 de setembres de 2025

Més informació

Crosslinguistic influence (i.e., CLI, henceforth) in trilingual speakers is multidirectional and shaped by factors such as the amount of exposure to, and use of, each of the speaker’s languages. This study investigates whether relative dominance explains progressive and regressive CLI in trilingual speakers. To this purpose, we examine the production of word-external vocalic sequences (i.e., /V#V/) in L3 English speakers who are Catalan-Spanish bilinguals. Participants completed a reading task in English, Spanish, and Catalan that elicited vowel-to-vowel sequences along four levels of stress (i.e., stressed-stressed, unstressed-stressed, stressed-unstressed, unstressed-unstressed). Alongside the production task, they filled out a trilingual version of the Bilingual Language Profile (i.e., BLP, henceforth) (Birdsong et al. 2012). The resulting vocalic sequences were classified as instances of glottal marking (i.e., creaky phonation or complete glottal stop) or modal phonation. To examine the role of dominance, we ran a Principal Component Analysis on the questionnaire data, identifying four principal components that explained 57.9% of the variance. We compared L3 English vowel-to-vowel sequences with those of Spanish-English bilinguals who speak Catalan as an L3, as well as with L1 English monolinguals. We ran dominance-based logistic regressions for each language. In English, our results show that L3 English speakers differ from their L3 Catalan Spanish-English bilingual counterparts in unstressedstressed vowel sequences, but differ across all four stress levels when compared to L1 English monolinguals. Dominance-related principal components do not predict the rate of glottal marking in L3 English. In L1 Catalan and L1 Spanish, the use of glottalization is predicted by the average rate of glottal marking in the speakers’ L3 English productions, as well as by higher scores on the principal component associated with L3 English dominance. In Spanish, vowel-initial glottal marking is predicted by scores associated with low Spanish dominance. These findings highlight that dominance mediates CLI in trilingual speakers, which in turn reflects the dynamic nature of CLI in multilingual speakers.

Russo Cardona & Villalba (2025). The interaction between clause size and Voice: Evidence from Catalan and Italian

Autors:

Russo Cardona, L. & Villalba, X.

Títol:

The interaction between clause size and Voice: Evidence from Catalan and Italian

Editorial: Glossa: a journal of general linguistics 10(1)
Data de publicació: 18 de setembre de 2025

Més informació
Text complet


We argue that in certain reduced embedded clauses Voice behaves differently from most other contexts, on the basis of tough-constructions (TCs) and modal passives (MPs) in Catalan and Italian. These constructions involve an A-dependency targeting only internal arguments of morphologically active transitive infinitives (unlike control, raising, and restructuring dependencies) because they involve a C/I-less VoiceP complement with a defective Voice layer (no accusative, no passive morphology, passive-like implicit agent). Thanks to the existence of a resumptive variant of TCs/MPs in Catalan, we propose a way to derive the distribution of defective Voice, which must be directly selected by a suitable lexical category, with regard to active/passive Voice, which must be directly selected by a functional head (at least in the languages at issue). Our findings bear on the broader theoretical debates about the typologies of Voice, clausal complements, and on the syntactic correlates of clause size.

Recasens (2025). The diachronic evolution of syllable-onset /Cl/ clusters in Romance revisited. An integrated account

Autors:

Daniel Recasens

Títol:

The diachronic evolution of syllable-onset /Cl/ clusters in Romance revisited. An integrated account

Editorial: Diachronica
Data de publicació: 23 de setembre de 2025
Pàgines: 53

Més informació
Text complet


This paper deals with the historical development of the syllable-onset clusters /kl gl pl bl fl/ in Romance languages and dialects and with their articulatory and/or perceptual motivations. Several diachronic pathways are identified which depart from articulatory unstable [Cʎ] sequences, most distinctively lateral vocalization (e.g., [kʎ] > [kj]) and obstruent lenition (e.g., [kʎ] > [çʎ]). Most sound changes are attributed to articulatory variation insofar as they require adjustments in constriction degree and location. In a few cases the replacement of one consonantal sound by another appears to have been induced by acoustic-perceptual equivalence, as for example the substitution of [θ] by [f] in Franco-Provençal and of palatalized labial stops by palatal stops in southern Italy. Of special interest is the one-to-many derivation problem by which a given phonetic outcome may be achieved through more than one pathway as exemplified by the two phonetic developments /kl/ > [kʎ] > [kj] > [c] > [tʃ] > [ʃ] and /kl/ > [kʎ] > [çʎ] > [çj] > [ʃj] > [ʃ].