IJCL

@ijcl.bsky.social

International Journal of Corpus Linguistics https://benjamins.com/catalog/ijcl

Jiří Milička, Anna Marklová & Václav Cvrček explore dimensions of variation in LLM-generated texts in Czech and English. They develop a reproducible method for measuring “register shift”, demonstrating the efficacy of different LLMs for producing linguistically diverse texts doi.org/10.1075/ijcl...

doi.org

OUT NOW: Qiao Gan offers insights into the variable realisation of there+BE+plural arguments Using Australian English language data collected in the 1970s and 2010s, Gan examines intersections of age, ethnicity and class alongside grammatical conditioning in use of THERE's doi.org/10.1075/ijcl...

doi.org

How reliable are dispersion measures? OUT NOW: Testing seven popular measures for sensitivity, Lukas Sönning & Jesse Egbert discuss the distributional and corpus design features that affect measures of the pervasiveness and evenness of text features in corpus analysis. doi.org/10.1075/ijcl...

Sensitivity of dispersion measures to distributional patterns and corpus design

Abstract Recent work has shown that dispersion measures respond to multiple features in the data: Juilland’s D varies systematically with the number of corpus parts, and all commonly used indices are ...

doi.org

OUT NOW: Felix Morger & Aleksandrs Berdicevskis compare how well variation can be predicted using logistic-regression models and BERT. Their example case studies include the English dative alternation and Swedish att-omission, to consider differences in predictability doi.org/10.1075/ijcl...

Not all linguistic variation is equally predictable

Abstract We compare to what extent the choice of a variant can be predicted from language-internal factors for two linguistic variables: the English dative alternation and the omission of the infiniti...

doi.org

OUT NOW: Liina Repo, Brett Hashimoto & Veronika Laippala look at extending register annotations from a corpus of historical documents using BERT-based deep learning models. Considering text-internal variation, they find that beginnings tend to support better predictions. doi.org/10.1075/ijcl...

doi.org

OUT NOW: @timfeld.bsky.social, Fabian Barteld & Alexander Ziem present and evaluate an approach for predicting typical fillers for the slots of grammatical constructions using BERT. They ask: How can language models be used to support the development of linguistic resources? doi.org/10.1075/ijcl...

Can BERT predict fillers for construction elements?

Abstract The starting point of this paper is a central problem in constructicography: on the one hand, constructicon projects aim at describing a broad spectrum of constructions based on usage data; o...

doi.org

OUT NOW: Yiğit Savuran & Stefanie Wulff introduce the Turkish Learner Corpus (TURLEC), comprising written and spoken texts of learners of Turkish L2 across CEFR proficiency levels. The corpus and metadata are available at: osf.io/bnv3p/files/... #OpenScience #LearnerCorpus doi.org/10.1075/ijcl...

Introducing TURLEC

Abstract This paper provides a detailed account of the Turkish Learner Corpus (TURLEC). Building on the first author’s doctoral dissertation project, which aimed to identify proficiency descriptors fo...

doi.org

OUT NOW: the next paper to appear in our Special Issue on Corpus Perspectives on #LegalDiscourse Le Cheng, Xiuli Liu & Jian Li present a continuum of stance as derived from a cross-genre examination of stance expressions in legislation, judgments and legal academic articles doi.org/10.1075/ijcl...

Continuum of stance in law

Abstract Stance is deep-rooted in law, where legal values can never stand in a vacuum. Despite a growing body of literature on stance in legal genres, cross-genre examinations conducted from a corpus-...

doi.org

OUT NOW: Davide Mazzi explores argument structures in a corpus of Supreme Court of Ireland’s judgments on human rights, based on indicators of pragmatic argumentation. This paper appears as part of our Special Issue on Corpus Perspectives on #LegalDiscourse doi.org/10.1075/ijcl...

“…animated by a number of fundamental principles”

Abstract The aim of this paper is to combine a quantitative analysis of indicators of pragmatic argumentation with a qualitative investigation of the argument scheme in a corpus of Supreme Court of Ir...

doi.org

OUT NOW – the next contribution to our forthcoming Special Issue on Corpus Perspectives on Legal Discourse: Edward Clay presents a systematic approach for identifying indicators of divergence, comparing terms relating to migration in EU legal documents and news articles doi.org/10.1075/ijcl...

doi.org

OUT NOW: Biel, Wasilewska & Koźbiał explore linguistic variation in the Polish Eurolect, applying MDA to a corpus of legal acts, judgments, administrative reports, and institutional websites. This paper will appear as part of our Special Issue on #ForensicLinguistics doi.org/10.1075/ijcl...

Dimensions of variation across institutional legal and administrative registers

Abstract This study applies full Multidimensional Analysis (MDA) to examine linguistic variation in the Polish Eurolect — a hybrid variety shaped by translation and institutional constraints within th...

doi.org

OUT NOW: Jia Li and Xianyao Hu compare human- and machine-translated texts from Chinese to English to evaluate features of conservatism across registers Their investigation sheds light on the potetnial of human-machine collaborative translation models doi.org/10.1075/ijcl...

Is human translation more conservative than machine translation? | John Benjamins

Abstract The present study investigates whether conservatism exists in human- and machine-translated texts from Chinese into English, and whether this tendency is consistently observable across differ...

doi.org

OUT NOW: Rose Stamp provides a comprehensive review of the current state of sign language corpora around the world – discussing video capture, transcription, and coding and how these relate to corpus compilation in terms of representativeness, searchability and open access doi.org/10.1075/ijcl...

Sign language corpora designed for sociolinguistic research | John Benjamins

Abstract Sign language corpora are generally under-represented in the field of corpus linguistics. Fortunately, in the last twenty years there has been a steady rise in their creation, following techn...

doi.org

la semaine prochaine but la raison suivante Looi, Riget, Boulton & Hassan discuss synonym alternation between French prochain and suivant, using corpus evidence and statistical methods to re-examine variables derived through introspection #OnlineFirst #FrenchLinguistics doi.org/10.1075/ijcl...

From theory to data | John Benjamins

Abstract This paper presents a corpus-based study that evaluates variables identified introspectively by Berthonneau (2002) in relation to the alternation between two French synonymous: prochain (‘nex...

doi.org