IJCL

@ijcl.bsky.social

International Journal of Corpus Linguistics https://benjamins.com/catalog/ijcl

How reliable are dispersion measures? OUT NOW: Testing seven popular measures for sensitivity, Lukas Sönning & Jesse Egbert discuss the distributional and corpus design features that affect measures of the pervasiveness and evenness of text features in corpus analysis. doi.org/10.1075/ijcl...

Sensitivity of dispersion measures to distributional patterns and corpus design

Abstract Recent work has shown that dispersion measures respond to multiple features in the data: Juilland’s D varies systematically with the number of corpus parts, and all commonly used indices are ...

doi.org

OUT NOW: Felix Morger & Aleksandrs Berdicevskis compare how well variation can be predicted using logistic-regression models and BERT. Their example case studies include the English dative alternation and Swedish att-omission, to consider differences in predictability doi.org/10.1075/ijcl...

Not all linguistic variation is equally predictable

Abstract We compare to what extent the choice of a variant can be predicted from language-internal factors for two linguistic variables: the English dative alternation and the omission of the infiniti...

doi.org

OUT NOW: Liina Repo, Brett Hashimoto & Veronika Laippala look at extending register annotations from a corpus of historical documents using BERT-based deep learning models. Considering text-internal variation, they find that beginnings tend to support better predictions. doi.org/10.1075/ijcl...

doi.org

OUT NOW: @timfeld.bsky.social, Fabian Barteld & Alexander Ziem present and evaluate an approach for predicting typical fillers for the slots of grammatical constructions using BERT. They ask: How can language models be used to support the development of linguistic resources? doi.org/10.1075/ijcl...

Can BERT predict fillers for construction elements?

Abstract The starting point of this paper is a central problem in constructicography: on the one hand, constructicon projects aim at describing a broad spectrum of constructions based on usage data; o...

doi.org

OUT NOW: Yiğit Savuran & Stefanie Wulff introduce the Turkish Learner Corpus (TURLEC), comprising written and spoken texts of learners of Turkish L2 across CEFR proficiency levels. The corpus and metadata are available at: osf.io/bnv3p/files/... #OpenScience #LearnerCorpus doi.org/10.1075/ijcl...

Introducing TURLEC

Abstract This paper provides a detailed account of the Turkish Learner Corpus (TURLEC). Building on the first author’s doctoral dissertation project, which aimed to identify proficiency descriptors fo...

doi.org

OUT NOW: the next paper to appear in our Special Issue on Corpus Perspectives on #LegalDiscourse Le Cheng, Xiuli Liu & Jian Li present a continuum of stance as derived from a cross-genre examination of stance expressions in legislation, judgments and legal academic articles doi.org/10.1075/ijcl...

Continuum of stance in law

Abstract Stance is deep-rooted in law, where legal values can never stand in a vacuum. Despite a growing body of literature on stance in legal genres, cross-genre examinations conducted from a corpus-...

doi.org

OUT NOW: Davide Mazzi explores argument structures in a corpus of Supreme Court of Ireland’s judgments on human rights, based on indicators of pragmatic argumentation. This paper appears as part of our Special Issue on Corpus Perspectives on #LegalDiscourse doi.org/10.1075/ijcl...

“…animated by a number of fundamental principles”

Abstract The aim of this paper is to combine a quantitative analysis of indicators of pragmatic argumentation with a qualitative investigation of the argument scheme in a corpus of Supreme Court of Ir...

doi.org

OUT NOW – the next contribution to our forthcoming Special Issue on Corpus Perspectives on Legal Discourse: Edward Clay presents a systematic approach for identifying indicators of divergence, comparing terms relating to migration in EU legal documents and news articles doi.org/10.1075/ijcl...

doi.org

OUT NOW: Biel, Wasilewska & Koźbiał explore linguistic variation in the Polish Eurolect, applying MDA to a corpus of legal acts, judgments, administrative reports, and institutional websites. This paper will appear as part of our Special Issue on #ForensicLinguistics doi.org/10.1075/ijcl...

Dimensions of variation across institutional legal and administrative registers

Abstract This study applies full Multidimensional Analysis (MDA) to examine linguistic variation in the Polish Eurolect — a hybrid variety shaped by translation and institutional constraints within th...

doi.org

OUT NOW: Jia Li and Xianyao Hu compare human- and machine-translated texts from Chinese to English to evaluate features of conservatism across registers Their investigation sheds light on the potetnial of human-machine collaborative translation models doi.org/10.1075/ijcl...

Is human translation more conservative than machine translation? | John Benjamins

Abstract The present study investigates whether conservatism exists in human- and machine-translated texts from Chinese into English, and whether this tendency is consistently observable across differ...

doi.org

OUT NOW: Rose Stamp provides a comprehensive review of the current state of sign language corpora around the world – discussing video capture, transcription, and coding and how these relate to corpus compilation in terms of representativeness, searchability and open access doi.org/10.1075/ijcl...

Sign language corpora designed for sociolinguistic research | John Benjamins

Abstract Sign language corpora are generally under-represented in the field of corpus linguistics. Fortunately, in the last twenty years there has been a steady rise in their creation, following techn...

doi.org

la semaine prochaine but la raison suivante Looi, Riget, Boulton & Hassan discuss synonym alternation between French prochain and suivant, using corpus evidence and statistical methods to re-examine variables derived through introspection #OnlineFirst #FrenchLinguistics doi.org/10.1075/ijcl...

From theory to data | John Benjamins

Abstract This paper presents a corpus-based study that evaluates variables identified introspectively by Berthonneau (2002) in relation to the alternation between two French synonymous: prochain (‘nex...

doi.org

spooktacular, momfluencer, pupperazi.. How do combining forms operate and what meaning is transferred? Jinhong Huang and Yongwei Gao examine the evidence for 10 combining forms in American English to map out their schematic extensions and stability #WordFormation doi.org/10.1075/ijcl...

A corpus-based study into new combining forms in American English | John Benjamins

Abstract This study examines 10 new combining forms (CFs) in American English from both diachronic and synchronic perspectives, based on data from the Corpus of Historical American English, the Corpus...

doi.org

OUT NOW: Fonteyn, Manjavacas & De Regt show how large predictive language models can be used to (semi-)automatically annotate corpus data. In this example, Early Modern English -ing forms are automatically classified by means of the historical English model MacBERTh. #BERT doi.org/10.1075/ijcl...

Using machine learning to automate data annotation in corpus linguistics | John Benjamins

Abstract A wealth of linguistic data has been annotated by corpus linguists, and this extant annotated data can be used to automatically replicate and apply the linguist’s annotation scheme by means o...

doi.org

This paper appears as part of our Special Issue: Reproducibility, replicability, and robustness in corpus linguistics, from guest editors Martin Schweinberger and Michael Haugh. A reminder of the contributions to this issue: 1. doi.org/10.1075/ijcl... 2. doi.org/10.1075/ijcl... ...

Reproducibility, replicability, and robustness in corpus linguistics | John Benjamins

Abstract This introduction to the special issue Reproducibility, Replicability, and Robustness in Corpus Linguistics calls for more transparent and robust research practices in the field. It situates ...

doi.org

IJCL@ijcl.bsky.social · 11mo ago

OUT NOW: Maud Reveilhac and @geraldschneider.bsky.social present a replication study, applying their approach to stance detection to social media data. Their model is shown to be transferable and performs competitively alongside other machine learning methods. doi.org/10.1075/ijcl...

That's right: remmeber life before the Covid-19 pandemic? @journolinguist.bsky.social explores potential nostalgic markers in news about Covid as a methodological reflection on hypothesis-testing in #corpuslinguistics What do we learn when we don't get expected results? doi.org/10.1075/ijcl...

doi.org

Anna Marchi 🌈@journolinguist.bsky.social · 11mo ago

thrilled to have my paper "Hypothesis-testing in corpus-assisted discourse studies A methodological exploration" published in @ijcl.bsky.social www.jbe-platform.com/content/jour... Thanks to colleagues, reviewers and editors for the precious feedback #corpuslingustics

OUT NOW: Alan Partington and @diegolieugenia.bsky.social advance Lexical Priming (LP) theory, responding to Michael Hoey's desire that the theory be tested on discourse types that go beyond newspaper texts and in languages other than English – in this instance, Japanese. doi.org/10.1075/ijcl...

Lexical Priming theory | John Benjamins

Abstract This paper is an early step in a wider project which, on the behest of the late Prof Michael Hoey, attempts to review the evolution of Lexical Priming (LP) theory since its first appearance i...

doi.org