Search 250M+ words of current Bulgarian, compare real-world usage, and uncover emerging language trends with the Bulgarian Trends corpus. sketchengine.eu/bulgarian-trends-corpus/ #NLP #CorpusLinguistics
Sketch Engine
@sketchengine.eu
Sketch Engine is a linguistic search engine and corpus query system with text analysis tools and corpora in 100+ languages. Concordance, n-grams, term extraction, co-occurrences (Word Sketch) are only some of its features.
📚 Built a dictionary, corpus, or language technology resource? The Adam Kilgarriff Prize is now open for applications. #DigitalHumanities #computationallinguistics #lexicography kilgarriff.co.uk/prize/
English Trends now contains 90+ billion words and grows by over 100 million words each week. Follow current English usage in the largest English monitor corpus. Free trial available. #CorpusLinguistics #bigdata www.sketchengine.eu/english-tren...
Search 88M+ words of current Croatian, compare usage, and spot changing language trends with the Croatian Trends corpus. www.sketchengine.eu/croatian-tre... #Croatianlanguage #LanguageData #textnalalysis
700+ participants. Since 2001. Join Lexicom 2026 in Palermo 🇮🇹 (14–18 Sept) for hands-on training in digital #lexicography, #corpuslinguistics, dictionary building, and lexical computing. 🔗 lexicom.courses/lexicom-2026...
The new Albanian Trends corpus is now available in Sketch Engine! With 91M+ words and daily updates, it helps you explore present-day Albanian and monitor how language trends evolve over time. www.sketchengine.eu/albanian-tre... #Albanian #NLP #corpuslinguistics
📚 Adam Kilgarriff helped many of us see language through data. The prize in his name recognises outstanding work in lexicography and language technology. Applications are open. #lexicography #languagetechnology #NLP kilgarriff.co.uk/prize/
From the ELEXAI kick-off meeting in Ljubljana. Lexical Computing is now part of this EU Horizon project, alongside 24 other partners. #ELEXAI builds on the previous ELEXIS project to improve #LLMs with lexicographic data and apply LLMs in #lexicography.
How many languages did we meet at #PolyglotGathering 2026 in Brno? Quite a few! Thanks to everyone who came by our booth to chat and find out how many languages you can explore in Sketch Engine. It was great being there. #polyglot #multilanguage
Our new Romanian Trends corpus is now available! With 370M+ words and daily updates, it lets you explore contemporary Romanian and follow language trends as they develop over time. ske.li/romanian_tre... #NLP #Romanianlanguage #CorpusLinguistics
Our new Turkish Trends corpus is out! With 170M+ words and daily updates, it helps you study current Turkish and track language trends over time. www.sketchengine.eu/turkish-tren... #Turkishlanguage #CorpusLinguistics #NLP
Greetings from #LREC2026 in Mallorca 🇪🇸 From poster sessions to our booth, it is great to connect with researchers, linguists, and the NLP community. Thanks to everyone stopping by for discussions on corpora and language technology. #NLP #ComputationalLinguistics
Save 20% with early-bird registration for Lexicom 2026 before 14 May 🚨 Join a 5-day workshop on digital lexicography, corpus-based dictionary building, lexical computing, and AI applications in lexicography. 🔗 lexicom.courses/lexicom-2026... #Lexicography #CorpusLinguistics #NLP
📚 The Adam Kilgarriff Prize recognises work connecting lexicography and language technology. If you’ve created a dictionary, corpus, NLP resource, or language tool, consider applying – or share with someone whose work deserves attention 👇 kilgarriff.co.uk/prize/
At last – a corpus with a bigger ego than a Sandalwood hero. 🎬 Kannada Trends in Sketch Engine now reaches 30+ million words. That’s a lot to explore. A true digital ocean of Kannada. 🌊 www.sketchengine.eu/kannada-tren... #corpuslinguistics #NLP
Join the 25th edition of Lexicom 2026 in Palermo 🇮🇹, 14–18 September! Be part of a workshop in #lexicography, #corpuslinguistics, #dictionaries, and lexical computing, attended by 700+ participants worldwide. 🔗 lexicom.courses/lexicom-2026...
Explore how Afrikaans words change through time with our latest corpus – ideal for discovering #neologisms and emerging language trends. ske.li/afrikaans_tr...
🔹A very small announcement: We’ve just published Dot Corpus, the tiniest text corpus in the world. You’ll read it in no time! 🌐 www.sketchengine.eu/dot-corpus/ 👉 app.sketchengine.eu#dashboard?co...
Explore how Malay words change through time with our latest corpus — ideal for discovering #neologisms and emerging language trends. ske.li/malay_trends
460M+ words of 🇲🇹 Maltese language data now available in one corpus. A useful resource for research and #NLP on this unique Semitic language written in the Latin script. Special thanks to the University of Malta for making this possible. www.sketchengine.eu/maltese-refe... #corpuslinguistics
📚🔎 The Adam Kilgarriff Prize is open for applications. If you created a dictionary, corpus, or language tool, consider applying or sharing the opportunity. #lexicography #NLP kilgarriff.co.uk/prize/
The Telugu Web 2021 corpus, with 100+ million words and part-of-speech tagging, is now available in Sketch Engine! #corpuslinguistics, #digitalhumanities, #linguistics www.sketchengine.eu/tetenten-tel...
Registration is open for Lexicom 2026 in Palermo 🇮🇹! Since 2001, this workshop in #lexicography and #corpuslinguistics has welcomed 700+ participants worldwide. Join the community and take part in the next edition, 14–18 September 2026. 🔗 lexicom.courses/lexicom-2026...
Our new Chinese corpus in Traditional Chinese (繁體字) is now available. It is part-of-speech tagged and partly annotated for topics and genres. A useful resource for research and language teaching. #corpuslinguistics #digitalhumanities www.sketchengine.eu/zhtenten-chi...
We’ve published a new Chinese corpus in Simplified Chinese (简体字). It is part-of-speech tagged and partly annotated for topics and genres. A useful resource for research and language technology. #corpuslinguistics #linguistics #nlp www.sketchengine.eu/zhtenten-chi...
An example of Sketch Engine used outside the field of pure linguistics. This study in media discourse analysis will be published in @nature.com www.nature.com/articles/s41... #MediaRepresentation #discourseanalysis #corpuslinguistics
We’ve published the Urdu Corpus 2021 in Sketch Engine, with 328 million words and topic and genre classification. Urdu is the 11th most spoken language worldwide (Ethnologue, 2025). 🔗 www.sketchengine.eu/urtenten-urd... #corpuslinguistics #TextAnalysis #اردو
The new Latvian Corpus 2021 now available in Sketch Engine. The corpus is enriched with part-of-speech tagging and lemmatization. Perfect for #corpuslinguistics, #digitalhumanities, #linguistics, #lexicography, and #nlp.
📢 Registration is open for Lexicom 2026 in Palermo 🇮🇹! Apply for this hands-on workshop on #lexicography, #corpuslinguistics, and #dictionaries. Learn from experts, explore new tools, and build your skills. 📅 14–18 September 2026 🔗 lexicom.courses/lexicom-2026...
You can search for multiple variants at the same time in the Word Sketch tool. Just add a comma between them – Christmas, Xmas – to see the results for both: ske.li/bav0