Here's my new blogpost on the dataset I have created alongside @ksolovev.com: we decided to solve the issue of data availability for research on news articles by processing, cleaning, tagging and indexing the whole news subset of the Common Crawl, resulting in cs2.uni-graz.at/blog/infini-...
Kirill Solovev
@ksolovev.com
Postdoc @ IDea_Lab, University of Graz. I study political communication and recommender algorithms in European multi-party democracies.
Check out the largest open multilingual preprocessed news corpus from Common Crawl News (built with @ruggsea.eurosky.social & @janalasser.eurosky.social). 1.36B articles | subsecond search | FOSS Demo: infini-news.uni-graz.at Paper: arxiv.org/abs/2605.18337 Code: codeberg.org/ksolovev/inf...
infini-news
infini-news: search 1.36B+ news articles with sub-second full-text search
infini-news.uni-graz.at