Heon Cha Haus was a great tip for Seoul @tedunderwood.com. I had their signature iced tea and an extremely tasty Red Bean & Cream Rice Cake. Not to mention the wonderful service! #DH2026
Sebastian Majstorovic
@storytracer.com
Managing Director of datovis.com. Open Data Consultant for @eleutherai.bsky.social. Technical Director of @datarescueproject.org. Personal website: storytracer.com.
Enjoying the weekend in Seoul before #DH2026 kicks off next week in Daejeon @dh2026daejeon.bsky.social.
Reading the Archive by Machine: An OCR Benchmark for Historians, 1612–1921 Here is version 1 of a working paper on the new OCR tools that are transforming digital history. working-papers-in-critical-search.github.io/paper-004-oc...
Reading the Archive by Machine – Working Papers in Critical Search
A benchmark of six OCR systems (Tesseract, olmOCR 2, Chandra 2, Infinity Parser 2, GLM-OCR, and Gemini 3.5 Flash) on human-transcribed archival documents spanning 1612–1921: early-modern print, ninete...
working-papers-in-critical-search.github.io
We are so very grateful for this honor! Thank you to our supporters and everyone who voted. Congratulations to everyone who was nominated -- we feel very privileged to be counted among your incredible work.
Congratulations to the 2026 Data Integrity Award winners - @datarescueproject.org - Gina Plata-Nino of @fracposts.bsky.social - @mapresearch.bsky.social & Williams Institute - @hudgov.bsky.social & FEMA - Disaster Recovery Data Portal Details in 🧵
Congratulations to the 2026 Data Integrity Award winners - @datarescueproject.org - Gina Plata-Nino of @fracposts.bsky.social - @mapresearch.bsky.social & Williams Institute - @hudgov.bsky.social & FEMA - Disaster Recovery Data Portal Details in 🧵
Congratulations Fireworks Display
ALT: Congratulations Fireworks Display
static.klipy.com
Read the feasibility study for the European Books Data Commons and discover the next steps! Explore how European libraries can make the full text of millions of digitised books available for research, innovation and contribute to a European AI infrastructure built on public values: bit.ly/4gs09tY
Feasibility study paves the way for a European Books Data Commons
The National Library of the Netherlands (KB) and the Europeana Foundation have just published a feasibility study for the European Books Data Commons (EBDC).
dataspace-culturalheritage.eu
@storytracer.com shows how volunteers or #GLAMS workers can be responsible actors in the current political climate by protecting datasets from war, political censorship and defunding! #glamlabsfutures
I think VLM-based OCR might finally be close to working on historic newspapers! Many models I've tried before failed i.e. hallucinations, repetition loops, context overflow. Surya OCR 2 (a 650M model!) does a very good job!
Ran the same 650M model across 7 languages: German, Polish, Finnish, Russian, Serbian, Estonian, Swedish via @europeana.bsky.social newspapers. Still some challenging layouts, but pretty mind-blowing to me that a 650M model is doing this when a year ago 72B VLMs failed very often.
I think VLM-based OCR might finally be close to working on historic newspapers! Many models I've tried before failed i.e. hallucinations, repetition loops, context overflow. Surya OCR 2 (a 650M model!) does a very good job!
Read our newly published paper on AI 'The case for Public AI: making it happen with cultural heritage.' 400+ professionals, one shared position - cultural heritage should shape AI, not just feed it. 👉 bit.ly/4xLdyU3 #PublicAI #AI #ArtificialIntelligence #CulturalHeritage #DataSpace
Making Public AI reality: How cultural heritage can lead the way
After months of collaborative development and iteration with the data space community and the Europeana Initiative, we are excited to share the newly published paper, ‘The case for Public AI: making i...
dataspace-culturalheritage.eu
Unfortunate to see this outcome but the fight for the panels at Washington’s house will continue in the community! This is our history and we want to #SaveOurSigns
The three-judge panel unanimously agreed to toss out an injunction that ordered the National Park Service to restore interpretive panels telling the story of nine people enslaved at the site.
In the 90 minutes after @theguardian.com article went live, BHL received more than US$2,500 in donations. Thank you to everyone helping keep biodiversity knowledge free and open. Please read, share, and help us build momentum.
A bonanza for fans of the natural world: the digital library sharing 64m pages of scientific knowledge with everyone
The Biodiversity Heritage Library is an invaluable online archive of historic texts on species living and lost supplied by the world’s leading museums and universities. Now its future is in doubt
theguardian.com
A federal judge in Massachusetts has ordered the Trump administration to restore all signs that were changed or removed at national parks across the country as part of President Donald Trump’s controversial directive last year.
Judge orders Trump administration to restore signs changed at national parks | CNN Politics
A federal judge in Massachusetts has ordered the Trump administration to restore all signs that were changed or removed at national parks across the country as part of President Donald Trump’s controv...
cnn.it
OCR for Ancient Greek is constrained by a lack of open training data. But the deeper challenge tackled by 2 papers from Inria is doing OCR and document structure recovery together: section hierarchies, milestone numbering, marginal references. Paper 1: arxiv.org/html/2603.02...
You can now run SQL over 2.19 BILLION web pages — zero download. @commoncrawl.bsky.social April 2026 crawl + URL index are on Hugging Face Storage Buckets. DuckDB reads it straight over hf:// — I counted all 2.19B in ~35s. Or point your own agent at it 👇 huggingface.co/spaces/davan...
NuExtract3 (4B, Apache-2.0) does OCR *and* structured extraction. Point it at a dataset of scanned index cards + a JSON schema → clean catalog JSON One command on huggingface Jobs. (or skip the schema for plain Markdown OCR) Script + dataset 👇
We're announcing two changes CourtListener API access: 1. Full API access is now open to everyone, including the PACER APIs that previously required a conversation with us. 2. Higher tiers are available through FLP memberships (including edu!) or commercial agreements. 👇 free.law/2026/05/07/a...
Full CourtListener Data Access via API Now Included with Membership
Researchers, journalists, developers, and vibe coders can now access the full CourtListener API, including PACER data, with a membership. No contact form. No waiting for approval.
free.law
The Guardian: “‘Things Were Going Dark Left and Right’: The Race to Save US Government #Datasets Before They’re Deleted” www.infodocket.com/2026/05/07/t... @datarescueproject.org @envirodgi.bsky.social @archive.org #data #libraries #librarians
Thanks to @theguardian.com for highlighting the people involved in the effort to save and demonstrate the importance of federal public data. ❤️🛟 www.theguardian.com/us-news/2026...
‘Things were going dark left and right’: the race to save US government datasets before they’re deleted
Group has banded together to rescue data as Trump administration has removed or altered data on climate change, reproductive health, LGBTQ people and more
theguardian.com
As we head into the weekend, and near the end of our #BobsBurgers month of memes, I wanted to share this joyous depiction of our truly amazing DRP volunteers. They give me hope without intending to; they're creating the good they want to see in the world; they are all gold stars.
“[T]he popularity of AskHistorians demonstrates that there can and should be more history-related jobs in both academic and public-facing settings. People want history to understand the world, and I can also tell you that those people do not want the AI-generated version of their answers.”
You may not immediately associate Reddit with academic rigor, but @askhistorians.bsky.social is changing that 🔥 We sat down with @dhowlett1692.bsky.social to talk about the subreddit, AMA's, public scholarship, combating misinformation online, and more! 👇 open.substack.com/pub/uncpress...
We’ve signed the #OpenHeritageStatement to support a global call for equitable access to #PublicDomain heritage in the digital environment! Discover more about the Statement, how it links to our advocacy for open #CulturalHeritage and how you can sign ➡️ bit.ly/4sB1Ur0 @creativecommons.bsky.social
The Europeana Initiative is proud to sign the Open Heritage Statement
Read on to discover more about the Statement’s unified vision for the public domain, and how it links to our work and advocacy for open cultural heritage!
bit.ly
After Trump took office last year, the US Holocaust Museum quietly removed educational material about American racism from its website and canceled a workshop on the "fragility of democracy," @iriesentner.bsky.social reports. www.politico.com/news/2026/04...
‘Proactively fall in line:’ Holocaust Memorial Museum quietly changed content after Trump returned to office
Two former employees said they believed the museum was altering its content preemptively to avoid unwanted negative attention from the Trump administration.
politico.com
You never know what data will be used for! I uploaded a @britishlibrary.bsky.social dataset to Hugging Face in 2022. IIRC one of my first PR to a HF repo! 4 years later, someone trains a Victorian chatbot on it More libraries should be sharing their public domain collections for AI to build on!
Want to talk to the past? Here' an LLM "trained entirely from scratch on a corpus of over 28,000 Victorian-era British texts published between 1837 & 1899, drawn from a dataset made available by the British Library" Quite different from an LLM roleplaying a Victorian. huggingface.co/spaces/tvent...
📣 New historical data visualization! "How Fast Was the Mail?" is an interactive map showing how long information took to travel across the US between 1882-1908: cblevins.github.io/mail-time/ +
How Fast Was the Mail?
Explore mail transit times via railway between major U.S. cities, 1882–1908.
cblevins.github.io
I’m at #HSP2026 at MIT this week! I’ll be giving a talk Friday at 5:25pm entitled “Structural Priming Effects in Language Models are Less Human-like in Languages Other Than English”. Looking forward to chatting to everyone!
Interesting new piece from @iatp.bsky.social showing how confidence in USDA data is eroding under the Trump Administration, with serious consequences. www.iatp.org/farming-risk...
Farming in the Dark: Unreliable USDA data jeopardizes a sustainable farming transition
As farmers struggle to adjust to a changing climate that has led to an increased frequency of droughts, floods, and other extreme weather events, incomplete or inaccurate public data has become a grow...
iatp.org
New DRP post: #Philly interpretive panels returned to the President's House. We congratulate the local community that fought HARD for their return. We hope this inspires other communities to #SaveOurSigns www.datarescueproject.org/independence...
Independence National Historical Park - A Hopeful Update
We recently posted about the takedown of signs at the President’s House site at Independence National Historical Park. We also posted a call for more photos for Save Our Signs, both before and after…
datarescueproject.org