Sebastian Majstorovic

@storytracer.com

Managing Director of datovis.com. Open Data Consultant for @eleutherai.bsky.social. Technical Director of @datarescueproject.org. Personal website: storytracer.com.

We are so very grateful for this honor! Thank you to our supporters and everyone who voted. Congratulations to everyone who was nominated -- we feel very privileged to be counted among your incredible work.

APDU@apduorg.bsky.social · 4w ago

Congratulations to the 2026 Data Integrity Award winners - @datarescueproject.org - Gina Plata-Nino of @fracposts.bsky.social - @mapresearch.bsky.social & Williams Institute - @hudgov.bsky.social & FEMA - Disaster Recovery Data Portal Details in 🧵

Read the feasibility study for the European Books Data Commons and discover the next steps! Explore how European libraries can make the full text of millions of digitised books available for research, innovation and contribute to a European AI infrastructure built on public values: bit.ly/4gs09tY

Feasibility study paves the way for a European Books Data Commons

The National Library of the Netherlands (KB) and the Europeana Foundation have just published a feasibility study for the European Books Data Commons (EBDC).

dataspace-culturalheritage.eu

I think VLM-based OCR might finally be close to working on historic newspapers! Many models I've tried before failed i.e. hallucinations, repetition loops, context overflow. Surya OCR 2 (a 650M model!) does a very good job!

A 1901 newspaper page overlaid with Surya OCR 2's detected layout: colored boxes around each block (text, section-header, picture) and numbered dots joined by a path showing the model's reading order across the columns.

Ran the same 650M model across 7 languages: German, Polish, Finnish, Russian, Serbian, Estonian, Swedish via @europeana.bsky.social newspapers. Still some challenging layouts, but pretty mind-blowing to me that a 650M model is doing this when a year ago 72B VLMs failed very often.

Daniel van Strien@danielvanstrien.bsky.social · last mo.

I think VLM-based OCR might finally be close to working on historic newspapers! Many models I've tried before failed i.e. hallucinations, repetition loops, context overflow. Surya OCR 2 (a 650M model!) does a very good job!

A 1901 newspaper page overlaid with Surya OCR 2's detected layout: colored boxes around each block (text, section-header, picture) and numbered dots joined by a path showing the model's reading order across the columns.

A federal judge in Massachusetts has ordered the Trump administration to restore all signs that were changed or removed at national parks across the country as part of President Donald Trump’s controversial directive last year.

Judge orders Trump administration to restore signs changed at national parks | CNN Politics

A federal judge in Massachusetts has ordered the Trump administration to restore all signs that were changed or removed at national parks across the country as part of President Donald Trump’s controv...

cnn.it

NuExtract3 (4B, Apache-2.0) does OCR *and* structured extraction. Point it at a dataset of scanned index cards + a JSON schema → clean catalog JSON One command on huggingface Jobs. (or skip the schema for plain Markdown OCR) Script + dataset 👇

A worked example of NuExtract3 turning a document image into structured data. Top: a terminal command — hf jobs uv run --image vllm/vllm-openai nuextract3.py index-cards cards-json --template schema.json. Middle (input): a scanned typewritten library index card reading "ABAD (Joseph), Captain, Spanish Army, letter of (1783), 5538, f.11." Bottom (output): the JSON extracted from the card — image_type "index_card", heading "Abad J.", heading_type "person", epithet "Captain, Spanish Army", and one entry with ms_no "5538", folios ["f.11"], description "letter of (1783)".

We're announcing two changes CourtListener API access: 1. Full API access is now open to everyone, including the PACER APIs that previously required a conversation with us. 2. Higher tiers are available through FLP memberships (including edu!) or commercial agreements. 👇 free.law/2026/05/07/a...

Full CourtListener Data Access via API Now Included with Membership

Researchers, journalists, developers, and vibe coders can now access the full CourtListener API, including PACER data, with a membership. No contact form. No waiting for approval.

free.law

“[T]he popularity of AskHistorians demonstrates that there can and should be more history-related jobs in both academic and public-facing settings. People want history to understand the world, and I can also tell you that those people do not want the AI-generated version of their answers.”

UNC Press@uncpress.bsky.social · 3mo ago

You may not immediately associate Reddit with academic rigor, but @askhistorians.bsky.social is changing that 🔥 We sat down with @dhowlett1692.bsky.social to talk about the subreddit, AMA's, public scholarship, combating misinformation online, and more! 👇 open.substack.com/pub/uncpress...