Josh Rhodes

@joshrhodes.bsky.social

Lecturer in Digital History at UCL. All things census related + agrarian/industrial development of Britain 16th-19th centuries. https://www.joshuarhodeshistorian.com/

Convincing you to buy my book, day 3/31. Let's talk about climate disasters -- how are they experienced by people in their work or everyday lives? Hobos -- American migrant workers in the late 19th and early 20th centuries -- show us one possible answer.

Recently, FamilySearch digitized and uploaded tons of microfilmed records, including many from 18th-century Massachsuetts. They're using some sort of AI to transcribe/summarize the handwritten documents. I've noticed that the AI strips out references to race and enslavement in 18thc documents.

What do we do with duplicates? The original and duplicate are placed in a Document Arena (or Docudome™), then we turn away from them and allow one to consume the other, unobserved. This is a double-blind trial, where we don't know which was the original. Thus, only the strongest document survives.

How are historians rethinking environmental and social history via improved OCR of imperial archives? Join us this Wednesday 3pm UK time to hear from @jimclifford.bsky.social and @historyjacob.bsky.social - registration link below.

Katie McDonough@kmcdono.bsky.social · 7mo ago

Join the Lancaster-Manchester Environmental #DH Seminar on March 11 @ 3pm UK (online) for a talk by @jimclifford.bsky.social & @historyjacob.bsky.social: "Solving OCR: Using olmOCR to Follow Commodities across the British World" www.eventbrite.co.uk/e/solving-oc... #dhist #ocr #envhist 🗃️

Ran the same OCR models on 68 pages of historic newspaper. Every model hallucinated or looped. DeepSeek-OCR-2, LightOnOCR-2, GLM-OCR – all melt down on dense newspaper columns. You can try yourself using this @hf.co dataset: huggingface.co/datasets/dav...

Image with historic newspaper on the left and output from OCR models sampled on the right. The output on the right shows okay starting text and then a lot of repetition.
Daniel van Strien@danielvanstrien.bsky.social · 8mo ago

Re-OCR'd the complete 1771 Encyclopaedia Britannica (2,724 pages) with a single command on @hf.co Jobs. - 0.9B model (GLM-OCR) ~$0.002/page ~$5 total on an L4 GPU Before (old Tesseract ocr) → After

Screenshot of old vs new ocr. 

old ocr text is garbled. New ocr much cleaner.

Any ideas what the word/s between 'necessary' and 'Registrar' are in this marginal note from the 1851 English census? I thought I was pretty sure but keep doubting if I've got it right!

Image of part of 1851 England Census return