Daniel van Strien

@danielvanstrien.bsky.social

Machine Learning Librarian at @hf.co

The Europeana Newspapers dataset on @hf.co now has an `alto` config: the raw ALTO XML for all 5.9M pages. The coordinates for every word, line and block, per-word OCR confidence and font info that the flattened text dropped are back. huggingface.co/datasets/biglam/europeana_newspapers

biglam/europeana_newspapers · Datasets at Hugging Face

We’re on a journey to advance and democratize artificial intelligence through open source and open science.

huggingface.co

Used Astra + Jobs to see how well this model performs on the full olmOCR-bench Unsurprisingly, it doesn’t do brilliantly overall: 36.8% But for a ~16M-parameter recogniser, I think 74.4% on long/tiny text and 57.9% on multi-column pages are pretty interesting.

Daniel van Strien@danielvanstrien.bsky.social · last mo.

A 15.9M-parameter specialist OCR model outperforms many VLMs hundreds of times larger. On a 2,165-page historical eval dataset, Kraken PP-OCRv6 ranks #4 on reading CER and #1 when long-s, ligatures and case are preserved. huggingface.co/spaces/fineb...

Plot showing parameters vs performance. Kraken is top left.

OCR for Japanese manga, Swedish handwriting or Arabic print? There’s a growing range of OCR models on @hf.co: VLMs, dedicated text recognisers and complete OCR pipelines. I’ve gathered 41 models into four collections, with short notes to help you choose: huggingface.co/collections/...

OCR on the Hub - a davanstrien Collection

Curated OCR models for documents, languages, handwriting and text in images. Browse four collections with short practical notes.

huggingface.co

📈 New blog post: Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers. As a practical example, I finetuned a ColBERT-style model for medical retrieval. 14.5 hours on one RTX 3090, and it beats every general-purpose retriever I could find. Thread 🧵

Bild

Every illustration in the Encyclopaedia Britannica dataset now has an instance mask: 411,385 cut-outs from 115,293 pages, 1768–1929, each linked to its full-resolution scan. Public domain, no image generation involved. Work in progress.

The left image shows an illustration surrounded by text. right image shows image cropped via mask prediction
Daniel van Strien@danielvanstrien.bsky.social · 2mo ago

Uploaded a dataset of 115,293 illustrated pages from the Encyclopaedia Britannica, 1st edition (1768) to 14th (1929) to the Hub huggingface.co/datasets/big...

grid showing examples from the dataset

Every illustration in the British Library's 19th-century book images dataset now has an instance mask: 1.02M clean cutouts of public-domain engravings, no image generation involved! One small segmentation model, $3.24 of compute. Model, masks and a search Space with cutout view all open.

Bild

A film catalogue tells you what a film is about, not what happens inside it. So: 1,864 public-domain Prelinger films, described every ~60 seconds by an open 2B video model. 23,148 searchable moments. Search "typing on a computer keyboard", land on the second it happens.

Bild