Josh Hadro

@hadro.bsky.social

Librarian and digital scholarship/data person working in cultural heritage. Also a former journalist and current local journalism booster in #Brooklyn / #NYC I work for the largest library in the world but all skeets and thoughts here are my own

Correct me if I’m wrong, but wasn’t James Carville’s last successful political insight closer in time to JFK’s assassination than it is to today?

Feels like we're in an era related to that old plumber anecdote about charging for years of experience, not the actual action: Using LLM tools to build interfaces and datasets: mostly simple Having 20 years of experience w/ common pitfalls, appropriate uses, and where to look for data: priceless

I don't generally trust automated historical geocoding because streets change often, but for a work-in-progress the NYC addresses are basically stable and useful as a broad indicator of where listings clustered.

A screenshot of a map data viz, with bubbles emanating from the center of NYC neighborhoods, sized by the number of entries contained in that neighborhood. Central Harlem (772), Midtown (255), and Bed-Stuy (196) stand out as the largest bubbles.

Dream event! I’ll be talking about my Directory Pipeline work, and some of the projects my data extraction tools have enabled — it will be fun, I promise

Paul Ford@ftrain.bsky.social · 2w ago

Library-focused event next week @ the Aboard office, featuring @hadro.bsky.social, regarding the intersection of The Most Forbidden Subject (LLM-based AI) and The Most Beloved Subject (librarianship and open access). Free, open, friendly, snacks. luma.com/aboard-p9g1

Hey, New Yorkers: Mayor Mamdani is recreating the city's digital service, after Mayor Adams destroyed it. They're hiring devs, designers, and product folks to work on-site in Brooklyn, starting now. You don't even have to be a city resident. Get on it!

NYC PIT Crew

nyc.gov

Getting pretty close to finishing a first pass on a fun personal project that brings together travel listing across a bunch of different publications from ~1930-1966

A data visualization of an entry for “Liberty Apartment Hotel” in Atlantic City listed in seven different travel guides

The bottleneck isn't model size, it's data. And pooling pays off: adding one collection made the model better at collections it had never seen. Every institution that contributes makes the shared model better for everyone. I'd love to see this kind of thing built by the GLAM sector!

Say, for example, if we had 3 examples of labeled city directory data from every state in the US, we could easily train a generic city directory parser! (I’m actually working on this, and I’ll have it done relatively soon, but @danielvanstrien.bsky.social’s point holds in so many other cases!)

Daniel van Strien@danielvanstrien.bsky.social · last mo.

If libraries, archives and museums pooled their (labelled) data, they could build state-of-the-art open models for the things they actually care about! I tried a small version: one open model (NuExtract-3, 4B) fine-tuned to read archival index cards across several collections.

I cannot urge you strongly enough to spend a random evening every now and then hand-transcribing 150 year old city directories from where you live, because you will find out that an inventor from 1876 lived around the corner from you!

A screenshot of a custom-built transcription tool for Brooklyn city directories, showing an entry for a Lucius K Gibbs who was in inventor in 1876

Calling library and general GLAM data people 📜📚: If you were looking for training data related to historical city directories, where would you look? I'm considering fine-tuning a small GLAM model to parse 19th and early 20th century city directories, but I don't want to reinvent the wheel!

I did this with Emily Wilson’s Odyssey when my kid was in the NICU for 30 days. Not sure the kiddo cared what I was reading given they were -2 months old, but I can confirm it was a fun version to read out loud

Ryan Moulton@moultano.bsky.social · 2mo ago

I've been reading Emily Wilson's Odyssey out loud to my kids and it absolutely rocks for that purpose. So much fun to read aloud and flows so smoothly. The kids would never have sat still for her wordier laborious predecessors.

Can an LLM agent help me with all the "class action notice" emails I get and wake me up if there's ever actually a genuine payout for the various ways technology companies have violated my rights in recent years?

My 2yo was really grooving to St Thomas when it came on the radio this AM. Only later did I realize why all the inter-segment music buttons were by Sonny Rollins. RIP to one of the true greats.

Thanks so much for putting "The Case for Boring AI" right up front! I've been carrying that banner for a long time -- but just in conversations and meetings -- so now I have this chapter and amazing book-in-progress to link to.

Daniel van Strien@danielvanstrien.bsky.social · 3mo ago

What is "AI for libraries" beyond a catalogue chatbot? IMO: design patterns (OCR, extraction, classification, search), and agents that both run them and develop the small models behind them. As part of work with @natlibscot.bsky.social started a book on this: danielvanstrien.xyz/ai-patterns-...