We're looking for a CV/ML Engineer to help us improve the machine learning systems that power iNaturalist's species identification and geographic range modeling. If you're excited to help build tools that help millions of people engage with nature, we'd love to hear from you! Apply: buff.ly/YZqaW6c
Mike Trizna
@miketrizna.bsky.social
Data Scientist focused on AI/Data Literacy, responsible applications of AI for Libraries, Archives, and Museums
Derived datasets are bigger on Hugging Face Hub than people realise. ~73% of analysed datasets on the Hub are derivatives of something else, i.e. cleaned, translated, extended, etc. Built an explorer that infers the missing lineage from content: huggingface.co/spaces/davan...
AI agents finally have a proper CLI for Jupyter notebooks. nb-cli lets agents read, write, execute, and search notebooks without a running server, built in Rust, optimized for LLM context windows. Read the blog: blog.jupyter.org/nb-cli-a-com...
nb-cli: A Command-Line Interface for AI Agents and Notebook Automation
The rise of AI coding agents has transformed how we think about developer tools. Large language models like Claude, GPT, and others are…
blog.jupyter.org
Thanks so much for putting "The Case for Boring AI" right up front! I've been carrying that banner for a long time -- but just in conversations and meetings -- so now I have this chapter and amazing book-in-progress to link to.
What is "AI for libraries" beyond a catalogue chatbot? IMO: design patterns (OCR, extraction, classification, search), and agents that both run them and develop the small models behind them. As part of work with @natlibscot.bsky.social started a book on this: danielvanstrien.xyz/ai-patterns-...
Exploring 218,567 pages of @bhl-au.bsky.social content using a "gilbert" curve. Just some of the content added by @nicolekearney.bsky.social and her team, HT to @cajunjoel.bsky.social for putting @biodivlibrary.bsky.social images on AWS which made this visualisation possible.
IBM just released the R2 generation of their Granite multilingual embedding models for retrieval, and the jump over R1 is very notable. Two models, both Apache 2.0: - granite-embedding-97m-multilingual-r2 (384-dim) - granite-embedding-311m-multilingual-r2 (768-dim) 🧵
New report is out with the latest open model adoption data we have gathered for Interconnects & The ATOM Project. At the surface level, we can see Chinese models continuing to accelerate in adoption. The report details much more. atomproject.ai/report
Good news for anyone working on their proposals for the next Fantastic Futures conference - the deadline is extended to April 16! ai4lam.org/submission-i... #FF2026 is in the US, but will be very hybrid so you don't need to travel there to present or attend many sessions #AI4LAM #MuseTech
Submission Instructions - Ai4lam
FF2026 will be both in‑person and hybrid! Submit your proposal via: Fantastic Futures 2026 – ConfTool Pro – Login The information below is for planning purposes and may change or expand. The Program…
ai4lam.org
EVoC is a library designed specifically for fast clustering of high dimensional embedding vectors. It can produce high quality clusters extremely efficiently, and requires little to no hyperparameter tuning. Better clustering than UMAP + HDBSCAN; faster clustering than KMeans.
You never know what data will be used for! I uploaded a @britishlibrary.bsky.social dataset to Hugging Face in 2022. IIRC one of my first PR to a HF repo! 4 years later, someone trains a Victorian chatbot on it More libraries should be sharing their public domain collections for AI to build on!
Want to talk to the past? Here' an LLM "trained entirely from scratch on a corpus of over 28,000 Victorian-era British texts published between 1837 & 1899, drawn from a dataset made available by the British Library" Quite different from an LLM roleplaying a Victorian. huggingface.co/spaces/tvent...
Now that AI Literacy Day is over: mail.cyberneticforests.com/human-litera...
Human Literacy
Something I Can Tell Students Now That I Am Not Teaching You and I probably both keep hearing that students should be working toward AI literacy. That you should know what to type into prompt windo...
mail.cyberneticforests.com
🚀 We've just open-sourced Embedding Atlas – a tool for exploring large embedding spaces through rich, interactive visualizations 📊.
The best path forward in AI requires technologists to be reflective/self-critical about how their work impacts society. Transparency helps this. Appreciate Bsky for flagging AI ethics &my colleague’s response. Let’s make informed consent a real thing. More later; Recommend: bsky.app/profile/cfie...
I've removed the Bluesky data from the repo. While I wanted to support tool development for the platform, I recognize this approach violated principles of transparency and consent in data collection. I apologize for this mistake.
Super excited to announce our best open-source language models yet. OLMo 2. These instruct models are hot off the press -- finished training with our new RL method this morning and vibes are very good.
I like this new analogy for working with LLMs by @emollick.bsky.social "treat AI like an infinitely patient new coworker who forgets everything you tell them each new conversation, one that comes highly recommended but whose actual abilities are not that clear" www.oneusefulthing.org/p/getting-st...
Getting started with AI: Good enough prompting
Don't make this hard
oneusefulthing.org
Big milestone for Project Jupyter 🚨: Jupyter has finalized the creation of the Jupyter Foundation, hosted by @linuxfoundation.org. www.linuxfoundation.org/press/linux-...
Linux Foundation Announces Formation of the Jupyter Foundation
Linux Foundation Announces Formation of the Jupyter Foundation
linuxfoundation.org
Bluesky uses AI internally to assist in content moderation, which helps us triage posts and shield human moderators from harmful content. We also use AI in the Discover algorithmic feed to serve you posts that we think you’d like. None of these are Gen AI systems trained on user content.