Marco Herzog

@marcoher.bsky.social

Connecting AI research & Applied AI since 2018 | Previously: ML, NLP, RecSys | Now: GenAI, LLMs, RAG | Natural science, tech and climbing is my thing.

If you are a leader, you shouldn't ignore this paper from MIT for your business and your workforce. This research from a large U.S. firm's R&D department shows real-world effects of AI — positive and negative. These findings can be adopted by other businesses and departments as well. 🧵 #AI #ML

If this is not telling an interesting story, than what does? Accenture makes more money with GenAI than OpenAI. GenAI revenue: OpenAI: $3.4bn Accenture: $3.7bn I’m still thinking what that means…

Bild

Unveiling the true inspiration behind the ‘attention’ operator in Transformers! From Bahdanau’s emails to Karpathy, we learn how attention transformed neural networks. Surprising that ‘Attention is All You Need’ outshines its predecessor by Bahdanau et al. #AI #ML #Transformers

Bild

By far the best visualization and walkthrough of a Transformer I've ever seen. You can explore the algorithm down to every add & multiply, seeing the whole process in action. With animation and explanation. That was definitely created with a lot of passion and work. By Brendan Bycroft. #AI #LLM #LM

@bsky.app Impression stats are really missing. Not necessarily visible for everyone on the post, but at least for the post creator. 1. To know how big the audience really is 2. To see the actual quality of post (engagement/Impressions)

Do you want to see the new Anthropic Model Context Protocol in action? It was announced just 3 d ago, and we already see great apps. This time build by Alec Velikanov. Take this project as an example to build something yourself. I'm off, need to cook 7 courses with the stuff Claude ordered. #AI

What can we do about the benchmark fatigue for #LLM? More people I speak to don’t take them seriously anymore, and I can’t blame them. I still hesitate, but I’m about to drop them as well. I think we need something new for #AI eval. Or is there something I’m not aware of?

Once moving away from naive RAG, chunking becomes pretty important. I use semantic chunking with BERT, and I'm quite happy with it. I didn't use this, but looks promising. Quite handy if you don't use bigger libs like Haystack, LangChain, LLamaIndex etc. github.com/bhavnicksm/c... #LLM #RAG #AI

GitHub - bhavnicksm/chonkie: 🦛 CHONK your texts with Chonkie ✨ - The no-nonsense RAG chunking library

🦛 CHONK your texts with Chonkie ✨ - The no-nonsense RAG chunking library - bhavnicksm/chonkie

github.com

I had a talk with Elena Samuylova from Evidently AI and she was super helpful. Evidently is an open-source framework to evaluate, test and monitor ML and LLM-powered systems. Does anyone have experience with it in prod env or an opinion on other eval tools?

Bild

Quite an unusual benchmark. RE-Bench tests AI Agents against human experts in ML research tasks. Key findings: • AI shines in sprints up to 4 h • Humans lead in longer projects • Sonnet 3.5 and o1-prev do substantially better than humans given 2 h • Human improvement is much steeper #AI #ML

Bild

Navigating Bluesky's content curation tools 🧭 • Lists: Organize accounts • Moderated Lists: Tighter content control • Starter Packs: Onboard new users (max 150) • Feeds: Custom algorithmic timelines I was lost at first, but here's what I learned! 🧵👇

Bild

Actually, an interesting thought for generalization in AI. It seems there is a connection between spacial location and latent context similarly to influence decision-making, despite the task being non-spatial, at least in the hippocampus. Is not Fei-Fei Li currently working on something similar?

Andrew MacAskill@macaskillaf.bsky.social · 2y ago

Huge congrats to @karyna-mi.bsky.social for her paper published today in Science! She found that the hippocampus is really important for a key strategy we use to make decisions called hidden state inference! 🧪 🧠https://www.science.org/doi/10.1126/science.adq5874 1/7

🔬 𝗔𝗜 𝗔𝗧 𝗪𝗢𝗥𝗞 — 𝗘𝗻𝗵𝗮𝗻𝗰𝗶𝗻𝗴 𝗘𝗳𝗳𝗶𝗰𝗶𝗲𝗻𝗰𝘆, 𝗗𝗶𝗺𝗶𝗻𝗶𝘀𝗵𝗶𝗻𝗴 𝗦𝗮𝘁𝗶𝘀𝗳𝗮𝗰𝘁𝗶𝗼𝗻 New MIT study. The numbers tell a striking story: • Overall 𝗽𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝘃𝗶𝘁𝘆 ⬆️ 𝟰𝟰% • Patent filings ⬆️ 39% • BUT 𝗷𝗼𝗯 𝘀𝗮𝘁𝗶𝘀𝗳𝗮𝗰𝘁𝗶𝗼𝗻 ⬇ 𝗳𝗼𝗿 𝟴𝟮% Here's the twist: • Top performers: 𝗣𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝘃𝗶𝘁𝘆 𝘀𝘂𝗿𝗴𝗲 ⬆️ 𝟴𝟭% • Bottom third: Minimal gains

BildBildBildBild

@bsky.app Seems like you are getting momentum. Congrats! Feels like a lot of excitement and genuine interactions are shared here. Would you mind sharing any features and other plans users can expect for the future?

You heard about DSPy? Let's talk TextGrad. The upgraded version in AI optimization. I've been diving into TextGrad lately, and I thought I'd share some insights. It's a Python framework from Stanford that's making waves in AI optimization. Here's the lowdown: 🧵