Matthew Kenney

@baykenney.bsky.social

Founder - Algorithmic Research Group. Previously: Senior ML Engineer at Apple, Asst. Research Prof at Duke University, Duke Data Science | PSU '15 | Cornell '11

There's rightly a lot of excitement around Karpathy's autoresearch. We've been studying at ARG for a couple years now: what happens when you put an AI agent in a loop and let it run experiments, evaluate results, and iterate without you. We've built a bunch of benchmarks and tools to measure this. 🧵

Over the past ~2 years I’ve been working hard on agents, models, and datasets to understand what recursive self-improvement might look like, and what path supporting infrastructure for this line of research might take. Very excited to open source some of that work starting today

A lot of ML tools help you implement. Not many help you think. When I’m exploring a new research direction, I don’t want another search engine or citation graph. I want something that’s actually read the literature, can suggest promising directions, and helps me reason through tradeoffs.

AI for science could be more impactful than chatbots. It is already helping win Nobel prizes and accelerating drug development and materials discovery. Today we published an essay about it: why it matters, how it’s happening and its implications. Here is a summary from an econ / social sci lens.

Bild

Important point that the open protocol makes extracting data from bluesky easy. Can't have it both ways. I like the protocol and think this site is well designed, but that means anyone can and will analyze these posts (if there is value to them, which I'm honestly less convinced of than some)

Stella Biderman@stellaathena.bsky.social · 2y ago

OpenAI doesn't want this dataset. Any company that wants Bluesky data will press a button and go get all posts on Bluesky. The very design of the website makes that trivial to do. This doesn't help OpenAI *at all*.

A dataset of 1 million or 2 million Bluesky posts is completely irrelevant to training large language models. The primary usecase for the datasets that people are losing their shit over isn't ChatGPT, it's social science research and developing systems that improve Bluesky.

Jeremy Howard @howard.fm · 2y ago

Did you know that 99% of email today is spam? Your inbox isn’t 99% spam because AI is used to filter it. The same 99% will happen here too, but if AI researchers continue to get perma-banned for making available the datasets needed to filter it, it’s going to make this platform unusable.

I've created an initial Grumpy Machine Learners starter park. If you think you're grumpy and you "do machine learning", nominate yourself. If you're on the list, but don't think you are grumpy, then take a look in the mirror. go.bsky.app/6ddpivr

Post nicht verfügbar.