@eleutherai.bsky.social

We at @cdt.org have an amazing lineup for our 3/26 event on AI and internet scraping with @gtowntechlaw.bsky.social, including @archive.org, @wikimediafoundation.org, @cloudflare.social, @sparcopen.bsky.social, @nytimes.com, @eleutherai.bsky.social, and more! RSVP: docs.google.com/forms/d/e/1F...

Kevin Bankston@bankston.bsky.social · 5mo ago

Internet + AI policy friends: RSVP to join @cdt.org at @georgetownlaw.bsky.social on Thurs March 26 for a lively morning of panels and talks on internet scraping and the future of the open web in the AI age--with perspectives from AI labs, web publishers, scrapers, & a range of experts.

We’re bringing back a Community Spotlight talk series, highlighting cool work being done by members of our community. We’re kicking it off with a talk on running diffusion-based world-models in real time on consumer hardware. Jan 9th at 2 pm US Eastern Time

Can you train a performant language model using only openly licensed text? We are thrilled to announce the Common Pile v0.1, an 8TB dataset of openly licensed and public domain text. We train 7B models for 1T and 2T tokens and match the performance similar models like LLaMA 1 & 2

Bild

How do a neural network's final parameters depend on its initial ones? In this new paper, we answer this question by analyzing the training Jacobian, the matrix of derivatives of the final parameters with respect to the initial parameters. https://arxiv.org/abs/2412.07003

Bild

The latest from our interpretability team: there is an ambiguity in prior work on the linear representation hypothesis: Is a linear representation a linear function (that preserves the origin) or an affine function (that does not)? This distinction matters in practice. arxiv.org/abs/2411.09003

Refusal in LLMs is an Affine Function

We propose affine concept editing (ACE) as an approach for steering language models' behavior by intervening directly in activations. We begin with an affine decomposition of model activation vectors ...

arxiv.org