NDIF Team

@ndif-team.bsky.social

The National Deep Inference Fabric, an NSF-funded computational infrastructure to enable research on large-scale Artificial Intelligence. 🔗 NDIF: https://ndif.us 🧰 NNsight API: https://nnsight.net 😸 GitHub: https://github.com/ndif-team/nnsight

Can you tell when an AI model is lying? Announcing Aletheia's Quest, an AI lie detection challenge running this summer, organized by Cadenza Labs and NDIF. Multiple model organisms to interrogate and probe, $50K prize pool, no local GPU required.

Bild

NNsight 0.6 is out now! We directly address your feedback in our biggest release yet. Pain points included cryptic errors, slow traces, no remote execution of custom code, and limited vLLM support. We tackle all of these and more in this new release. 🧵 Here's what changed:

Watch Sam Marks present his work "Auditing Language Models for Hidden Objectives" in our new YouTube video! Sam's team ran a blind auditing game to assess efficacy of black box and white box techniques for LLM alignment auditing. 🔗 youtu.be/jZiOJTHqB6M

Auditing Language Models for Hidden Objectives with Sam Marks

Sam Marks leads Anthropic's Cognitive Oversight team, a subteam of Alignment Science. Sam's research focuses on settings where understanding something about ...

youtube.com

🔥I am super excited for the official release of an open-source library we've been working on for about a year! 🪄interpreto is an interpretability toolbox for HF language models🤗. In both generation and classification! Why do you need it, and for what? 1/8 (links at the end)

Bild
D

What does an LLM do when it translates from Italian "amore" to Spanish "amor" or French "amour"? That's easy! (you might think) Because surely it knows: amore, amor, amour are all based on the same Latin word. It can just drop the "e", or add a "u".

Bild

Do you wish you could run experiments on any model remotely from your laptop? In a future release, NDIF users will be able to dynamically deploy any model from HuggingFace on NDIF for remote experimentation. But before this, we need your help!

D
D

We’re excited to announce a new series of applied "mini paper" tutorials! The goal of this series is to help researchers get hands-on experience with findings, methods, and results from recent papers in interpretability using NNsight and NDIF.

Bild

We are prereleasing NNsight 0.5 today! We've added many highly-requested new features, including easier access to intermediate values within the forward pass, ability to run any code within NNsight’s tracing context (no more proxies!), and easier debugging.

Bild