Banning all teens from social media hasn't ever seemed like a great idea to me. A much healthier approach is for platforms to build smoother on-ramps for teens that are developmentally appropriate for each child's age. 🧵 [1/n]...
Samidh
@samidh.bsky.social
Co-Founder at Zentropi (Trustworthy AI). Formerly Meta Civic Integrity Founder, Google X and Google Civic Innovation Lead, and Groq CPO.
I continue to be impressed with Zentropi. My first thought is that they were a replcement for Perspective API that Jigsaw was ending. And yes, but so much more. So far they can answer virtually any question about a text of almost any size my work cares about that CAN be answered with a Yes or a No.
Today, in a single day, @dwillner.bsky.social and I had meetings with people in England, Turkey, NYC, SF, Argentina, and Australia. Very inspiring to see four continents of folks all using Zentropi and united in the earnest work of building a better internet.
The real question is whether these classifiers can find all the dogs that Dave has at home.
Zentropi can now label video, not just text and images. You simply write a policy in plain language and the model applies it to clips up to 60 seconds. Our optimizers also work with video, taking your initial policy and dialing it in on your data. Check it out! blog.zentropi.ai/zentropi-now...
I distinctly remember being at Meta in the wake of the Christchurch massacre, when horrific videos were circulating across Facebook without end. The technologies we had for being able to accurately classify videos just didn't exist. [1/n]
Today we're releasing Coop 1.0, the world's first free, open source content review & enforcement system any org can self-host and build on. For the first time, any org, whatever its size or budget, can review, act on, and report CSAM end to end, for free. roost.tools/blog/coop-1-...
Coop 1.0: World’s First Free, Open Source Child Safety Infrastructure for Every Platform
Robust Open Online Safety Tools or ROOST is a new non-profit entity designed to address the urgent need for accessible, high-quality safety tools in the rapidly evolving digital landscape.
roost.tools
Today, in a single day, @dwillner.bsky.social and I had meetings with people in England, Turkey, NYC, SF, Argentina, and Australia. Very inspiring to see four continents of folks all using Zentropi and united in the earnest work of building a better internet.
Can your LLM actually follow your content policies, or will it just revert to its own static training? Today we're introducing a new benchmark that we call policy steerability that tries to measure this concept: blog.zentropi.ai/introducing-...
Beyond Static Accuracy: Introducing the Policy Steerability Benchmark
A new way to measure how likely a model is to accurately follow your rules
blog.zentropi.ai
Very cool! @julietshen.bsky.social's independent tests show that our new model CoPE-B cooks :-) Direct link to her results: github.com/julietshen/c...
github.com
I am NOT an AI engineer or AI researcher, but I tried to do a little evaluation of CoPE-B vs CoPE-A vs gpt-oss-safeguard github.com/roostorg/mod... lmk what you think, and we'd love for more evaluations to be part of the ROOST Model Community! cc @samidh.bsky.social
The ROOST Model Community is growing! Today we welcome Zentropi's CoPE-B-A4B, a bring-your-own-policy model that's got the power of 25B but runs on only 4B active parameters. It's a fast, low-cost model that can be used on its own or as a first pass before larger models! roost.tools/blog/welcomi...
Welcoming Zentropi's CoPE-B-A4B to the ROOST Model Community
Robust Open Online Safety Tools or ROOST is a new non-profit entity designed to address the urgent need for accessible, high-quality safety tools in the rapidly evolving digital landscape.
roost.tools
New in the ROOST Model Community: zentropi.ai 's CoPE-B-A4B is here. roost.tools/blog/welcomi...
Welcoming Zentropi's CoPE-B-A4B to the ROOST Model Community
Robust Open Online Safety Tools or ROOST is a new non-profit entity designed to address the urgent need for accessible, high-quality safety tools in the rapidly evolving digital landscape.
roost.tools
a few use cases I can think of for CoPE-B (and BYOP models): if you're a platform that's ok with NSFW role play but not age play, you can create a custom CoPE model that looks just for that. Many free text classifiers come with baked-in ideas of morality that might not fit your community
We're releasing CoPE-B in collaboration with @roost.tools and will support it through the ROOST Model Community. It's a great forum for advancing AI-powered trust & safety tooling and we're excited to be a part of it. We'll be active there if you're building with CoPE-B and have questions! 🧵 9/9
Exactly right. Speech shouldn't be ruled by the platform hegemons. Your rules should rule.
For anyone in the trust and safety space this is big. A powerful, but small, *open source* classifier not weighed down by what some giant AI company decides you must want in your classifier ...
We are THRILLED to announce that @roost.tools is jointly releasing CoPE-B with the Zentropi team. We believe that everyone should have access to openly licensed models designed for safety use cases that can be tuned to their community's norms. Join us at RMC office hours next week to learn more!
We're releasing CoPE-B in collaboration with @roost.tools and will support it through the ROOST Model Community. It's a great forum for advancing AI-powered trust & safety tooling and we're excited to be a part of it. We'll be active there if you're building with CoPE-B and have questions! 🧵 9/9
Super pumped to release CoPE-B, our latest policy-adaptive content classification model. It delivers frontier-level accuracy in a self-hostable package that's orders of magnitude cheaper to run-- opening up new possibilities in trustworthy platform design. Details: blog.zentropi.ai/meet-cope-b-...
Meet CoPE-B: Frontier-Quality Content Classification You Can Self-Host
TL;DR: * Today we're releasing CoPE-B, our next-gen small language model for policy-adaptive content classification * CoPE-B-A4B (text-only) is open weights under Apache 2.0 and free to use * CoPE...
blog.zentropi.ai
@samidh.bsky.social and I are releasing CoPE-B today, the next version of our policy-adaptive content classifier. It delivers at-or-better-than-frontier classification while being self-hostable, faster, and cheaper to run. Full writeup with benchmarks at blog.zentropi.ai/meet-cope-b-.... 🧵 1/9
Meet CoPE-B: Frontier-Quality Content Classification You Can Self-Host
TL;DR: * Today we're releasing CoPE-B, our next-gen small language model for policy-adaptive content classification * CoPE-B-A4B (text-only) is open weights under Apache 2.0 and free to use * CoPE...
blog.zentropi.ai
Very excited to have this public. The @oversightboard.bsky.social's use of Zentropi to study child marriage content on Meta platforms was very cool to support: blog.zentropi.ai/how-the-over...
How the Oversight Board uses Zentropi to study policy impact at scale
The Oversight Board used Zentropi to analyze a large dataset of content that potentially violated Meta’s policies on human exploitation. The tool helped the Board reduce the project timeline from week...
blog.zentropi.ai
Check out how the @oversightboard.bsky.social used Zentropi to better understand how child marriage-related content manifests on Meta's platforms. Fantastic example of how advanced content labeling technologies can strengthen both our online and offline world. blog.zentropi.ai/how-the-over...
How the Oversight Board uses Zentropi to study policy impact at scale
The Oversight Board used Zentropi to analyze a large dataset of content that potentially violated Meta’s policies on human exploitation. The tool helped the Board reduce the project timeline from week...
blog.zentropi.ai
It has been incredible partnering with character.ai since the very start of zentropi.ai. We're excited to share some details of that partnership with this case study. Anyone creating AI-powered systems might find it interesting! blog.zentropi.ai/how-zentropi...
How Zentropi partners with Character.ai
Character.ai takes safety seriously. With millions of users creating and chatting with AI characters every day, the team invests heavily in systems that help protect their community — and they're alwa...
blog.zentropi.ai
In 2017, it took us months to define “political ad” at Facebook. Recently, I built two political content classifiers in an afternoon using Zentropi AI created by @dwillner.bsky.social and @samidh.bsky.social Why AI content moderation is good, actually — in this week’s newsletter 👇
Why AI Makes Content Moderation Better, Not Worse
Building a political content labeler with AI — what actually works
open.substack.com
One of the things we've been thinking about a lot at Zentropi is: what happens when AI agents need to make judgment calls about content — not humans reviewing a queue, but agents acting autonomously?
There's a major gap in content safety tooling: classifiers typically only score complete text. When you're working with generative AI, "complete text" means the user already saw it. That's too late. So we built a streaming classifier that we're releasing today! Here's what we did and why. 🧵...
If you're not using either tool yet, now's a good time to try both! Zentropi's Community Edition is free and gives you unlimited labelers. Coop is fully open source and runs on your infrastructure. :D
@dwillner.bsky.social and I have spent years watching T&S teams rebuild the same infrastructure from scratch. This is what it looks like when open tools actually work together instead. Really proud of this one and appreciative of @roost.tools's leadership! Details: blog.zentropi.ai/zentropi-now...
Zentropi is now integrated into Coop, @roost.tools's open source moderation platform. You can write a content policy in plain English on Zentropi, plug it into Coop as a signal, and have a moderation pipeline running in minutes.
Dave Willner, who led trust and safety at major tech firms and has cofounded a company that is developing an AI content classification platform, says LLM-driven technology can now accomplish classification at the scale necessary for moderation on large platforms. That has substantial implications.
AI is Removing Bottlenecks to Effective Content Moderation at Scale
Zentropi's Dave Willner says LLM-driven technology can now accomplish content classification at the scale necessary for moderation on large platforms.
techpolicy.press
I can has cats.
We just shipped image classification on Zentropi! You can write your criteria in plain English, then analyze images against them at scale using cope-b-12b, a new multimodal model we trained Critically, @samidh.bsky.social used it to make a cat detector: blog.zentropi.ai/zentropi-now-labels-images/
Just shipped Zentropi's most requested feature: image classification! Now analyze images against your own policies, at scale. To power it we built cope-b-12b, a new multimodal model w/ native vision. Check out the cat detector we made in < 1 min. 🐱 blog.zentropi.ai/zentropi-now-labels-images/
Zentropi Now Labels Images
Building guardrails for visual content just got a lot easier. Today we're launching image classification on Zentropi and announcing cope-b-12b, a multimodal model that powers this experience.
blog.zentropi.ai
If you are looking for a technical description of how X rots your brain, look no further than their github post on the 'X algorithm'. It is pure, unadulterated behavioral engagement maximization that amplifies the very worst human impulses. github.com/xai-org/x-al...
GitHub - xai-org/x-algorithm: Algorithm powering the For You feed on X
Algorithm powering the For You feed on X. Contribute to xai-org/x-algorithm development by creating an account on GitHub.
github.com
Why are we just giving away all our secrets? Well, it is our hope that it helps the ecosystem further advance the state of the art in policy-steerable content classification, which is foundational to a more trustworthy internet.
We just published the methodology behind CoPE, our 9B parameter model that matches GPT-4o at content classification at 1% the size! The model is already open source, but now we're sharing our training technique. blog.zentropi.ai/how-we-built... 🧵 1/6
Dave just published a Zentropi labeler that can precisely identify requests at prompting an AI model to undress a person in a photo. The tools exist to easily deal with this problem -- platforms just need to choose to use them. If you are the developer of an AI system, please use this guardrail!
Over the weekend I used Zentropi to build a labeler that blocks requests to use AI to undress or sexualize real people. The labeler itself took maybe 30 minutes to a first solid draft, with another 30 minutes of tweaking. In my testing it's got an F1 of .98 on real examples pulled from X. 🧵 1/6