Juliet Shen

@julietshen.online

product @roost.tools. just another extremely earnest and Very Online idealist. long time Trust & Safety person full of personal opinions like the skeets below https://julietshen.online/

And we have basic machine-vision in place to help identify harmful content. We understand discomfort with these systems, but there's great potential for harm in live video; machine-assisted moderation has long been considered a T&S best practice. Humans will always make the final rulings on bans.

Most tooling for fighting fraud and abuse was built for 2010s social media. The abuse evolved. The tools mostly didn't. @julietshen.online at Marketplace Risk NYC on why open source and open weight models help: you run them yourself, so sensitive data never leaves your systems. github.com/roostorg

ROOST Head of Product Juliet Shen speaking at Marketplace Risk in NYC.

(1) This is amazing 👏 (2) The better automated content moderation gets, the more reasonable it is for govts to mandate automated censorship 🤐 (2) Having smaller/cheaper/more customizable como means there’s less reason to have centralized speech controls in the first place 🤓

Dave Willner@dwillner.bsky.social · 2d ago

At TrustCon this year I talked about a technique we’ve developed for automatically optimizing content-moderation policies, using an inversion of the binocular labeling approach Zentropi had already pioneered. Today we're shipping the tool that technique became. blog.zentropi.ai/optimizing-o...

FORTUNE WEAVE COMES OUT TOMORROW FORTUNE WEAVE COMES OUT TOMORROW FORTUNE WEAVE COMES OUT TOMORROW FORTUNE WEAVE COMES OUT TOMORROW FORTUNE WEAVE COMES OUT TOMORROW FORTUNE WEAVE COMES OUT TOMORROW FORTUNE WEAVE COMES OUT TOMORROW FORTUNE WEAVE COMES OUT TOMORROW FORTUNE WEAVE COMES OUT TOMORROW

My talk for the Trust & Safety Professional Association APAC Summit was accepted! I'll be presenting on "Open Safety Models and the New T&S Tooling Stack" I'm poking around at Asian language support and performance in bring-your-own-policy models and fine tuned ones: github.com/julietshen/v...

vibecheck/MULTILINGUAL.md at main · julietshen/vibecheck

Playground for evaluating various safety tools. Contribute to julietshen/vibecheck development by creating an account on GitHub.

github.com

As a result, this new optimizer has three distinct capabilities: policy-only correction rewrites the policy text to better match a labeled dataset. Label-only correction flags and fixes labels that no longer match the policy. Auto-optimization runs both in a loop until the whole stops improving.

To summarize the point - a policy and a set of examples labeled against it aren't independent things you can fix one at a time, because a given label can only be said to be meaningfully right or wrong relative to a specific policy text. So, to improve either you have to work on both simultaneously.

We were interested in studying the Bsky custom feed ecosystem, with a focus on feed creators, as arguably the first real instantiation of the vision of middleware providers for social media services like recommendation. The idea always seemed great in theory, but how sustainable is it really?

Tony Zhou@tyzhou.bsky.social · 4d ago

Bluesky gives users the power to choose their algorithms, but how does this actually work in practice? In our new preprint, we explore the ecosystem of Bluesky custom feeds and the feed creators behind them! 🧵 arxiv.org/abs/2609.12958

How should we think about user data, access control, and accountability? What tools or infrastructure + other support do middleware providers need to thrive? These questions have not all been sorted out yet since it’s still fairly new. But our interviews point to ways to improve on the status quo.

When we presented our findings to child safety experts, they were shocked that we were persistently finding NCMEC images. Even though X said it reported 1.3 million images to NCMEC this year, some were clearly slipping through — X did not answer specific questions about how it blocks abuse material.

We also partnered with the Canadian Center for Child Protection to analyze Grok images produced during the nudification episode in Jan. It found images of known CSAM victims that were modified by Grok, and new CSAM that it created in response to requests from users. www.nytimes.com/2026/09/11/t...

Elon Musk Has Pledged to Rid His Platform, X, of Child Sexual Abuse, but It Persists (Gift Article)

Reviews by the Canadian Center for Child Protection and The New York Times found that explicit images of children remain on the social media site owned by Elon Musk.

nytimes.com

Over six months, we tracked the appearance of child sexual abuse material on X. We found images from the National Center for Missing and Exploited Children's database, which is considered one of the most highly vetted and which many tech companies block on upload. www.nytimes.com/2026/09/11/t...

Elon Musk Has Pledged to Rid His Platform, X, of Child Sexual Abuse, but It Persists

Reviews by the Canadian Center for Child Protection and The New York Times found that explicit images of children remain on the social media site owned by Elon Musk.

nytimes.com

All this is important to note because unfortunately, midterms news coverage frenzy clashes directly with H1 product roadmapping season, so take a minute to digest what exactly is changing before rushing to overhaul your roadmaps.

(3) More importantly, this is catch-up regulation, racing to match what litigation outcomes have already set for many companies. In fact, the remedies out of litigation have gone much further, which means this doesn't really change much for companies already tracking and updating their roadmaps.

(2) This also means that any personalization of a feed is now at risk, which is not necessarily a good thing. For example, it's personalization that lets you surface age-appropriate content, or make sure that the product is actually relevant to your interests.

(1) "Addictive features" in this @nytimes.com headline only refers to (a) a feed ranked using data about you, (b) autoplay, (c) anything the AG later decides is "addictive". None of the earlier language in the bill around engagement-optimizing features made it in.