Becca Cohen

@beccacohen.bsky.social

IS PhD Student at UIUC studying digital humanities, language, cultural analytics and ethical AI

OpenAI posted the terms of the deal. Reveals that it absolutely does allow for domestic surveillance. EO 12333 is how the NSA hides its domestic surveillance by capturing communications by tapping into lines *outside the US* even if it contains info from/on US persons. openai.com/index/our-ag...

2. Our contract. Here is the relevant language:

The Department of War may use the AI System for all lawful purposes, consistent with applicable law, operational requirements, and well-established safety and oversight protocols. The AI System will not be used to independently direct autonomous weapons in any case where law, regulation, or Department policy requires human control, nor will it be used to assume other high-stakes decisions that require approval by a human decisionmaker under the same authorities. Per DoD Directive 3000.09 (dtd 25 January 2023), any use of AI in autonomous and semi-autonomous systems must undergo rigorous verification, validation, and testing to ensure they perform as intended in realistic environments before deployment.

For intelligence activities, any handling of private information will comply with the Fourth Amendment, the National Security Act of 1947 and the Foreign Intelligence and Surveillance Act of 1978, Executive Order 12333, and applicable DoD directives requiring a defined foreign intelligence purpose. The AI System shall not be used for unconstrained monitoring of U.S. persons’ private information as consistent with these authorities. The system shall also not be used for domestic law-enforcement activities except as permitted by the Posse Comitatus Act and other applicable law.
Mike Masnick@masnick.com · 5mo ago

I saw some folks asking what the difference was between what OpenAI signed with the DoD and what Anthropic said they wanted, and Sam more or less admits here the key point: OpenAI's deal requires them to trust the NSA. Anthropic's contract had real safeguards.

User Chris: What was the core difference why you think the DoW accepted OpenAI but not Anthropic

Sam Altman: 
I can't speak for them, but to speculate with the best understanding of the situation.

*First, I saw reporting that they were extremely close on a deal, and for much of the time both sides really wanted to reach one. I have seen what happens in tense negotiations when things get stressed and deteriorate super fast, and I could believe that was a large part of what happened here.

*We believe in a layered approach to safety--building a safety stack, deploying FDEs and having our safety and alignment researcher involved, deploying via cloud, working directly with the DoW. Anthropic seemed more focused on specific prohibitions in the contract, rather than citing applicable laws, which we felt comfortable with. We feel that it it's very important to build safe system, and although documents are also important, I'd clearly rather rely on technical safeguards if I only had to pick one.

*We and the DoW got comfortable with the contractual language, but I can understand other people would have a different opinion here.

*I think Anthropic may have wanted more operational control than we did

Making a CLEAN, SHAREABLE dataset is fucking hard! I'm super proud, then, to publish this one, on a team led by @sdileonardi.bsky.social and @beccacohen.bsky.social, with @post45data.bsky.social. It has more than a decade of 21C int'l bestseller data, revealing how popular world lit circulates....

Post45 Data Collective@post45data.bsky.social · last yr.

New dataset on bestsellers from 40+ countries, with consistent coverage for France, Germany, Spain, Italy, and the U.S. Congrats to the authors @sdileonardi.bsky.social, @beccacohen.bsky.social, and @dan-sinnamon.bsky.social on this major contribution! 🎉 🔗: doi.org/10.18737/386...

Hoo boy! I can’t believe it’s happening. We’ve been working on this for years. Thanks to @post45data.bsky.social you can now search our IB database, with a very cool interface! Please share with anyone who might be interested Special thanks to @ninasabak.bsky.social for aiding with the data source

Post45 Data Collective@post45data.bsky.social · last yr.

New dataset on bestsellers from 40+ countries, with consistent coverage for France, Germany, Spain, Italy, and the U.S. Congrats to the authors @sdileonardi.bsky.social, @beccacohen.bsky.social, and @dan-sinnamon.bsky.social on this major contribution! 🎉 🔗: doi.org/10.18737/386...

I’m going to need Mother Nature to knock it off with all these tornadoes. How am I supposed to write my dissertation while I’m stressing out about corralling my cats into the basement?

I need to hear this right now, which means that others probably do too: However you're getting through this, you're doing a good job. This month's been fucked up and intense and I know I'm not the only one who's overwhelmed. It's exhausting and cruel, and the important part is getting through 🩵

If Stephen King is a brand name, what’s the diff bw him and more recent authors who embrace self-branding? It was more difficult to answer this than I first thought It was a pleasure getting here with my coauthors (& the ed.s & Post45 2023 crew). Please read and let us know what you think

Dan Sinykin@dan-sinnamon.bsky.social · 2y ago

With @sdileonardi.bsky.social and @beccacohen.bsky.social, I wrote a shortish essay on The Girl on the Train and how self-branding is changing authorship. For a fascinating special issue of Studies in the Novel ed by @megaplex.bsky.social and @sarahdallison.bsky.social muse.jhu.edu/pub/1/articl...

Maybe one of the reasons there are so many unhinged academics is that in order to keep your job, you’re constantly required to tout yourself as a Scholar of World-Historical Significance, even as your office’s ceiling tiles slowly dislodge themselves to fall on your head.

There are many ways to identify texts that seem ahead of their time. Our CHR 2024 paper asks which measures of textual precocity align best with social evidence about influence and change.

Abstract: Measures of textual similarity and divergence are increasingly used to study cultural change. But which measures align, in practice, with social evidence about change? We apply three different representations of text (topic models, document embeddings, and word-level perplexity) to three different corpora (literary studies, economics, and fiction). In every case, works by highly-cited authors and younger authors are textually ahead of the curve. We don't find clear evidence that one representation of text is to be preferred over the others. But alignment with social evidence is strongest when texts are represented through the top quartile of passages, suggesting that a text's impact may depend more on its most forward-looking moments than on sustaining a high level of innovation throughout.

The Gemini debacle showed how AI ethics *wasn't* being applied with the nuanced expertise necessary. It demonstrates the need for people who are great at creating roadmaps given foreseeable use. I wasn't there to help, nor were many of the ethics-minded ppl I know.

Lukasz Olejnik@lukaszolejnik.bsky.social · 2y ago

My comments for Telegraph about Google Gemini hiccup: generation of weird, falsified images of human history. When AI Ethics and risk-assessment goes really bad. I’m concerned also as a person with a disability. www.telegraph.co.uk/news/2024/02...