Ian Sample

@iansample.bsky.social

Guardian Science Editor and co-host at Science Weekly. Author of Massive | Sony Gold Award | PhD in Biomedical Materials https://www.theguardian.com/profile/iansample

Great article from @iansample.bsky.social. Starts with info on the significant harms to oral health from smoking and moves on to give nuanced assessment of the more limited evidence on vaping. If all coverage was like this then a majority of Brits would not think that vaping was as bad as smoking.

Action on Smoking Health (UK)@ashorguk.bsky.social · 10mo ago

There's been a few articles this week about vaping and oral health. This excellent article sets out the facts and crucially also highlights the extreme damage that smoking does to oral health. www.theguardian.com/society/2025...

I have stacks of questions about how the model’s working: how is the grammar intact, why are certain words repeated, are any lowest probability? But I like it. AI threatens to reduce human experience by steering our choices to popular ones. This is feeble, silly pushback. Or token subversion? 🤖🧠🧪

Using least probable tokens, the answer was a hoot: “A great city emerges through kaleidoscope sandwiches of perpetual thunder. It integrates gelatinous traffic systems of whispering algorithms. It preserves holographic fountains of magnetic jam and encourages staircases of wandering moons.” 🤖🧠🧪

The question I posed was: What features make a great city? The normal (most probable, single token) response highlighted the richness of a city’s culture, the diversity of its people and the strength of its public spaces. Fair enough. But a bit dull.

I asked ChatGPT to emulate Least Probable Token Selection across single, double and triplet tokens. Instead of building sentences from the most probable next token(s), it chooses them from the long tail of lowest probabilities. The results are fun. 🤖🧠🧪

Accordion staircases of wandering moons AI chatbots churn out answers by repeatedly predicting the next token, be that a word, subword or character. Building sentences from highly probable next tokens isn’t a bad way to extract consensus from language. But it's boring. So I had a word. 🤖🧠🧪