Lana Tikhomirov

@lanatikhomirov.bsky.social

PhD Student at @aimlofficial. Using cognitive psychology to understand and safeguard clinician-AI decision-making 🥼🩻. She/her. AI ethics, human factors, algorithmic safety.

Our main findings from the scoping review are: 1. Most silent trials do not report model metrics outside of AUC (rarely reporting bias testing, failure modes, and data drift) 2. The evaluator of the silent trial is often unnamed or underspecified- human factors and stakeholder engagement is rare

Lana Tikhomirov@lanatikhomirov.bsky.social · 12mo ago

Our preprint is up! Ever heard of the silent phase of AI evaluation for medical AI? Well, now we’ve summarised the current state of research!

Holy moly. I'm trying to write an academic paper, and nearly every application I'm using is not only offering Generative AI as an option for writing, but *pushing it* -- pervading the design to the point where a simple misclick would make my content AI-generated. Here's why that's a problem. 🧵

Screenshot of Google docs generative AI agent asking me to Rephrase, Shorten, Elaborate, More formal.

The fact that Deepseek R1 was released three days /before/ Stargate means these guys stood in front of Trump and said they needed half a trillion dollars while they knew R1 was open source and trained for $5M. Beautiful.

Trump announces 500B in AI funding. Five days ago. Deepseek r1 release. 8 days ago.

We propose that ethical governance for health institutions should be grounded in these local evaluations (not just AI vibes 😎) to ensure that when we say a tool ‘works,’ it means it works for US, for OUR patients, for OUR staff, and we have the evidence to say that.

a woman is standing on a balcony holding a piece of paper and saying i have the receipts .

ALT: a woman is standing on a balcony holding a piece of paper and saying i have the receipts .

media.tenor.com

What does this look like? It will be different for each tool, depending on many things including how much works has been done before on similar tools, how different the local context is, etc. How do we decide? That’s what we want to figure out, and we need your help!

Bild

Fairness evaluations, human factors, cognitive science, patient engagement, environmental considerations (+++ to this one!), economics, and more ➡️ taking this global perspective from the get-go we think will reduce wasteful translation and optimize the benefit!

Many know this already. But we can be doing better with how silent trials are practiced. It’s not just about the model - so we introduce ‘translational trial’ to signal the need for a more holistic evaluation during this silent stage, recognizing that AI is a sociotechnical tool

a man in a suit is standing in a doorway and saying `` what you 're thinking but way more ''

ALT: a man in a suit is standing in a doorway and saying `` what you 're thinking but way more ''

media.tenor.com

"When trying to develop a measure of intelligence, it’s essential to avoid Goodhart’s law: “When a measure becomes a target, it ceases to be a good measure.” As a community of AI researchers, we really need to figure that one out." @melaniemitchell.bsky.social aiguide.substack.com/p/did-openai...

Did OpenAI Just Solve Abstract Reasoning?

OpenAI’s o3 model aces the "Abstraction and Reasoning Corpus" — but what does it mean?

aiguide.substack.com