Grigori Guitchounts

@guitchounts.bsky.social

AI + biology, venture creation @FlagshipPioneer | Neuroscientist @Harvard | Writer, pianist, runner, painter, dilettante

I wrote an essay for @noemamag.com about AI consciousness, animal research, and the limits of what we can know about another mind. My argument: the lessons neuroscientists learned hunting for neural markers of consciousness in animals are exactly the ones we'll need as we start asking...

Noema Magazine@noemamag.com · 2mo ago

“The science of consciousness & the ethics of moral consideration — what we have learned from rats & monkeys & humans — will have to guide us as we meet artificial minds,” @guitchounts.bsky.social writes. #ai #aiethics #consciousness

A few years ago @jesseba.bsky.social and I started experimenting on large language models the way a neuroscientist probes a brain. We’re trained as neuroscientists, but converged on the problem from a few different directions: dynamical systems, graph theory, stat phys.

Jesseba Fernando@jesseba.bsky.social · 2mo ago

Check out the first in a 3-part series of posts explaining @guitchounts.bsky.social and I's latest paper exploring dynamics of transformers! Watching a transformer spin open.substack.com/pub/oftwomin...

Title of article with a background depicting a spiral

A lot of biology’s “intelligence” shows up before anything like a brain: cells negotiating, wounds closing, embryos turning a noisy soup of parts into a reliable body. This paper argues that if we want robust machine intelligence, we should copy more of that— not just neurons.

If you ask a general LLM to “summarize the recent literature” and tack on citations, it’ll often confidently hand you references that… don’t exist. This paper actually measures it: for recent, cross-field queries, GPT-4o’s citations were fabricated ~78–90% of the time.

A lot of “AI for science” eval still sounds like a school exam: answer the question, show your work. Real research is messier—dig up the right papers, read figures in context, debug a protocol that’s failing, pull sequences, plan experiments, defend the choices.

A lot of biology is still “glue work.” Ten tabs open, CSVs shuttled around like buckets in a leak, half-manual analyses, scripts you swear you’ll clean up later, tables reformatted… again. The claim here: an “integrated biology environment” + agents becomes the default interface.

Touch feels like the missing sense in a lot of “dexterous” robotics. You can watch contact, sure—but you can’t reliably feel slip, pressure spreading across a fingertip, shear, or tiny surface texture. That’s how human hands modulate force without crushing or dropping.

I keep coming back to this: one protein sequence doesn’t cash out to one behavior. Even “the same” molecule can hop between shapes, with different timing. And the weird, rare states—the ones you’d bet against—can still matter.

Your blood is full of cell-free DNA—millions of tiny shards, like confetti after a rough party. An Alzheimer’s classifier trained on them ends up leaning hard on a blunt signal: fragment length patterns, not just sequence or methylation calls.

People say “emergence” in LLMs like it’s a magic trick: nothing… nothing… then—poof—capability. In complexity science it’s stricter. Emergence is when you can describe the system in a new, lower-dimensional way that makes the messy micro-details irrelevant.

KG retrieval has an irritating failure mode: either you cast a wide net (nice coverage, but it’s all a bit mushy) or you commit to edge-walking (great multi-hop… unless you picked the wrong starting node and everything collapses). Real queries usually want both, in one pass.

LLMs can do a decent impression of almost anyone… until they can’t, and you feel the rubber band snap back to “Helpful Assistant.” This paper tries to locate that snap-back in the model’s activations—and finds what looks like a single direction for “Assistant-ness.”

Hypotheses are getting cheap. Lab time isn’t. A lot of “AI for science” feels like moving the traffic jam: from dreaming up ideas to the grimy work of checking them—what you test, how quickly, and what you do when the first run faceplants.

Consciousness theories have a very dull way of failing. Either they can’t be falsified, or they end up “trivial”—basically restating whatever our test already measures (report, behavior). Hoel argues a lot of popular theories land on one of those horns.

One coding agent can be perfectly competent and still be the wrong “unit of work” for a big project. It moves like a single person trying to renovate a house alone: slow, snag-prone, and it forgets what it was doing. So: can you run many agents without summoning chaos? Not really—not yet, at least.

A lot of “medical AI regulation” in the US isn’t one big rulebook. It’s happening hospital-by-hospital: local committees kick the tires, watch the numbers, and decide if a tool is safe enough to touch patients. Flexible, yes. Also uneven.

Harder reasoning problems make humans slow down. What caught my attention here: a “large reasoning model” seems to slow down the same way—when it burns more tokens on chain-of-thought, humans also take longer on that exact item.

Agent failures usually don’t look like “the answer is wrong.” They look like a domino run: a tool call slips, state gets a little haunted, and the agent walks to the finish line anyway, smiling. So you test the trajectory *and* the final world state.

LLMs got better in 2025 less by bulking up, and more by being pushed longer through reward loops you can actually score. That changes what “progress” feels like: more optimization, more test-time “thinking,” and a little more jaggedness at the edges.

Chemistry and biology models live in different “languages”—SMILES strings, graphs, 3D coordinates, protein sequences. On paper they don’t rhyme. What caught my attention: they might still compress “matter” into similar latent features. This paper tries to test that.

Hyper-Connections widen the residual stream in transformers, boosting accuracy without extra FLOPs. But at scale the identity path disappears, gradients blow up or vanish, and memory traffic jumps. DeepSeek’s new paper asks: can we keep the new width and stay stable?

New PNAS Nexus study presses a basic question: can large language models reason by analogy, or just regurgitate patterns? The authors test it with counterfactual tasks—changing one fact and seeing if the model adjusts the rest.

Getting to the new release from Arc: Stack, a foundation model for single-cell RNA-seq. Feed it drug-treated immune cells as a prompt; it guesses how, say, lung epithelium would respond—even if that combo never showed up in training.

Residual nets succeed because the identity shortcut keeps gradients alive, but that same identity forces every layer to be a pure addition. Deep Delta Learning replaces the fixed shortcut with a learnable, data-dependent rank-1 transform.

Really cool idea: a single night wired up for polysomnography (EEG/EOG, ECG, EMG, breathing belts—the whole spaghetti situation) contains enough structure to say something about future disease risk. That’s really intriguing

Really cool idea: a single night wired up for polysomnography (EEG/EOG, ECG, EMG, breathing belts—the whole spaghetti situation) contains enough structure to say something about future disease risk. That’s really intriguing