Burny

@burnytech.bsky.social

On the quest to understand the fundamental mathematics of intelligence and of the universe with curiosity. http://burnyverse.com Upskilling @StanfordOnline

Claude formalized proof of Fermat's Last Theorem! Nice! Next step is formalizing full Weil conjectures proofs and Inter-universal Teichmüller theory!

BildBild

These models are also deceptively elegant in important ways. Their complexity largely reflects the complexity of their environment. The fact that they can pull this off with “just a bunch of matrix multiplications” is precisely what makes them so relevant for human neuroscience.

Finally, we argue that the apparent complexity of LLMs should not discourage us from trying to understand them. They are not black boxes! They are fully transparent—and we hope the explosion of interpretability work continues to cross-pollinate with computational neuroscience.

In LLMs, syntax and semantics are processed through the same neural mechanisms and both are encoded geometrically in a neural population code. If the brain follows a similar strategy, we should not be surprised that syntax/semantics are co-localized and functionally entangled.

For example, years of neuroimaging work tried to disentangle the cortical bases of syntax and semantics based on their separability according to linguistic theory—but more recent work indicates that syntax and semantics are highly coextensive throughout the language network. Why?

To the LLM, these are all just patterns in context—it will learn them insofar as they are *useful* for producing natural language. This context-first approach to language renders different types of linguistic structure learnable by a simple, general-purpose learning algorithm.

Certain patterns are highly regular; these are the kinds of patterns that are more easily detected and have been richly described by linguists. Other patterns are more abstract, graded, or multidimensional, and therefore harder to describe.

This dense, geometric representational format also yields a surprisingly powerful form of generalization: LLMs can interpret (and produce) novel linguistic contexts by interpolating the meaning of new locations in this geometric landscape (within the bounds of their prior experience).

Principle 1 has important implications: for example, it helps us understand how many seemingly different structures of language (e.g., syntax trees, semantic categories) can be unified in a geometric lingua franca that supports neural computation.

Principle 2: They replace rule-based learning with a generic, self-supervised, context-driven learning algorithm for reproducing the statistical structure of real-world language.

Bild

Principle 1: They encode all of the structures of natural language (syntax, semantics, pragmatics, etc.) in the geometry of a continuous, high-dimensional embedding space—i.e., a neural population code.

Bild

One problem is simple: brain activity and LLM states are both responses to the same structured linguistic input. Alignment can therefore arise from shared stimulus structure without implying shared internal computations.[3/8]

Flow Reasoning Models. They developed a recurrent flow-based architecture to efficiently solve structured reasoning problems (e.g., Sudoku). It applies continuous flows to discrete data and recurrently refine their past mistakes through self-conditioning.

AAII gave Astra the same score as GPT-5.6 Sol. It's pretty clear somebody has messed up something. Such as testing the wrong model for whatever reason, having a whole lot of API errors, or using some minimal reasoning effort (they say it used only 42M tokens with max reasoning).

https://artificialanalysis.ai/

An example of useful knowledge work: I assigned GPT-6 to read through tens of thousands of my emails, my writings, my calendar appointments and more to assemble a personal knowledge base of research, contacts, ideas, relationships, and tasks over my recent career.1/

Bild

we really need a way to talk about entities that are narrative-shaped (text elementals, jpeg-artifacted reasoning-capable shards of the akashic library, whatever) that acknowledges how weird they are without papering that over with the implication that they're just little guys inside the computer