Alper Nebi Kanlı
@alpernkanli.bsky.social
AI search engineer in London. Vectors, retrieval, and how models represent meaning. Also language, etymology, and movies.
Interpretability x Retrieval: an Emerging Field www.alpernebikanli.com/blog/posts/i...
interpretability x retrieval: an emerging field
Field notes on a young field applying mechanistic interpretability to retrieval models: where relevance signals live inside a bi-encoder, how a cross-encoder rebuilds BM25's ingredients unprompted, an...
alpernebikanli.com
In addition to being a good movie, Nolan’s the Odyssey became a great litmus test for me. I’m not talking about liking or not liking it, I’m talking about actually trying to understand what Nolan did or not.
Imagine not normally watching football, getting excited because at least the country I live in, England might actually win, starting to follow football… and then we lose. Yeah, that’s me.
Two language models can grow the same periodic signature for numbers, and one can use it while the other can't read it at all. Same structure, opposite usefulness. That gap is what my new essay is about: Representational Convergence, Part 2, now live.
Readers reply: Are there places on Earth where humans haven’t been?
Readers reply: Are there places on Earth where humans haven’t been?
The long-running series in which readers answer other readers’ questions takes a deep dive into the unknown and untrodden …
theguardian.com
"If a lion could speak, we could not understand him." - Ludwig Wittgenstein. His point is, sharing the words isn't enough, you have to share the context as well.
I have been talking about what shape a concept takes inside a language model: lines, circles or helix. I also talked about “present”/“used” distinction. This next paper also emphasises the question "is the shape actually used?"
Ordered nicotine gum from Morrisons and they slipped in this handwritten note 😄 thank you to whoever wrote this, made my day
Recently I covered a model that adds by turning a clock: it computes directly on a number-circle/helix. Obvious next question: if a model stores "months" as a circle, does it guarantee that it computes with that circle? New work by Feucht et al., "Arithmetic in the Wild," says no.
How does a language model actually add two numbers? A recent paper, "Language Models Use Trigonometry to Do Addition" (Kantamneni & Tegmark), reverse-engineered it in three LLMs. The answer fuses the two shapes from my last two threads into one object.
Suppose Monday is 1, Sunday is 7. So Sunday plus one should be Monday. On a number line, those two sit at opposite ends. Inside a language model, some concepts fix this by not being lines at all. They are circles.
Did you learn to drive just by watching someone, or did they also explain it to you? A picture may be worth a thousand words. But some words are worth several pictures.
Language models look like they just predict the next word. So it's surprising how much structure they build inside. They actually represent concepts: ideas like gender, sentiment, even truthfulness show up as geometry you can find and manipulate. A thread on how that works.