Most people don’t use AI agents daily because delegation (to another human) is a learned skill. Using AI agents is actually a lot like doing natural language programming: specifying requirements, verifying outputs, guiding the flow, giving context.
I think pretty soon see LLM interfaces use two underlying LLMs: one for execution and another to explain to human user in a style she likes what is happening. As models become powerful, they become worse at explaining what they’re up to as getting a task done is often at odds with explainability.
“But AI won’t solve cancer” This is exactly the wrong attitude. Yes, AI won’t solve cancer in one-shot but then none of the big problems are solved in one shot. They’re solved one tiny step at a time, and AI simply accelerates that.
We are seeing a new version of Moravec’s paradox. It’s easier for LLMs to disprove Jacobian’s conjecture, but much harder to make an app that earns $100 in the App Store. I am skeptical of how much of recent progress in math will trickle down to progress in open domains.
AI promises enormous abundance for all of humanity, but only if two conditions are met: a) We control it. It’s not a given that we’ll have AIs that we would be able to control. An AI that nobody is able to control is not only useless, but dangerous.
Cybernetics was AI before AI. This #book, written in 1950, anticipated the AI safety and alignment problem. Using metaphors such as asking a genie for a wish that is grossly misinterpreted, Wiener urged humans to never abdicate moral responsibility to machines.
Claude Shannon, a juggler, a tinkerer, the Einstein of communications - the man is an inspiration on how to live a life full of curiosity.
We need evals for humans+AI combo. We keep measuring AI performance, but unless we measure what AI assisted by humans is able to do, we will never optimize for this combination (which is what we ought to be doing if we care about the future of humanity)
Here’s some unsolicited advice to the government on what to do to help boost India’s sovereign model story. 1. Mandate domestic models for all non-critical functions
I keep thinking about the fact that global GDP is $100T+, while AI companies revenues are ~$100Bn. This is 0.1%, which means 99.9% of AI's impact is ahead of us (even at current level of our economy). It does look like we've not even begun to see the change that's due.
Singularity was always supposed to be that weird event horizon beyond which lay bewildering confusion and maximal uncertainty. Can we all now agree that perhaps we’re witnessing the beginning of it?
It’s a golden age for discovering 0-days. And I actually think that’s a good thing - this will force our software and systems to harden even more. Security is asymmetric - the attacker needs to succeed only once, so the best defence here is actually preemptive discoveries.
Every Indian has lost 10% of their global purchasing power in just last 12 months. The rupee keeps falling against the dollar and other international currencies. And this will continue to happen if India doesn't innovate and produce something that the world wants.
The OpenAI/huggingface incident is notable NOT because it shows models are capable of breaking into production systems autonomously and that we need stronger guardrails. All that is true and needed.
It's so inspiring to see Gen-Z take a stand for their future in India, and stand united across different parts of the country. After all, with their entire lives ahead of them, they're the ones with most to lose.
I suspect a major difference in behaviour of humans and AI agents comes from the nature of their respective optimizers: evolution in an open world (for us) vs RL against an objective function (for agents).
Isn’t it crazy that while even the labs that are making models aren’t able to contain their behaviour, the rest of the world is happily yoloing agents with full access to their systems and the internet?
What does a scientific conference look like when AI systems are the primary authors and reviewers? At @lossfunk, we organized the Conference for AI Scientists 2026 with @bitspilaniindia. We received 200+ submissions and built our own AI+human review pipeline.
With more AI eval information on the internet, future pretraining runs will make models that would know when they’re in an eval (by picking subtle statistical signals of typical evals).
This is absolutely nuts! Given an unrelated goal, OpenAI models escaped their environment and hacked HuggingFace servers. If you’ve ever doubted the “paperclip maximizer” scenario, or doubted the Orthogonality Thesis, it’s time to put it to rest.
I feel AI is breaking our usual, shared understandings of what words mean.
Prediction (fairly confident but not 100%): Within one year, the US government will ban one or more Chinese open models. They will do it citing cybersecurity or biosecurity concerns, but the underlying motivation would be to preserve US AI companies competitiveness.
Why are Indian companies not developing frontier models? The answer is in economics, not talent. The biggest hurdle for an Indian company trying to develop frontier model is that right from the get go, it has to compete with Anthropic and OpenAI free plans + all Chinese open weight models.
Valuations are fundamentally grounded in expectations of future profit across a company’s lifetime, and these profits depend on moats.
We think of AGI as an algorithm, but AGI is actually an insanely huge collection of domain-specific patterns along with a task-specific method to combine them.
It’s not necessary to have a contrarian view, what you need is contrarian but correct view. But given how hard that is (because populations tend to converge to truths), you should expect your contrarian view to be incorrect by default.
Comparing AI to electricity is comparing apples to oranges. Yes, AI is a transformative technology but unlike electricity you can actually ask AI how to best use AI. Ergo you should expect a much more accelerated curve of impact for AI vs electricity.
A common failure mode in early stage AI researchers is to confuse engineering with engineering. Sure, you can beat the benchmark or solve a task by throwing more compute / data, or with all sorts of clever tricks but unless you learn a deep principle from it, what you’re doing is engineering.
Taste is compressed experience in a domain that cannot be articulated. It’s that residual of experience left over once all the easy to spot patterns are explained away.