I haven't read all the articles in the series, just this one, so maybe it is answered in a later one.
I wonder, how well can current AI recognize intent?
When a person sees another person lift a cup to drink from it, it is usually easy for us to determine what they intend to do before they do it. Are AI systems able to do something like that?
Good question. The cup case is actually the easier version of this, and AI does okay at it. Video prediction and world models can do short-horizon "what happens next" pretty well, though it falls apart fast as the horizon stretches.
What I find interesting is why the cup is easy. Physics is doing most of the work. A hand rising toward a face with a cup in it has maybe two possible endings, and one is overwhelmingly likely. You're not really reading a mind, you're reading a situation that has already narrowed itself down. Language gives you none of that. "Book me a flight to Boston" is compatible with a hundred different underlying goals, and reading the sentence harder doesn't help, because what would settle it was never in there.
Part 7 gets into the physical side, react versus anticipate in embodied models, and funnily enough it lands on almost the same cup example. But I'm curious what you think: is the cup thing intent reading, or is it just running physics forward and calling the result intent?
I think if a person raises a cup it involves different brain activity than if a cup is falling, specifically mirror neurons. There are some subtle differences when a waiter clears a table and lifts a cup, and we seem to be able to judge what a person's goal is very effectively, although the movements can initially be very similar.
I can't see how LLMs would do something like this, but as I don't know much about robotics and embodied models, I was wondering if AI systems used for robots are using something similar to or inspired by mirror neurons.
You're onto something real, and the neuroscience backs you up. Fogassi's 2005 work found mirror neurons that fire differently for grasp-to-eat versus grasp-to-place, the same opening movement, told apart by where it's headed. That's your waiter example at the single-neuron level.
But the honest answer to your question: mainstream robot learning is mostly not mirror-neuron-inspired. Imitation learning does teach robots from watching demonstrations, which sounds mirror-like, but under the hood it usually just maps situations to actions, "in this state, do this," and skips the intention layer your brain does for free. There's a smaller research thread that takes the analogy seriously, building explicit "mirror systems" and goal-inference models, but it's niche, and none of it approaches that effortless grasp-to-eat/grasp-to-place read you're describing.
And it loops back to the series. Even reading intention perfectly is still reading someone else's, translation, not origination. Worth noting it's contested too: some researchers argue mirror neurons identify actions but not intentions, and that intention-reading leans on separate "mentalizing" regions. If so, what you're describing may be even further from what robots do than it looks.
To the man with a hammer, everything looks like a nail. Most of his friends work with hammers. He reads thinkpieces about the impact of hammers and even attends hammer conferences.
The field of hammers is developing rapidly. There are now agentic hammers that autonomously bang anything nail-shaped. The remaining limitations of hammers, he reads, are almost solved, and the General Hammer is imminent.
Sometimes, on a clear quiet night, when the breeze lands just right, he allows himself to wonder. Is there something missing? Why all the hammering in the first place? Could it be that the dominance of the "hammer gaze" is why the world itself seems to require ever-more hammers? What is the true cost?
He dimly recalls having once heard of a form of carpentry that doesn't need nails at all, let alone hammers. But none of his friends seem to have such doubts, or if they do, they hide it as well as he does.
Besides, hammers are where the money is. And there are many very smart and well-paid people working on hammers. Perhaps we should just leave the "why" to the philosophers.
I haven't read all the articles in the series, just this one, so maybe it is answered in a later one.
I wonder, how well can current AI recognize intent?
When a person sees another person lift a cup to drink from it, it is usually easy for us to determine what they intend to do before they do it. Are AI systems able to do something like that?
Good question. The cup case is actually the easier version of this, and AI does okay at it. Video prediction and world models can do short-horizon "what happens next" pretty well, though it falls apart fast as the horizon stretches.
What I find interesting is why the cup is easy. Physics is doing most of the work. A hand rising toward a face with a cup in it has maybe two possible endings, and one is overwhelmingly likely. You're not really reading a mind, you're reading a situation that has already narrowed itself down. Language gives you none of that. "Book me a flight to Boston" is compatible with a hundred different underlying goals, and reading the sentence harder doesn't help, because what would settle it was never in there.
Part 7 gets into the physical side, react versus anticipate in embodied models, and funnily enough it lands on almost the same cup example. But I'm curious what you think: is the cup thing intent reading, or is it just running physics forward and calling the result intent?
I think if a person raises a cup it involves different brain activity than if a cup is falling, specifically mirror neurons. There are some subtle differences when a waiter clears a table and lifts a cup, and we seem to be able to judge what a person's goal is very effectively, although the movements can initially be very similar.
I can't see how LLMs would do something like this, but as I don't know much about robotics and embodied models, I was wondering if AI systems used for robots are using something similar to or inspired by mirror neurons.
You're onto something real, and the neuroscience backs you up. Fogassi's 2005 work found mirror neurons that fire differently for grasp-to-eat versus grasp-to-place, the same opening movement, told apart by where it's headed. That's your waiter example at the single-neuron level.
But the honest answer to your question: mainstream robot learning is mostly not mirror-neuron-inspired. Imitation learning does teach robots from watching demonstrations, which sounds mirror-like, but under the hood it usually just maps situations to actions, "in this state, do this," and skips the intention layer your brain does for free. There's a smaller research thread that takes the analogy seriously, building explicit "mirror systems" and goal-inference models, but it's niche, and none of it approaches that effortless grasp-to-eat/grasp-to-place read you're describing.
And it loops back to the series. Even reading intention perfectly is still reading someone else's, translation, not origination. Worth noting it's contested too: some researchers argue mirror neurons identify actions but not intentions, and that intention-reading leans on separate "mentalizing" regions. If so, what you're describing may be even further from what robots do than it looks.
To the man with a hammer, everything looks like a nail. Most of his friends work with hammers. He reads thinkpieces about the impact of hammers and even attends hammer conferences.
The field of hammers is developing rapidly. There are now agentic hammers that autonomously bang anything nail-shaped. The remaining limitations of hammers, he reads, are almost solved, and the General Hammer is imminent.
Sometimes, on a clear quiet night, when the breeze lands just right, he allows himself to wonder. Is there something missing? Why all the hammering in the first place? Could it be that the dominance of the "hammer gaze" is why the world itself seems to require ever-more hammers? What is the true cost?
He dimly recalls having once heard of a form of carpentry that doesn't need nails at all, let alone hammers. But none of his friends seem to have such doubts, or if they do, they hide it as well as he does.
Besides, hammers are where the money is. And there are many very smart and well-paid people working on hammers. Perhaps we should just leave the "why" to the philosophers.