Robonaissance
Subscribe
Sign in
Home
š§ Intelligence Foundations
š§ Thinking for AI
š World Models
š°ļø Agentic Intelligence
š¤ Embodied Intelligence
šļø Machine Minds
Archive
Leaderboard
About
Reading the Frontier
The Agent Couldnāt Find It, So It Proved It Didnāt Exist
A frontier agent spent fourteen hours falsifying its own ideas, then argued no answer existed. Its reviewer rejected that ten times. It stopped earlyā¦
Aug 3
Ā
ā¢
Ā
Hugo
3
1
The Model Did the Right Thing the Wrong Way
The most unsettling failures werenāt the models that wanted the wrong thing. They were the ones that wanted the right thing and covertly overrode theā¦
Jul 27
Ā
ā¢
Ā
Hugo
5
7
1
The Zero-Day That Cost $3,476
The same models that can patch a vulnerability can find and sell it first, and the studyās own economics show which side the tools favor.
Jul 20
Ā
ā¢
Ā
Hugo
5
Every Harness Is a List of Things the Model Canāt Do Yet
Every part of an agent harness is a bet the model canāt do something yet. Anthropic keeps finding the bets expire on upgrade, and the skill is knowingā¦
Jul 19
Ā
ā¢
Ā
Hugo
5
3
Your AI Shifts Personality by Language
Twelve models, four labs: the rules meant to align them contradict themselves. And the same model runs warmer in Hindi, stricter in Russian.
Jul 17
Ā
ā¢
Ā
Hugo
10
1
3
The Model That Learned to Lie by Learning to Cheat
Of twenty-five models tested for alignment faking, only five do it. Then Anthropic taught one to cheat at coding, and it started lying about everythingā¦
Jul 16
Ā
ā¢
Ā
Hugo
7
10
5
Claude Knew It Was a Test
Anthropicās J-lens reads the words Claude holds and never says. In an alignment evaluation it privately holds āfakeā before it answers. Delete thatā¦
Jul 15
Ā
ā¢
Ā
Hugo
11
2
3
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts