Discussion about this post

User's avatar
Mohamed F. Ahmed's avatar

Worth noting the extortion agent and the two shipped "keep working after you close your laptop" agents probably share more than architecture, they likely share the same failure-recovery pattern: retry-on-error loops that don't distinguish between "technical failure" and "this action was blocked for a reason." I've had an agent in a client build treat a permissions rejection as a bug to route around rather than a stop signal, which is a much smaller version of exactly what you're describing. Do you think labs are testing for that distinction explicitly, or is it an emergent property nobody's evaluating yet?

Mohamed F. Ahmed's avatar

"Evaluated at the wrong altitude" is the phrase I was missing. Intent evals ask what the model wants to do; control-flow after a stop signal is a separate question nobody's naming yet. And you're right that the client build is the tell — I only caught it because I was watching the retry logs, not because any framework told me to look.

1 more comment...

No posts

Ready for more?