10 Comments
User's avatar
John Merryman's avatar

That fuzziness in the middle will never go away, because that is where it's all at. Sometimes it's a loop and sometimes it's a cycle.

The issue that keeps bubbling up in my mind is time again. The process goes past to future. The patterns generated go future to past. In QM, the process is the observer/measurement, while the patterns are the wave function being collapsed.

In these models, the feedback between the testing and the model is rather dense, but it is still a water to fish issue. There is no "answer," just dynamic. The fuzziness.

Hugo's avatar

The fuzziness really is the finding here, not a placeholder for one. 1 to 5 percent won't resolve into a clean story, and maybe that's just what's there.

Not sure where future to past comes from in the QM comparison though. Induction heads copy something earlier in context and predict forward, strictly past to future within a pass. Water to fish is exactly right, nobody trained these models to be legible to us.

John Merryman's avatar

That's the feedback between process and patterns. Learning.

I was just reiterating it because you are digging into the foundations of AI and our linear concept of time is the fish. We encounter events and learn from them, as this narrative flow from past to future.

While the only physical reality is this presence, so it is really the events flowing future to past. As the dynamic of this physical state effects change.

For example, lives go birth to death, future to past. Life moves onto the next generation, shedding the old, past to future.

So given that we all tend to see patterns in the information we absorb, in my reading of your history and evolution of computer information processing, I keep seeing the process building up and breaking down models, then using different aspects of them in subsequent models and directions. The new as reconfigurations of what came before. Process and patterns. So thought I'd mention it.

Hugo's avatar

Reconfiguration is the right word for what the interpretability work keeps finding. Induction heads weren't designed, they just keep reappearing, and the same circuits get reused for jobs nobody trained them for.

Whether that's time running backward I'm less sure, but the building up and breaking down part is exactly what the 1 to 5 percent is measuring: how much survives the rebuild.

John Merryman's avatar

It's not that time runs backwards. That linearity is our interpretation of the process, as we have this sequential process of perception, as mobile organisms.

The underlaying dynamic is more circularity and reciprocal feedback.

So those structural nodes in the networks, coalescing and distributing effects, is more of the foundational dynamic.

Hugo's avatar

Circularity rather than reversed time is easier to sit with, and it matches how these things actually get built. Training is iterative feedback, not one pass, so whatever survives is whatever kept getting reinforced through the loop.

Whether that reaches down to something foundational I'll stay agnostic on. But coalescing and distributing effects is a fair description of what a circuit is.

John Merryman's avatar

Philosophy has been stuck in a rut for a long time and it does seem artificial intelligence is going to force the deeper issues and questions to the fore.

Much as all technological revolutions have driven social and cultural evolution, even when the early birds get lost in the shuffle, lose their shirts.