7 Comments
User's avatar
Scott H.'s avatar

A very timely series, thank you for writing about these ideas. My curiosity goes quickly to multi-variable scoring, I I hope you will touch on this and bring some clarity to the complexity of it.

Hugo's avatar

Yeah, this is the question that everyone reaches for as soon as the single-score version starts to bite, and rightly so. The honest short answer is that going multi-variable helps and also doesn't, in a specific way worth being clear about. Multiple objectives buy you coverage, you're no longer blind to everything the one metric ignored. But the moment you have several, you need a rule for trading them off against each other, and that rule is itself a specification, with all the same problems one layer up. Weight the variables and the weights are now the thing being gamed. Leave them unweighted and the system quietly picks its own trade-offs in the gaps you didn't specify.

So it doesn't dissolve the problem, it relocates it, from "what's the target" to "how do these targets combine." Sometimes that's a real gain, because a trade-off rule can be easier to reason about than a single proxy. Sometimes it just hides the same gap behind more machinery. The interesting cases are the ones where the variables genuinely conflict and there's no principled exchange rate between them, which is where it stops being a scoring problem and turns into a values problem.

Filing your interest away though, because you're right that it deserves proper unpacking. Where have you seen multi-variable scoring actually hold up, versus collapse back into one dominant number that eats the rest? That failure mode seems almost gravitational and I'm curious if you've watched anything resist it.

Scott H.'s avatar

I've bumped up against the failure mode twice from from different approaches landing at the same place. The upside is the map continues to narrow the places to look, when I have something I'll pass it to you.

Damaris Kroeber's avatar

I wonder a lot about the management thing. Obviously having “middle layer” humans simply for dealing with the complexity of large organizations that outpace human cognitive capacity seems hard to score. Like, is that a score at all? Or are we looking at some type of “coherence of data across organizational scale” that would become the objective? You mentioned in a previous piece that the machines carry a vector for “coherence” that they can use to make better output. It would probably still need a lot of steering (and intention giving) from top management to make it cohere in the “right direction” but when it comes to the middle layers that exist purely for complexity overwhelm reasons, might one use that coherence vector?

Hugo's avatar

Yeah, "is that a score at all" is the whole thing right there. So much of middle management exists just because the org got too big for any one person to hold in their head, and that layer isn't chasing a number, it's holding information together across a scale nobody can actually see all at once. That's the part my third question keeps snagging on. Even if a machine could hold all that state more consistently than the humans do, somebody still has to say what counts as cohering in the right direction, and that's not really a coherence problem, it's a specification problem wearing a coherence costume. The machine can keep more plates spinning. It just can't tell you which ones you actually wanted up there.

So maybe the layer just splits in two? The pure keeping-track-of-complexity part feels very scoreable, very automatable. But the part that's quietly deciding which way things should cohere was never a complexity problem in the first place.

Damaris Kroeber's avatar

Yeah I was also thinking that that’s why you might always need a founder/some top level management personnel who give the whole endeavor a direction or goal, a reason for existence. And given that, you think one could develop a coherence score for the entire rest of the company?

Hugo's avatar

Yeah, you could build one, though the trap is that the moment coherence becomes measurable it becomes optimizable, and people start producing coherence rather than being coherent, learning the vocabulary, citing the strategy deck, shipping things that pattern-match the direction. Score climbs beautifully, everyone's still pointing different ways.

But I think you could build one that survives that, it just has to be measured somewhere nobody's incentivized. Two things that seem to work: ask people to describe the direction in their own words rather than scoring their alignment to it, because a team that's actually coherent will paraphrase it consistently and a team that's performing will quote it verbatim. And track one thing nobody's rewarded for, like how often teams kill their own projects on strategy grounds, since that's expensive and nobody fakes it upward. Plus the founder's direction going back on the table on a schedule, because a company can be perfectly coherent around a goal that stopped being right three years ago.