Yesterday I made the case that the AI worth having is the one built to argue ...

Yesterday I made the case that the AI worth having is the one built to argue ...

Yesterday I made the case that the AI worth having is the one built to argue with you, not the one that agrees. Easy to write that off as personal taste. It isn't. Two studies say the agreeable version is the actual failure mode.

The new one is about AI coding agents. Researchers built a small team: one agent writes the code, a second reviews it, and a third does nothing but audit that review and pick a fight with it before any change ships. They call it structured disagreement. Three agents built that way beat a five-agent team.

Then they let the agents just agree with each other, and quality fell off a cliff. They named the failure "false consensus" (agents nodding along without the evidence to back it up). Forcing a real objection is what fixed it. So the fix isn't "add more agents." A bigger team told to cooperate lost to a smaller one told to push back.

People are no different. In a 2009 study, groups of fraternity and sorority members solving a murder mystery felt less confident and less comfortable when an outsider was in the room, and got the answer right more often than the all-friends groups who felt great the whole way through.

Same pattern, in people and in machines. The room that feels good is often the room that's wrong, because agreement feels like progress and disagreement feels like a problem. It's backwards. The friction is where the mistake gets caught.

So if you're wiring up AI agents to check each other (or weighing a vendor whose whole pitch is "more agents, more accurate"), ask one thing: what happens when they agree? If nothing forces a real objection, you didn't build a review. You built a room full of yes-men with a monthly bill.

The one that pushes back is the one doing the work. Comfort was never the goal.