AI Is Rediscovering Management

The Smarter the Machine, the Older the Management Problem

I have this hypothesis that somewhere around Claude Opus 4.6 or 4.7, we have already surpassed “AGI.” Forget the metaphysical debate about superintelligence and Artificial General Intelligence. By “AGI” I mean Average General Intelligence. As of the writing of this blog post, you can already give the machine any knowledge work, and it fully understands the task and most likely does it better than the average new hire could.

The frontier models may continue to get 2x better every three months, but much like computer chips, which keep getting better and faster, to the average person it doesn’t matter, because the average use case doesn’t need the extra power and capability. A computer chip from 10 years ago opens email in a web browser about as fast as the latest and greatest computer chip.

A newer AI model might do the current office job a little better, but I don’t think newer models will make that much of a difference to most businesses anymore, because the current AI already does the job well enough.

The economic use case of AI for most people has already hit kind of an asymptote of capability. The models can already do the job, and while they might get incrementally better, it really doesn’t matter how much better they can do it — it’s more or less already at the functional limit of what the job requires. The use case reaches an asymptote before the model capability does. i.e., the capability continues to go up, but the economic value of the work being done has already plateaued and flattened out. Doing the work any better doesn’t make the company any extra money.

The use case saturates before the model does: economic utility flattens against perfect job performance while model capability keeps rising

The job to be done has a fixed ceiling. It doesn’t need to go above that ceiling (and can’t anyway). As an oversimplification: if the job is to answer 10 math problems correctly, there’s no extra credit for the model to earn. Replying to emails to confirm an upcoming lunch meeting, or creating an internal status-update PowerPoint, has the same problem. For most knowledge work, sufficient intelligence in the models has already arrived.

Over the last several years, the idea was basically: better model = better AI product. But now I think the actual edge for everyday business will come from the ‘harness’ design.

There’s a recent research project from NVIDIA showing how the same model can either fail or pass a task depending on the harness around it.12 What computer scientists are discovering about the agent harness looks suspiciously like what organizations have spent a century discovering about management.

That kind of makes intuitive sense, because if you hire an 18-year-old with no experience and then refuse to give them the tools to do their job — the firm’s Excel templates, access to the secure servers with the historical data, or any instructions on how to do the job — they probably aren’t going to do very well.

Since technical terms still sometimes have differing definitions, the ‘harness,’ as I understand it from the research, is surprisingly all the mundane things that business is about anyway. It’s the same stuff that lets the top investment banks, accounting firms, and consulting firms continually bring up new recruiting classes and regenerate talent every year.

Nobody says Goldman has a sustainable moat because it can hire people with an IQ of 130 while Morgan Stanley is limited to 125.

They compete on their ability to organize capable humans into a high-performing institution.

And once you get the tools and training out of the way, the things that matter are all old management principles: providing clarity to the people (or machines) getting the job done.

  • What does done look like?
  • How do you choose between two great ideas (or, sometimes, the lesser of two evils)?
  • What can the team decide without asking me?
  • What quality level gets us fired by the customer?
  • How do we know when an idea is working and deserves more resources? And, more challengingly, how do you know when it’s time to pull the plug on something that’s delivering some results but not great ones — versus when it’s too early to quit?

The failure points are likely no longer about intelligence, but about classic delegation-design challenges.

Bonus, as a fun aside: I remember studying a case in university where a police candidate was denied an interview on the grounds that he was too intelligent.3 The candidate sued and lost, because the police department was able to show a rational reason that being too smart was bad for the job. This is of course a gross oversimplification, but as I weigh whether Fable 5 gets the assignment or an old standby like Opus 4.6, I couldn’t help but remember that lecture.


Sources


  1. Terry Chen, Yeyin (Eva) Zhu, Zhifan Ye, Jean-François Puget, and Humphrey Shi, “NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents,” NVIDIA Technical Blog, August 21, 2026. In the ARC-AGI-3 evaluation the same base model (Claude Opus 5) scored roughly 30% on its own but 100% inside the AVO harness — though NVIDIA cautions the two configurations aren’t directly comparable. https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/ ↩︎

  2. Terry Chen et al., “AVO: Agentic Variation Operators for Autonomous Evolutionary Search,” arXiv:2603.24517, March 25, 2026. https://arxiv.org/abs/2603.24517 ↩︎

  3. Jordan v. City of New London, 225 F.3d 645 (2d Cir. 2000). The applicant, Robert Jordan, scored a 33 on the Wonderlic Personnel Test (about an IQ of 125) and was screened out because the department only interviewed candidates scoring between 20 and 27, on the theory that overqualified hires would grow bored and leave soon after costly training. The Second Circuit affirmed summary judgment for the city: the policy survived rational-basis review because it was applied to every applicant. ↩︎