‘You just hired a million bad employees’: How the brief tokenmaxxing era delivered the opposite of what it promised

0
1

George Sivulka, the 27-year-old Stanford grad who founded and runs the AI enterprise startup Hebbia, has put his finger on the defining anxiety of the AI agent era: companies raced to deploy AI workforces without building any of the management infrastructure to run them.

Hebbia already serves clients including BlackRock, KKR, and the U.S. Air Force, giving Sivulka a front-row seat to how badly this is going inside real enterprises. His essay, published on a16z’s newsletter, argues that AI didn’t cut labor costs—it inverted the equation entirely: “for the first time in history, humans are cheaper than software.” And for all those CEOs who rushed headfirst into the agent era briefly known as “tokenmaxxing,” he offered a warning: “you just hired a million bad employees.”

The era ended when Amazon famously disclosed a $500 million loss in one month alone as agents ran wild to little effect, while Ford Motor Company actually hired back a new force of human engineer “graybeards” to work hand in hand with AI augmentation efforts. To Sivulka, the moment is like a mostly forgotten railroad crash in 1841 that ended one era, and began another.

The railroad crash analogy

Sivulka reached back to the 1830s and 1840s, when American railroad track mileage exploded roughly 120-fold in a decade with no coordination systems to match the growth, until a fatal train collision in Massachusetts in 1841 forced the industry to invent modern management—defined roles, reporting lines, hierarchies.

That crisis, he argued, is what turned railroads into one of the first billion-dollar industries, “at its peak representing roughly 60% of the stock market.” Just as railroads unleashed the appetite to travel across the country, he argued that agents have done something similar to work on the web: “We just gave every employee, even the worst ones, effectively unlimited headcount and budget. Managing AI is harder than managing people, because AI scales dysfunction instantly.”

Why every company “hired a million bad employees”

The essay’s title captures Sivulka’s core diagnosis: AI agents don’t fail because the models are weak—they fail because almost nobody in a company can articulate a task clearly enough for an agent to execute it well. He estimates just “1 in 100 employees knows how to give AI context,” calling that skill “a rare breed” of clear thinking that most workers simply don’t have. The result is what he calls “looping”—agents calling themselves over and over to self-correct for bad instructions, which he bluntly describes as “spending tokens on spending tokens.”

UBS Global Research hosted “many of the highest-profile AI-native firms” at its 5th Annual UBS Private AI and Software event in Menlo Park at nearly the same time, and nearly every executive privately confirmed Sivulka’s thesis, of agentic failure happening at industrial scale. One firm executive told UBS’ analysts: “Internally, we don’t have token budgets, it’s not something that we, our engineers, have been trained to think about. But every single customer conversation is about this.”

The real cost wasn’t the tokens

Sivulka argues the industry misdiagnosed its own hype cycle: “tokenmaxxing” spending exploded and then collapsed within a month, but “the amount of tokens spent was never the real problem” — the problem was that people didn’t know how to use them efficiently.

UBS’ sourcing puts real numbers behind that claim: one unnamed AI firm disclosed, “our spend on Anthropic was $20k in December and we’re about to cross $1m in July, a 50x increase in 7 months,” adding that despite the surge, “we’re not throttling back, we don’t want people to stop using it.” That firm is nonetheless installing what amounts to Sivulka’s missing management layer after the fact: “we’re now alerting if you hit a certain threshold on a monthly basis, we’re going to start rolling out governors internally, like G&A staff should not be using the frontier models.”

Public statements from OpenAI (“AI costs have now become a huge issue that never came up at the start of the year“) and enterprises like Uber installing spend guardrails confirm this isn’t isolated—UBS estimated last month that token-cost anxiety had become “a real concern for ~60% of organizations,” and “that figure now feels higher.” Sivulka’s point directly echoes Palantir CEO Alex Karp’s public complaints that AI labs have “completely, irresponsibly, oversold” their models while enterprises burn money on token consumption without real ROI discipline. “Something has gone completely wrong,” Karp told CNBC’s Squawk Box earlier this month as he vented his spleen over misguided token usage. “The basic view among enterprises in this country is I’m going to chillax and waste my time with tokens.”

Sivulka systematically reframes AI marketing claims by testing them against a workforce lens, and finds every one breaks down under scrutiny:

He extends this to a broader indictment of bloated org charts, noting most companies are already “mismanaged” with workers who function as “cogs in the machine”—and that Elon Musk’s 80% staff cut at X performed better precisely because the cuts removed dead weight, which he says is mirrored by AI: “Just like 80% of employees do nothing, 80% of tokens today do nothing.”

His fix: The “100x token”

Rather than concluding AI is broken, Sivulka argued that the solution is the same one railroads found in the 1840s: better management, not less technology. He predicted the defining skill of the next decade won’t be engineering talent but context engineering: “The 10x engineer built the last era of companies. The 100x token will build the next.” His summary of the economic shift happening now is that “humans are cheaper than tokens on average, but good tokens are cheaper at scale. Management converts one into the other.”

UBS’s sources describe working toward this exact fix independently through “model routing”—matching specific tasks to specific models rather than treating AI as one undifferentiated tool. One AI firm executive explained the shift: “About six months ago, we’d take a whole task and say, ‘all right, this model is probably the best model for it.’ Now, the individual sub-tasks within that project will go to different models because we know exactly which models are good at which tasks.”

That same executive described the payoff in terms almost identical to Sivulka’s “100x” framing: “There’s a cost arbitrage opportunity when you can source from a bunch of different models… we have all this data on what the models are good at, what specific tasks they’re good at, and so there’s a lot of arbitrage on the price side that we can have and that we can pass on to our clients.”

Another firm described the split even more explicitly, telling UBS that for routine workflows “where the human-like experience doesn’t matter as much,” they now use “cheaper or faster models,” while reserving frontier models for “core use cases” that are “our bread and butter”—a real-time version of Sivulka’s argument that management, not raw model power, converts wasted spend into leverage.

Sivulka also warned of organizational friction ahead, as employees start resisting handing over their institutional knowledge to AI systems that may eventually replace them. He pointed to Meta, where equity-holding employees—despite being financially incentivized to want AI to succeed — have pushed back against the company using their own work context as training data, calling it “a microcosm of what is about to happen across every industry.”

“Context hoarding,” he warned, is emerging as “the latest job security tactic,” describing it as a “massive political problem with AI” inside companies that will only get worse. “Employees don’t want to teach AI systems their secret sauce,” and now that they know their management isn’t smart enough to use tokens cheaply, they have leverage.

Disclaimer : This story is auto aggregated by a computer programme and has not been created or edited by DOWNTHENEWS. Publisher: fortune.com