Same Game, Faster Clock
Why the economy's oldest bottleneck just became its most dangerous
Every company is a way of asking a question to the universe. The East India Company: can a corporation organize trade across oceans? Ford: can mass production turn a luxury into an everyday good? SpaceX: will life be multiplanetary? A couple months ago I joined Mercor because I thought the labor matching problem was the most important question of our generation. But I couldn’t have told you what Mercor’s question was. Not cleanly.
Then yesterday, during an all-hands, something shifted. Our leadership was talking about defining a new category of work – that the economy is becoming a reinforcement learning environment, one where humans work to teach models, and that this creates an entirely new type of job. One of our founders had written an essay making that argument before I joined. It's why I wanted to come here in the first place. At a time where everyone, optimists and pessimists alike, is talking about automation and labor displacement, getting the human side of this right seemed like the most important problem to work on. But sitting there, I kept turning over another question.
Hasn’t it always been?
That turns out to matter quite a bit. Because if the economy is becoming an RL environment, the job is to prepare for a transition – to build something new for unfamiliar territory. But if it always was one, the job is something different: figuring out what’s actually changing, and what that demands from us now.
And the question I think Mercor is really asking – the one I couldn’t articulate until yesterday – is this: as the economy gets faster, can humans keep up? And if so, what does it take to make that possible?
Reinforcement learning sounds technical but the core idea is almost embarrassingly simple. You have an agent. It exists inside an environment. It takes an action, the world responds, and it receives a signal telling it how well it did. It updates. It tries again.
That’s it. And if you stare at that for a second – that’s just markets.
A firm tries a pricing strategy, customers respond, revenue goes up or down, the firm updates. A worker develops a skill, the labor market responds, wages move, the worker updates. Every economic agent is an agent. The economy is the environment. Profit is the reward signal.
The loop has been running for as long as markets have existed. So what’s new?
To answer that, I think you have to go back to 1945.
Friedrich Hayek’s “The Use of Knowledge in Society” started with a simple premise: the information required to run an economy efficiently isn’t sitting anywhere. It’s scattered across millions of people in forms that are local, tacit, and constantly changing.
A baker knows what flour costs this morning in his neighborhood. A farmer knows her soil. A trader knows the mood of the dock. None of them could fully explain what they know – and even if they could, by the time you collected it all, it would already be out of date.
He said this is why central planning fails. Not just politically, but computationally. A Soviet planning committee trying to allocate resources across an entire economy was attempting to approximate a value function it could never fully observe. The state space is too large, the data too local and perishable, the environment shifting too fast. When you can’t observe the real value landscape, you optimize a proxy instead – tons of steel produced, quotas met – and destroy the actual value you were trying to create. This is Goodhart’s Law. Hayek diagnosed it in 1945 without the vocabulary.
(We call it the “sim-to-real gap” – benchmarks get saturated but don’t reflect what actually matters in the real world. Narrowing that gap is Mercor’s core focus now. But it’s worth noting that the economy already has something most benchmarks don’t: a reward signal grounded in reality. Profit isn’t a perfect signal – short-term profit can diverge from real value creation – but over long enough horizons, it’s the tightest feedback the system has. No benchmark invented in a lab comes close. Hayek understood this. The plan is never the territory, but the P&L always is.)
His solution was prices. Prices don’t require anyone to understand the whole picture. They aggregate distributed information automatically, transmitting just enough signal for each agent to act from their own position without needing to see the whole board. In RL terms: prices are the reward signal of a system where no one holds the full value function. That’s exactly why they work.
Hayek’s argument isn’t just that aggregate knowledge is distributed across society. It’s that individual knowledge is unique. Every person holds special knowledge of their own circumstances – their skills, their position, their situation – that no one else has and no institution can fully replicate.
A nurse who has spent twenty years in emergency rooms knows things about patient triage that no textbook captures. A comedian can feel laughter before it’s audible. A teacher knows which kid is about to disengage before the kid does. That knowledge is real, it’s economically valuable, and it exists nowhere else.
What this means is that the economy has always been trying to route to unique human value. The matching problem – finding the right person for the right task at the right moment – is fundamental to how the system creates value. It has just, historically, been very slow and very bad at it.
That's the knowledge problem. Not just that information is distributed, but that matching people to the work they are uniquely positioned to do well is the central bottleneck of the whole system. Energy, compute, models, the app layer – all of these are being optimized at extraordinary speeds. But none of that matters if the human on the other end is in the wrong role, working on the wrong problem, deploying knowledge the system doesn't need while the knowledge it does need sits undiscovered in someone else. The bottleneck was always matching.
So what actually changed?
Not the mechanism. The clock speed.
Humans have always been the agents hill-climbing this RL environment. A skill becomes valuable, people notice, they acquire it, new job categories crystallize – and this process plays out over years or decades. The economy updates, but at human speed.
Once everything else gets optimized – energy, compute, models, the app layer – agents will do this in hours. Same environment, same reward signal, but a radically faster optimizer running inside it.
And this is why the “always was” part matters. If the economy is becoming an RL environment, the implied frame centers AI as the novel thing and humans as the thing that must adapt to it. The question becomes: what new system do we build for this new reality?
But if it always was one, the question flips. Humans have always been doing RL – just slowly. The arrival of AI agents doesn’t change the game. It changes the clock. And that means the real question isn’t “what new thing do we build?” It’s “what existing bottleneck just became fatal?”
The answer: the speed at which humans can discover, match, and deploy their unique knowledge.
This is also not the first time new agents have entered the economy. The rise of multinational corporations, the formalization of institutional investors, the expansion of government agencies – each introduced a new kind of economic agent that the system had to absorb. Each time, the question was the same: how do humans relate to these new agents, and infrastructure do we need to keep the machine running?
AI agents are the latest entrants. But they’re different in one critical respect: they don’t just participate in the environment. They optimize it faster than the human agents who have been running the loop since the beginning.
When the optimizer speeds up by several orders of magnitude, something that looked stable starts to look fragile. Feedback loops that felt tight enough suddenly aren’t. Humans who were keeping pace find themselves falling behind – not because they got worse, but because the clock sped up.
So where does that leave humans?
There’s a tempting answer I think is wrong: that humans will do the “creative” work and AI will do the “routine” work. That framing is too vague to be useful and probably too optimistic to be honest.
Here’s what I think is more honest. Models and humans aren’t really doing different kinds of reasoning. They’re reasoning over different data.
Models are extraordinary at inference over known data – everything already captured in their training set. If A implies B and B implies C, a well-trained model will find that A implies C faster than any human, at scale, across every domain it’s been trained on. This is enormously powerful. And most of what most humans do most of the time is the same thing, just slower: chains of implication, patterns recognized across known information, inference through existing structure. Think of compliance teams keeping up with cascading regulations. That work is real, but it’s largely inference over public data – and models are going to eat it.
The human comparative advantage is something different – but not in the way people usually think. It’s not that humans reason in some fundamentally novel way that models can’t replicate. It’s that humans have access to inputs no model can see.
Every person is situated somewhere unique in the world, holding knowledge no one else holds. The nurse who catches what the chart missed isn’t doing magic – she’s making an inference from twenty years of embodied experience, from something she noticed in the room that never made it into any dataset. The founder who sees a market gap isn’t thinking in a way a model couldn’t. She’s thinking from a vantage point the model has never occupied.
This is Hayek’s point, precisely. The advantage isn’t unique reasoning. It’s unique knowledge – observations, context, felt signal that hasn’t entered any training set. The model can’t get there not because the logic is beyond it, but because it can’t see what you see from where you’re standing.
When a human makes an inference from private knowledge and that inference enters the system, the output is novel even if the reasoning isn’t. That’s what makes it exploration – not the logic, but the fact that new signal entered the loop. And the economy needs that signal to find new ground.
The question is whether the infrastructure exists to find that signal and route to it fast enough.
Most conversations about AI and labor end up at fairness: how do we make sure people aren’t left behind? That’s a real concern. But leading with fairness gets the stakes wrong.
Any RL environment needs two things to function well: exploitation and exploration. Exploitation is optimizing what you already know – extracting value from discovered structure. Exploration is finding new structure – discovering value that didn’t previously exist. Both are essential.
AI agents will become extraordinary exploiters. Given a known value landscape, they will find the optimum faster than any human. But exploration requires generating genuinely novel signal – new relationships, new structure – from positions in the world the model hasn’t seen. That’s the mechanism humans provide.
If humans can’t keep pace – if the infrastructure for human discovery breaks down – the economy doesn’t just lose some workers. It loses its exploration mechanism. The system keeps optimizing existing structure, gets trapped on local maxima, and stops finding genuinely new ground.
This is not a fairness argument. It’s a question of whether the system collapses. There are important discoveries that need to be made, important problems that require unique human knowledge to even see, and making that happen requires solving the matching problem – routing people to the work they’re uniquely positioned to do, not the work they happen to have stumbled into.
That’s Hayek’s knowledge problem. And it’s now urgent in a way it wasn’t before, because the clock is running fast enough to make the bottleneck fatal.
Right now the infrastructure for human development is badly mismatched to the speed of the environment it’s operating inside. Education is front-loaded – it happens once, early in life, for job categories that already exist. By the time a curriculum reflects a new economic reality, that reality has often moved on. Labor market matching is slow and noisy – people struggle to articulate their unique position, and the system struggles to find it. Feedback loops are long. It takes years to know if you’re developing in a direction the economy will actually value.
In a slow RL environment, this is just inefficient. In one running at the speed we’re entering, it’s dangerous.
What needs to exist is a system that helps people identify their unique position in the value landscape faster, reduces the cost of trying new directions, and tightens the feedback loop between human knowledge and economic signal – continuously, not just at the start of a career.
Building the environments where humans teach models is one critical piece of this. When a nurse corrects a model’s triage recommendation and that correction becomes training data, something real happened. Unique local knowledge entered the system in a way that pure model-to-model iteration could never have produced. The human discovered something. The machine learned it. The loop stayed alive.
But right now, most of this work doesn’t look like that. Even if you’re a great domain expert – a finance professional, a physician, a senior engineer – working as a human data contributor for an AI lab looks a lot like remote piecework. It’s isolated, decontextualized, and stripped of the factors that make real-world knowledge real: motivation, social context, gut instinct, the felt weight of consequences.
This matters for a reason that has nothing to do with worker satisfaction. A 2012 experiment found that when crowdworkers were told their image labeling was helping medical researchers, they produced more output. When told their work would be discarded, quality collapsed. Meaning isn’t a perk – it’s a data quality problem. When the human generating signal doesn’t experience the work as real, the signal degrades. You get out-of-distribution data. And if you optimize on that upstream, nothing downstream matters.
This is the sim-to-real gap applied to the human side of the pipeline. If the environments where humans train models don’t preserve the conditions that make human knowledge valuable – real stakes, real context, real feedback – then you’ve built a system that optimizes on a simulation of human judgment rather than the thing itself. Hayek’s warning again: the plan is never the territory.
The fix is generating training data in real-world settings where those factors actually exist. Profit as verifiable reward. High-fidelity environments. Work that feels meaningful because it is meaningful – because it’s real work, producing real outcomes, with the texture and complexity that simulated environments strip away.
That’s the loop that has to stay alive. And that’s what Mercor is working on – the matching infrastructure that routes humans to the work they’re uniquely positioned to do, in environments real enough that their knowledge actually survives the trip.
But this is bigger than any one company. The matching infrastructure, the real-world training environments, the tighter feedback loops – this is what solving Hayek’s knowledge problem actually looks like. Every person is already sitting on unique knowledge the system needs. The problem, now acute in a way it wasn’t before, is routing to it fast enough.
The infrastructure we build around human discovery isn’t a safety net. It’s the engine.
Get it right and the upside is staggering. Diseases get cured because the researcher with the right intuition got matched to the right problem at the right time. Epidemics never become pandemics. A kid who never had a teacher gets one. Every person's unique position – more legible, more reachable, more valuable than ever before. Not because we built better AI, but because we built the infrastructure that let human knowledge flow where it was needed.
Get it wrong and nothing else matters. It’s quiet. Models train on increasingly stale data. The economy optimizes the same shrinking landscape – efficiently, impressively, and with diminishing returns. New industries don’t get created. New categories of value don’t emerge. The system feeds on its own outputs, gets narrower, and calls it progress. This is model collapse applied to the real economy: not a crash, but a slow hollowing out where everything looks like it’s working until you realize nothing new is being found.
This has happened before. The Egyptians built the pyramids and we still don’t fully understand how. Civilizations don’t just fail to advance – they forget what they could already do. Not because anyone chose to forget, but because the infrastructure for transmitting knowledge broke down. Unique knowledge that doesn’t get routed into the system doesn’t wait around. It disappears.

