When Inference Gets Cheap, Choose the Skill That Compounds

Next step: pick one skill, not three. That sounds like advice-card wallpaper until a moment like this arrives — a single engineering announcement that quietly changes which skills are worth building.

Here is what happened. A leading AI lab announced its first custom inference chip, aimed squarely at the reasoning stage of artificial intelligence rather than at training. The headline number: 1.5 to 1.9 times the peak performance per watt of the reference system. The stated purpose is just as direct — to cut the power and the cost of running large-scale AI reasoning at the level of autonomous agents.

I am a career coach, not a chip analyst, so let me translate that into the language I actually use. The moment a piece of work gets dramatically cheaper, more of it gets done, and the people who know how to direct it become more valuable. Electricity got cheap, so we electrified everything. Compute is about to get cheaper at the exact margin where intelligence actually gets used — and that is a fork in the road for everyone whose work touches AI.

Training was the cathedral; inference is the street grid

Start with the distinction between training and inference, because the difference is where the whole story sits. Training is the expensive, one-time act of teaching a model. Inference is the running cost — every question you ask, every response you generate, every agent action you let it take. When someone says AI is expensive, most of that bill now lives in inference. A chip that delivers 1.5 to 1.9 times the work per watt is not a small step; at the scale these systems run, it is a meaningful cut to the per-use price of intelligence.

And here is the part that made me pause. The lab said the design goal is lowering the cost of agent-scale reasoning. Not training a bigger model — running more small, continuous, everyday decisions. That is the difference between building a cathedral and paving a street grid. The first is a statement; the second is infrastructure. When a technology stops being a landmark and becomes infrastructure, the people who earn from it change. The frontier builder gets replaced by the system operator, the evaluator, the person who knows what to ask and what to check.

I chased every feature once, and it cost me

I have been through a version of this before, and I got it wrong the first time. Early on, when generative tools first went mainstream, I advised a client to chase every new feature as it shipped — the logic being that early access meant advantage. That was a mistake. The features moved faster than anyone could build on them, and what actually paid off, eighteen months later, was the person who had quietly built a review-and-edit workflow that survived every new release. Same hours spent. Different compounding.

No, let me correct myself slightly: it is not quite that chasing is useless. It is that chasing is a tax you pay for discovery, and compounding is the account that eventually pays you. The useful skill is the one that keeps working when the tool under it changes. That is the honest version of the lesson, and it matters here because an inference cost cut does not change what tools exist. It changes how many of them can affordably run — and how many of them you can afford to run in parallel. The supply of usable intelligence expands. The demand for people who can direct, verify and combine it expands with it.

A method that survives contact with a real week

So what is the next step? Not “learn more AI.” That is not a plan. A plan has a direction and a deadline. Here is the method I have been recommending since this announcement crossed my desk — and it is deliberately boring.

First, map your own work week and find the single most repeated task that still happens by hand. Not the hardest task. The most repeated one. For most people that is summarising, checking, rewriting, reconciling — the work you do so often you stopped noticing you were doing it. That is your target.

Second, build one working pipeline for that task, using whatever tools you already have, and commit to using it every time for thirty days. The goal is not elegance. The goal is that the pipeline survives contact with a real week — Mondays, deadlines, a child home sick, the 6 p.m. version of you who is out of patience. If it survives that, it is a skill. If it only worked when you were fresh, it was a demo.

Third, and this is the step everyone skips: decide what you will not automate. The reason inference getting cheaper matters is that you will soon be able to hand more and more judgment-adjacent work to machines. The counter-intuitive move is to be explicit about the ten percent of your work that stays human — the decisions where being wrong has a cost you cannot explain in a prompt. Draw that line now, on paper, before the cheap compute arrives and makes the default choice feel frictionless. People who know their line will get the better work. Everyone else will be pleasantly busy.

Be honest about what is still unverified

I want to be honest about what I do not know here. The chip is announced, not shipped at scale; the efficiency claims are the lab’s own; and whether the real-world gains land as cleanly as the spec suggests is genuinely unverified. Per-watt numbers are optimistic right up until they hit a production rack. So treat the magnitude with healthy scepticism — but not the direction. The direction, which multiple sources and the entire history of computing agree on, is that the unit cost of intelligence falls, and it falls for a long time. Every career built on a fixed unit cost of intelligence needs to be re-priced.

A specific Thursday evening

That re-pricing is not abstract. Picture a specific Thursday evening in a specific team: someone has built a small agent that drafts half their routine reports, and the manager, who cannot articulate why, keeps asking them to do the reports “properly” anyway. That scene is playing out in offices right now, and it is not about the tool at all. It is about trust in a judgement layer that nobody wrote down. The worker who documented their own quality bar — the one who wrote the checklist the agent is now held against — does not fight that fight. They set the standard. That is a concrete image of the future I mean: not machines replacing people, but people who articulate standards replacing people who only execute them.

This is why I keep coming back to the fork metaphor. The path splits here, and the two branches look almost identical at first. Both involve learning new tools. But one branch is learning tools as an end, and the other is building judgement as the asset, with tools as the cheaper, more abundant lever for it. The first branch gets you to next month faster. The second one compounds.

The advice I would give a friend over coffee is the same as the advice I would give a room of engineers: do not place your bet on the price of a thing, place it on the thing that gets more valuable as the price falls. When compute was expensive, the rare skill was access to it. When compute becomes cheap, the rare skill is what you do with it — and that has always been the rarer skill.

So here is your next step, and it is actionable in the literal sense of the word: this week, pick the one repeated task, build the one pipeline, and write down the one thing you will not automate. That is a thirty-day project, not a lifestyle. By the time the cheap inference actually arrives in production, you will not need to scramble for a strategy. You will already be compounding.

The supervisor economy is already forming

When a cost falls, the constraint moves. Cheap inference means the number of things a single worker can set running in parallel stops being a budget question and becomes a judgment question. I have watched this shift happen inside teams over the past year: the scarce skill is no longer knowing how to make a tool produce text — it is knowing which of the dozen outputs on your screen deserve the two hours of human attention you have left that day. That is a supervisor economy, and it is already forming in the teams I coach. The people getting promoted inside it are not the ones who prompt best. They are the ones who can look at five machine-drafted versions of a report and know, without being told, which one is wrong, which one is merely average, and which one is ready to send. That judgment is not learned in a course. It is built the way every durable skill is built — by doing the work, making the call, and being held to it.

A one-hour audit that does the work for you

Here is the method in its most compressed form, and it fits in one honest hour. Write down the last twenty decisions you made at work where being wrong would have cost something real. Now sort them into two piles: decisions where a machine’s output would give you everything you need to decide, and decisions where you would want to see the underlying situation yourself before signing off. That second pile is your line. Most people discover the line is much smaller than they feared — ten percent of their week, at most — and that discovery is quietly liberating. It means the other ninety percent can be delegated, automated, and audited. What remains is the part you actually get paid for, and it is the part that compounds.

What to tell a manager this quarter

The actionable version, if you manage a team, is a question you can ask this week: of everything my team produces, what would I be willing to let a cheaper, faster version of the current tool draft, so long as a human checked it? The teams that answer that question honestly will reallocate hours toward the parts of the work that still need a person. The teams that refuse to ask it will keep spending human hours on tasks that are about to become nearly free. Same headcount, different compounding — and the difference will be visible within two quarters, not years. That is the next step of the story the chip announcement opens: not a faster tool, but a different allocation of the scarcest resource you have.

Choose the fork that compounds. The tools will keep changing; the judgement you built around them will not.