Imagine the scenario, somebody in your organisation likely finished their week on Tuesday and did not mention it.

They were right not to. Every incentive above them is arranged to make sure they never do, and that silence is the most economically interesting thing happening in white-collar work right now.

It has a two-hundred-year backstory, so we should start there. The working week everyone has been indoctrinated into following is not a fact. It is a settlement dating back to 1817 when Robert Owen, running cotton mills at New Lanark, proposed a formula against the 12 and 14 hour days that defined British industry: eight hours labour, eight hours recreation and eight hours rest.

Just over 100 years later Henry Ford announced a five-day week in 1922 and implemented it in 1926. He did not invent the weekend, as departments had moved between five and six-day schedules over several years, and Ford said in an authorised interview that 40 hours had matched or beaten 48 from a productivity perspective - an early version of "following the data", even if the data itself was thin. What he did do, which mattered enormously, was demonstrate at industrial scale that hours and output were not interchangeable. Just over 12 years later the United States wrote the 40 hour working week into federal law in 1938.

Keynes, and why you should distrust everything that follows

In 1930 Keynes published his famous essay "Economic Possibilities for Our Grandchildren", predicting that within a century technological growth would yield a 15-hour workweek and vastly higher wealth by 2030. Productivity per worker-hour has risen roughly eightfold since he wrote that. The average week fell from 47 hours in 1930 to about 39 by the 1970s.

Then it stopped, with the best part of fifty years of no movement.

The productivity output was real however it was largely absorbed into consumption, into rising standards, into the quiet expansion of what a week is expected to contain. But it was never converted back into time to the employee.

Which brings us to the four-day week, and why I think it is the last argument of its type. The UK 2022 pilot ran 61 companies and 2,900 people on the 100-80-100 model. Fifty-six carried on. Turnover fell 57%, burnout fell 71%, revenue rose slightly, and 15% of participants said no sum of money would return them to five days.

Two hundred guineas

On 25 and 26 November 1878, at the Old Bailey, James McNeill Whistler sued John Ruskin for libel.

Ruskin had reviewed Whistler's Nocturne in Black and Gold: The Falling Rocket and accused him of asking two hundred guineas for flinging a pot of paint at the public. He saw the painting as art made with minimum effort for maximum output. Capitalist production applied to something that ought to have required struggle.

Under cross-examination, Ruskin's counsel asked Whistler how long the painting had taken. Whistler said two days. Counsel pressed: two days of labour, and for that you ask two hundred guineas. Whistler's reply is the most useful sentence in this essay. "No; I ask it for the knowledge of a lifetime."

That exchange is where the pricing argument of the next decade sits. The counsel at the time was stating time and materials in its purest form: labour multiplied by duration equals price. Whistler was stating outcome pricing: you are not buying the two days, you are buying the compressed lifetime that made two days sufficient. Thus, the painter, with a lifetime of experience wielding a brush against canvas, is analogous to a seasoned professional wielding a state of the art (SOTA) LLM and harness (e.g. Claude / ChatGPT).

Whistler won the case, incidentally. The jury awarded him a farthing, the costs bankrupted him and he ended up leaving for Venice. So being right about how the work should be priced didn't do a great deal for him at the time!.

What Ford proved, and what he did not

Ford's decoupling was asymmetric and bounded. What he was really describing is fatigue, in that the marginal hour of a tired worker is close to worthless, and fatigue arguments have a floor to them. You cannot get from 48 hours to eight by removing tiredness, because below a certain point the hours genuinely are the input required to complete the task. So hours and output are loosely coupled, but only at the tail.

What is happening now is not that. AI consequently removes hours from the middle of the distribution rather than the tail, and it has no fatigue floor (other than compute capacity). Retrieval, synthesis, first drafts, document review, model building, deck production. None of that was ever limited by how tired you were. It was limited by how long it takes a human to assemble the context and produce the output.

And here is the part that answers the question properly. Ford's claim was about the worker: the same person, fewer hours, same output. This claim is about the work. The output no longer requires the hours, but it requires a very particular person to produce it, and the scarcity has relocated. It has moved out of time and into judgement: the accumulated sense of what matters, plus enough recent fluency with AI tools to direct them, plus the ability to know when the output is wrong.

Why we cannot help pricing time

If the value is the accumulated judgement, why does anyone still care that it took five hours instead of five days?

Because we are wired to. In 2004 Kruger and colleagues gave people an identical poem and told half it had taken eighteen hours to write and half that it had taken four. The eighteen-hour group rated it better and put a higher monetary value on it. The same held for paintings and for suits of armour. They called it the effort heuristic, and its most important feature is that the effect strengthens when quality is hard to assess.

Consulting deliverables, strategy documents, architecture decisions and creative work are precisely the category where quality is hard to assess, which is where the heuristic does the most damage.

There is also a hard-wired human trait that studies have picked up on, which is that people tend to prefer slower advice from a human, reading the delay as care, and faster output from a machine. The AI-assisted expert therefore falls into the gap between both intuitions. They look like a person and perform like a system, and the heuristic penalises them for it.

So when you ask whether I would be comfortable paying the same for five hours as for five days: economically, obviously yes, because I am buying the artefact and the judgement behind it. But psychologically, for most of us it's still a no, because we've been trained on that heuristic for generations.

The driveshaft

That is the individual version. The organisational one is older, and it starts with the fact that factories in the age of steam were built around a central driveshaft. One engine, one shaft running the length of the building, belts dropping to each machine. The floor plan was not designed around the work. It was designed around the shaft.

Every organisation I know is currently bolting agents onto the driveshaft, and only a handful (facing existential demise, see SaaSpocalypse…) have managed to reconstruct from the ground up.

If you were to try to liken the driveshaft in a modern company it would be the layer that moves context between people and applies governance to it, aka Middle Management.

Middle management is not the villain here. It's existed as a layer of necessity because for a century there was no cheaper way to hold organisational state and stamp it approved. It was rationally priced: the expected cost of an error exceeded the cost of a review.

That ratio has now flipped though. When a review can run continuously across everything at close to zero marginal cost, the case for holding it as a scheduled human gate gets a lot weaker, and most organisations haven't been back to look at the pricing since.

Which is the same mistake, in a different place. If judgement is the scarce input and hours are the residual, then anything priced on hours is measuring the wrong quantity.

The bill

Time and materials is the purest version of that.

Consulting economics has always run on leverage. A small number of expensive partners standing on a wide base of capable, cheaper juniors doing the research, the models and the slides. Margin therefore comes from the spread between what a junior costs and what a junior bills. That is the pyramid, and it has funded elite professional services since the 1960s.

AI eats the base of the pyramid first. A junior who billed sixty hours to research a market and build a deck now does it in six (or likely less). Which produces a genuinely perverse result: a firm that bills by the hour is financially punished for deploying the tool properly. Full deployment cannibalises the revenue line. The rational move for an incumbent is therefore partial deployment. Enough to look modern in the pitch, not enough to compress the bill.

As a result that pyramid base is already thinning. Graduate postings across accounting and consulting fell 44% year on year by 2024. KPMG UK cut its graduate intake by 29% in a single year, from 1,399 to 942. Which raises a question the industry has not answered. Juniors acquired judgement by grinding through the volume work, so you cannot automate the base of a pyramid and expect the apex to keep replenishing itself.

The token question

Here is the version of this I find hardest to argue against, and it applies wherever the work happens inside your own environment.

You provision the tenancy. You pay for the compute. You pay for the tokens. The models are grounding on your data. The artefact that emerges belongs to you. And then you also pay an hourly rate for the person supervising the run.

What exactly are you buying at that point? (a) Judgement about where to point the machine and whether the answer is right. (b) Accountability, meaning somebody carrying professional risk and signing their name against an outcome. And (c) pattern recognition across dozens of engagements or workstreams. But note, none of these is time.

If the tokens are yours, the data is yours and the artefact is yours, the only question left is whether the judgement is worth what it costs, and that is a very different negotiation from a rate card.

The reason none of this shows up

Around 91% of firms now use AI "somewhere". About 89% of managers report no change in output per employee. Only 39% can trace any effect to earnings and MIT's research puts 95% of enterprise pilots at no measurable P&L impact.

The tasks are where the impact is likely being seen, in reduction of time to analyse the data, in producing the deck that would've taken days, in pushing out PRs for product and backend development. But when you dig deeper within organisations, it's rare that any of these processes were ever measured in the first place. There is little solid ground on which to build a defensible ROI statement, and the most likely result of presenting one is that it's ripped apart at presentation.

The other side of this is misaligned incentives. An organisation values time, as do leaders, and nobody in the chain is rewarded for declaring that something took less of it. It takes a bold legacy organisation to break that chain and value based solely on output x quality.

However, where you can see it in clear light is at AI Native startups. Having met with numerous founders to see their product, watched podcasts and been to industry events, I'm convinced that the new class of organisation is becoming an amplification of productivity vs said incumbents. Nobody is measuring that properly either, because the focus is on Enterprise productivity and thus whether the AI Capex run is a bubble, but the startup evidence is anecdotal or only seen through a VC lens, but it’s there, and that native speed with the ability to automate as they go, unencumbered by legacy processes and middle management that will win out in many cases over Enterprises.

Within Enterprises small AI workflow change is not going to show up as a defensible ROI that can be banked for many months down the line, and this is where belief and trust that this is the right investment choice at board level matters.

I have watched teams accelerate their drafting and succeed only in building a bigger backlog in front of someone who still meets fortnightly. This is where proper support is needed for teams going through AI transformation change. Identifying how to automate away processes along the way is critical to not just creating a bigger set of regular tasks, but amongst all of this argument three critical questions remain, which will play out over the coming 12+ months:

  • How do we value time, and will we continue to believe in a 5 day working week, or is that up for reimagining? How does that play against the previous generational transitions.

  • How will legacy companies survive against AI Native upstarts that can output at many orders of magnitude faster?

  • What happens to the consultancy market, where Time and Materials becomes Tokens + Experience?

I write about AI transformation, product and innovation. Looking at what survives contact with a large organisation and opportunities for startups. Subscribe if that is useful.

Sources

Owen's 1817 formula. Standard labour history.

Ford: five-day week announced 1922, implemented 1926. Teaching American History; Silicon Canals on the limits of Ford's evidence.

Fair Labor Standards Act, 1938.

Keynes, Economic Possibilities for Our Grandchildren, 1930. Productivity roughly 8x per Gordon (2016); week 47 hours in 1930 to about 39 by the 1970s, flat since.

UK four-day week pilot, June to December 2022: 61 companies, about 2,900 workers, 100-80-100. Autonomy Institute, University of Cambridge, Boston College, February 2023.

Whistler v Ruskin, Old Bailey, 25-26 November 1878. Ruskin's review appeared in Fors Clavigera, July 1877. Whistler's own account: Whistler v. Ruskin: Art and Art Critics (1878), reprinted in The Gentle Art of Making Enemies (1890). Damages: one farthing.

Effort heuristic: Kruger, Wirtz, Van Boven and Altermatt (2004). Mixed replication: Ziano, Yeung, Lee, Shi and Feldman, Collabra: Psychology (2023).

Consulting: graduate postings down 44% year on year by 2024; KPMG UK graduate intake 1,399 to 942.

AI paradox: MIT NANDA GenAI Divide; NBER firm tracking.