How Your Firm Should Price, Staff, and Govern Legal AI as the Cost of Intelligence Splits in Two
A practitioner’s guide to the bifurcation reshaping legal pricing, and what to do about it before your clients, or a court, do it for you.
The short version. The confident forecasts that the billable hour is dying all rest on one premise: that AI is, and stays, cheap. That premise is only half true, and the half it gets wrong is the half that matters for your firm. Commodity legal work is racing toward near-zero cost, which is real and will reshape the bottom of your practice. But the most capable models, the frontier, are getting more expensive and scarcer, not less, and a large share of firms and clients will not be able to afford them. The market is not flattening. It is stratifying. The hour does not simply die; it splits into a fixed or per-outcome price at the bottom and something new at the top: a reasoning budget, the governed allocation of scarce frontier capability to a matter in proportion to what is at stake. This piece is about what that split means for how you price, staff, and govern, and what to do now.
I. The assumption your competitors are pricing on
Read the 2026 predictions and you find the same logic every time: AI compresses the time a task takes, compressed time cannot sustain hourly pricing, therefore fixed fees, subscriptions, and the slow end of the six-minute increment. The logic is sound. The premise beneath it, that AI is cheap, and cheap for everyone, is not. “Cheap” describes only one layer of the market, and the best models for serious legal work will not be affordable to every firm or every client. A firm’s competitive position will increasingly depend on whether it can afford the frontier. Many cannot, and will not. That is a stratification story, and it runs in the opposite direction from the consensus.
The consensus is already cracking, and open source is the leading edge. In spring 2026 a former Latham & Watkins associate released Mike OSS, a free, self-hostable, bring-your-own-key legal AI, and reproduced the core of what Harvey and Legora sell at enterprise prices in roughly two weeks. It need not displace the incumbents to matter. As one former law-firm IT director observed, once a working open-source alternative sits on GitHub, the renewal conversation shifts from “is this magic?” to “what exactly am I paying enterprise prices for?” Open source drags the commodity layer toward zero and forces every provider above it, including your firm, to name the value it actually adds. The practical question this poses is not “will AI make us cheaper,” but “which layer do we compete in, and can we afford the one we are aiming at?”
II. Two cost curves, not one
The unit cost of intelligence is genuinely collapsing, per-token prices have fallen by factors ranging from roughly nine-fold to several hundred-fold per year. That deflation is the engine of access: when the marginal cost of a competent contract review or a benefits appeal approaches zero, legal help becomes reachable for people who could never previously afford it, the low-income Americans who receive inadequate or no help for 92% of their civil legal problems. The floor of the market rises, and it rises for everyone.
But total spending on intelligence is rising, not falling, because consumption, agentic systems that reason in long chains and call models thousands of times per task, is scaling far faster than unit prices decline. Gartner puts the point sharply: chief product officers should not confuse the deflation of commodity tokens with the democratization of frontier reasoning. Cheaper commodity tokens and a more expensive, faster-moving frontier are not a contradiction; they are the same fact seen from two ends. Three forces keep the frontier scarce. It moves, today’s most advanced reasoning becomes tomorrow’s commodity, and a costlier frontier always opens above it. The bottleneck migrates rather than vanishes, when raw reasoning is cheap, the scarce inputs become verification, judgment, and accountability, none of which ride the compute curve down. And affordability stratifies, accessing the best of the frontier stays expensive enough that many firms will run on the tier below, the same way the best of anything has always been rationed by price.
III. Where your firm competes
At the lower tiers, capability becomes abundant, commoditized, and broadly affordable. Routine drafting, first-pass review, standard diligence, and regulatory tracking, work that filled much of a junior associate’s week, collapse into supervised, near-free workflows available to anyone. At the upper tiers, the opposite happens: the most advanced reasoning, fused with a firm’s proprietary data and validated workflows, becomes a widening and expensive source of advantage.
The counsel advising a private-equity acquirer on a contested, bet-the-company matter will not merely have more experienced partners than the firm serving a middle-market buyer; it will deploy frontier reasoning and proprietary capability the middle-market firm cannot afford to run. Nowhere is this starker than in high-stakes litigation, where the advantage compounds. A firm that pairs elite trial lawyers with its own corpus, decades of briefs, outcomes, judge-specific tendencies, and a record of what actually worked, and runs the best models over that private context will systematically outperform an opponent renting commodity capability over public data. The edge widens with every matter, because each case sharpens the corpus.
Capability is not the only axis; speed is the other. In a contested matter, advantage goes to the firm that reaches the right decision first, to preempt a filing, move inside a closing window, or respond before opposing counsel has finished reading. The analogy is high-frequency trading, where funds spend fortunes on co-located servers and low-latency links to buy an edge unavailable to anyone who cannot pay for it. This should not read as dystopian: legal quality has always been stratified by the ability to pay, and elite counsel was never evenly distributed. AI does not create that stratification; it adds two new dimensions to it, capability and speed. The strategic imperative for your firm follows directly: decide, explicitly and practice group by practice group, which tier you compete in. You cannot credibly be both the cheapest commodity provider and the frontier-premium shop, and the firms that drift without deciding will be undercut from below and outclassed from above.
IV. Why the billable hour bifurcates, and why the data already shows it
The hour was never really a measure of time. It was a bundle doing three jobs at once: recovering the cost of producing the work, pricing the expertise applied (a partner’s rate and an associate’s rate are an expertise tier, not a stopwatch reading), and absorbing risk by shifting the uncertainty of unclear scope onto the client. It survived for a century because labor was the dominant cost, a fair proxy for expertise, and correlated with effort-at-risk, so one number could do all three jobs. AI severs the three correlations at once, and the bundle comes apart along the fault line between the two cost curves.
At the bottom, where work is bounded, repeatable, and poolable, the hour genuinely dies. Clio’s data already shows up to 74% of hourly billable tasks exposed to automation, the average lawyer recording just 2.6 billable hours in an eight-hour day, and flat-fee billing up by more than a third over the prior decade. This work moves to fixed, per-matter, and subscription pricing, delivered at scale.
At the top, the hour persists, and not out of nostalgia. The work there is where risk cannot be pooled: a bet-the-company matter is a sample of one, its scope unknowable and its downside asymmetric, and no honest fixed fee can absorb that uncertainty without either gouging the client or ruining the firm. The market confirms it. Thomson Reuters found that whether firms discount aggressively or hold firm, they collect roughly the same per hour, and that despite years of heavy AI investment, roughly 90% of legal dollars still flow through hourly billing. The hour survives at the top of the complexity curve. The question is not whether it survives, but what it becomes.
V. The new unit: the reasoning budget
What the hour becomes at the top is a reasoning budget, not a measure of human time, and not a meter on raw tokens, but an allocation of governed reasoning capacity to a matter in proportion to what is at stake. Governed reasoning capacity is reasoning a licensed professional can stand behind: it carries provenance and chain-of-custody, runs under a controlled data regime, is validated against a quality bar, passes through human checkpoints with a named accountable person, and leaves an audit trail. It is metered and allocable like a budget, but unlike raw compute it comes with a quality-and-accountability guarantee, and that guarantee is what the client is paying for.
Two capabilities sit on top of the raw models and constitute the actual product. The first is orchestration, deciding which model handles which step, routing across a portfolio of systems, sequencing agentic workflows, and spending scarce frontier reasoning only where it earns its cost. Firms will differ enormously here, and no one yet knows what compute will truly cost: between the price war among the labs, the open-source floor, and the unwinding of below-cost subsidies, today’s token prices are an unreliable guide to tomorrow’s, and a firm that hard-codes them into a fixed fee is building on sand. The second capability is governance, the controls, auditability, and accountability that turn a powerful output into one a general counsel, or a court, will accept. The binding constraint on serious AI in regulated work has never been raw capability; it is accountability: who supervises the algorithm, who is liable when it is wrong, and what record proves compliance. Orchestration and governance are the scarce layer, because they are human, institutional, and regulatory work that does not ride the compute curve down.
This is also where proprietary capability becomes genuinely hard to copy. When a single client is served by a panel of firms, co-counsel on a deal, several firms across a litigation portfolio, federated learning allows those firms to improve a shared model on the client’s matters without any of them surrendering privileged data: the data never leaves each firm’s control, only the learning does. Privilege is preserved, confidentiality holds, and the model still improves for everyone working the client’s problems. The reasoning budget is the price of the whole bundle, senior judgment, frontier access, proprietary capability, and the governance to stand behind it, and it will not simply diffuse to every firm the moment the models are available to every firm. The precedent is electrification: the dynamo was available to nearly every factory by the 1890s, but the productivity gains took almost forty years to arrive, because they required redesigning the factory around the electric motor rather than bolting it in beside the steam engine. The firms that merely buy AI will get a faster steam engine. The firms that rebuild around it will get the next era — and they will be few.
VI. Pricing in practice
Start from an uncomfortable truth the hourly model spent a century hiding: legal services have never really been priced by cost. They have been priced by return. A board does not pay a premium on a multibillion-dollar acquisition because the work takes more hours; it pays because the downside of error is catastrophic and the value of getting it right is enormous. Price at the top has always floated toward the ROI of the outcome, with the hour as the polite accounting fiction connecting that price to something measurable.
AI removes the fiction, at least in part. Once the cost of producing the work decouples from its value, collapsing toward compute at the bottom, concentrating in scarce capability at the top, price can no longer pretend to track cost, and the honest anchor at the top becomes outcome. Success- and outcome-based pricing become not only defensible but, for a widening band of matters, estimable: litigation analytics now quantify motion outcomes, settlement likelihood, and judicial tendencies well enough to price risk that was previously unpriceable. So AI squeezes the hour from both ends, commoditizing the bottom into fixed fees and making the top estimable enough to price by outcome, and leaves it occupying a shrinking middle.
This forces a question the profession has deferred: is the cost of intelligence firm overhead, or a billable disbursement? The last technology to pose it was online legal research, and client pressure eventually pushed Westlaw and Lexis charges from pass-through line items into overhead. AI compute reopens the fight at far larger scale, and the volatility of frontier pricing makes it acute. Treat compute as overhead and quote fixed fees on today’s subsidized prices, and the firm is exposed when the false floor lifts. Pass compute through as a disbursement, and the firm offloads that risk to clients but collides with their demand for cost certainty, and with the settled principle under Model Rule 1.5 that a firm may not mark up disbursements above actual cost. The deeper consequence is that AI strips the hour of its ability to hide margin. When a model did the work in minutes, the firm can no longer bury its expertise premium inside a time entry; it must name what it is actually selling, orchestration, governance, judgment, outcome, and price each explicitly. Practically, that points toward blended structures: a capability or subscription fee for access to the firm’s governed stack, a success component tied to outcome, and metered compute passed through at cost. And it points toward transparency clients now expect as a matter of course, what the AI did, what it saved, and how that shows up on the bill.
VII. Build, buy, or orchestrate
For the general counsel, and for the firm competing to keep the work, the same stratification reframes the buying decision. The commodity layer becomes a contest to be cheapest-compliant, fought among legal-tech products, in-house teams, technology-enabled alternative providers, and increasingly self-hosted open-source tools a firm can run behind its own walls. A traditional firm rarely wins that contest on price and should not try; the sound move is to cede the bottom and move up.
What moves up is harder to commoditize. As routine work disaggregates to whoever is cheapest, the client’s legal-AI stack fragments into many tools, models, and providers, each with its own cost, quality, and compliance profile, and fragmentation manufactures a scarce role: the single accountable party who orchestrates that stack and owns the cost, quality, and compliance tradeoff end to end. The client’s real question shifts from “why isn’t this cheaper?” to “who will orchestrate this and stand behind it?” A firm can claim that role or cede it, to a Big Four entity, a legal-tech platform, an in-house operations function, or a new kind of governance intermediary. Claiming it means becoming the institution that supplies governed reasoning capacity at scale, sold as accountability rather than as time. (I set out the value-capture side of this argument for investors last week, in the companion edition of this piece in The AI Investment Briefing but the operative point for firms is simpler: this is a role you either build the capability to hold, or watch someone else take.)
VIII. The professional-responsibility landscape
For the profession, the governed layer is not merely a business preference; it is where the ethical duties now concentrate, and courts and bar regulators have moved faster than many firms realize.
The ABA has confirmed that a lawyer’s core duties apply in full to generative AI: competence under Model Rule 1.1 (which now includes understanding the capabilities, limits, and, increasingly, the cost profile of the tools you deploy), confidentiality under 1.6, supervision under 5.1 and 5.3, communication under 1.4, and reasonable fees under 1.5. Confidentiality deserves particular attention: self-hosting a tool narrows the privilege surface but does not eliminate it, because the moment a prompt leaves the firm for a third-party model API, privileged communications are being transmitted to an outside service, a point the ABA and several state bars now address directly. Supervision is the other pressure point: an autonomous or agentic workflow is a non-human agent whose output the responsible lawyer must still supervise and verify.
That verification duty is now enforced, repeatedly and without regard to a firm’s size or the sophistication of its tools. What began with Mata v. Avianca has hardened into a body of law: the Ninth Circuit has sanctioned and suspended counsel for undisclosed AI-hallucinated citations and warned its entire bar; the Sixth Circuit, in Whiting v. City of Athens, imposed Rule 38 sanctions for more than twenty fabricated citations; and the Sixth and Seventh Circuits, while agreeing that AI does not dilute the duties of competence, candor, and verification, have split on how harshly to punish lapses. Tracking databases now count well over a thousand hallucination incidents, with double-digit sanction decisions issued on a single day. The decisive fact for practitioners is that courts draw no line by firm size or tool quality: an AmLaw 100 partner using a law-trained model bears the same non-delegable duty to verify as a solo using a consumer chatbot, and elite firms have been caught alongside everyone else. Capability does not discharge the duty; governance does. Enforcement is therefore not a headwind for the governed approach — it is the strongest argument for it, pushing the market toward systems with citation provenance, retrieval grounding, and human-in-the-loop control.
Finally, structure. The rewriting of unauthorized-practice and alternative-business-structure rules in jurisdictions such as Utah and Arizona determines where the highest-value move, owning the service rather than merely selling software to it, is even permitted. For most firms, in most states, the traditional prohibitions still bind; where they do not, the strategic option set is meaningfully larger. Either way, the firm that treats governance and privilege as a competitive capability, rather than a compliance cost, is the firm positioned to hold the scarce layer.
IX. What your firm does now
If the analysis is right, the response is not to wait for prices to settle or for a consensus billing model to emerge. It is to act while the market is still forming. Six moves, in rough order:
Audit your matters onto the complexity curve. Sort each practice group’s work into commodity (bounded, repeatable, poolable) and frontier (bespoke, high-stakes, un-poolable). The mix tells you which tier each group actually competes in, which is rarely the tier it thinks it does.
Decide the tier, per practice, and commit. Cede the commodity bottom rather than bleeding margin trying to out-price legal tech and in-house teams; invest to hold the frontier top where you can. Drifting between the two is the one losing strategy.
Build or buy governed reasoning capacity, orchestration across models plus the provenance, audit, and accountability that let you stand behind the output. Treat compute cost as a variable to manage, not a constant to assume, and do not commit to fixed fees that depend on today’s subsidized token prices holding.
Stand up a citation-verification protocol and an AI-use policy now. More than half of firms still have neither, yet the verification duty is being enforced against firms of every size. A documented protocol, tiering tools by confidentiality risk, requiring human verification of every citation and quotation, and logging AI use, is the cheapest insurance available against a sanction.
Have the pricing conversation before your clients force it. Be ready to state plainly what AI did on a matter, what it saved, and how you priced it, capability fee, success component, pass-through compute, rather than burying the change inside an unchanged hourly rate that invites a Rule 1.5 challenge.
Keep faith with elevation, but see it clearly. AI in essential services is a capacity-expansion play, not an automation play: it lifts lawyers from gatherers of information to makers of judgment and expands the demand for that judgment. But the elevation is uneven, the floor rises for everyone while the ceiling rises fastest for the firms that rebuild around the frontier. Both halves of that sentence are true, and a firm that reads only the first has misjudged the decade ahead.
X. The pattern, at every scale
One pattern is worth naming, because it repeats at every level you look. Within a single matter, commodity tasks fall to near-zero cost while a few moments of frontier judgment carry the value. Within a firm, a commodity practice and a frontier practice diverge in their economics. Across the market, the floor and the ceiling pull apart on the same logic. And the pricing is self-similar all the way down: cost decouples from value, and price migrates from effort toward outcome, wherever you look. The billable hour is not dying so much as splitting, fixed and per-outcome below, a reasoning budget above, and the firms that recognize the same shape repeating, and position themselves at the scarce, governed top of it rather than the abundant, commoditizing bottom, are the ones that will define the profession’s next decade.
