essays & teardowns

Writing

One real cost pattern at a time, with numbers and primary sources. New issues weekly via the newsletter.

Essay 001 · August 2026

Why I don't do pay-from-savings — and what I do instead

Almost every firm in cloud and AI cost optimization makes the same offer: we only get paid if you save. A percentage of savings, no risk, what's not to like? It sounds unbeatable. I've decided to build my practice on the opposite model — fixed fees, agreed before we start — and I want to explain why, because the reasoning tells you a lot about how this industry actually works.

Pay-from-savings selects for shallow work

A contingency firm is paid per dollar of attributable savings, as fast as possible. So it rationally optimizes for the savings that are easiest to find and easiest to attribute: unused instances, commitment coverage, storage tiers. Those are real — and your cloud provider's free tooling already surfaces most of them. What a contingency fee structurally cannot reward is the hard 30%: the architecture that generates waste by design, the AI workload running a frontier model where a distilled one would do, the agent pipeline whose context grows quadratically, the organizational habits that regenerate every optimization within two quarters. None of that fits a savings-attribution formula. So it doesn't get done.

The attribution fight is built into the contract

Every gainshare deal eventually reaches the same conversation: what was the baseline? Would the provider's own recommendations have caught this anyway? Your usage grew — what's the counterfactual bill? Whose idea was the fix your own engineer implemented? These are not edge cases; they are the predictable endgame of tying fees to a number both parties have incentives to define differently. The most respected practitioners in cloud economics have refused percentage-of-savings work for years for precisely this reason. I watched a decade of enterprise cloud deals from the inside at AWS; the pattern is not subtle.

Sophisticated buyers already know

If you manage $10M of cloud spend, you know nothing is free. "No savings, no fee" reads either as desperation or as a signal that the work will be a shallow scan followed by an invoice for things you could have clicked yourself. A fixed fee, quoted up front, is the opposite signal: I know what this work is worth, I know what it takes, and I'm accountable for delivering it regardless of how the attribution argument would have gone.

What I do instead

Firms that get paid from savings are incentivized to find the easy 10% and leave. I charge a fixed fee to find the hard 30% — and to teach your team so it never comes back.

Concretely: a fixed-fee diagnostic with every finding quantified in dollars and ranked, whether or not it flatters the fee; recommendations that include the ones where the answer is "your provider's free tool does this — turn it on"; and capability transfer as a deliverable, because a cost function that depends on my continued presence is a failure of the engagement. Where automated, machine-measurable optimization genuinely fits — commitment management is the classic case — I'll point you to the software that does it well and takes its percentage honestly, because that's the one place gainshare pricing and reality agree.

One more thing this model buys, and it's the part I care about most: independence. I never have to inflate a baseline, claim credit for your engineer's idea, or steer you away from a fix because it's hard to attribute. The fee was settled before we started. The only thing left to optimize is your bill.

— Hussain Sehorewala. I advise companies on cloud & AI costs at fixed fees; here's how that works. Disagree with any of this? Tell me why — the best outcome of publishing a position is finding out where it's wrong.

In the pipeline

NEXT
The AI bill, from electrons to your P&Lthe cost cascade in one essay (Curriculum M0 artifact)
SOON
What a GPU actually costs per hour — the spreadsheet nobody publisheswith the open TCO model (M1)
SOON
I measured what a million tokens actually costsa weekend, an H100, and vLLM (M4)
SOON
The actuarial problem inside your AI agent30x run-to-run cost variance, measured (M6)

The Bookshelf

The five physical books worth owning in a field that doesn't have its own book yet — one per layer of the stack.

01
Cloud FinOps, 2nd ed. — Storment & Fuller (O'Reilly, 2023)
The canon of the discipline this field extends. Read it to learn it — and to study the seams, because it barely mentions tokens.
02
AI Engineering — Chip Huyen (O'Reilly, 2025)
The best modern application-layer textbook: model selection, evals, inference optimization. An engineer's book, in an engineer's language.
03
Chip War — Chris Miller (Scribner, 2022)
The silicon layer as narrative: TSMC, fab economics, why compute is geopolitically scarce. The history under every GPU price.
04
Large Language Model-Based Solutions — Shreyas Subramanian (Wiley, 2024)
The only book in print explicitly about LLM cost-effectiveness. It predates the agent era — its gaps are the field's open questions.
05
Hands-On LLM Serving and Optimization — Wang & Hu (O'Reilly, 2026)
Serving systems in book form: batching, quantization, deployment economics. Pair with the weekend-H100 lab.

And a sixth you assemble yourself: print the canonical papers — Kipply, Chinchilla, the Epoch economics papers, Mooncake, DistServe, FrugalGPT, the FinOps AI working-group papers — into a binder and annotate by hand. In a field this young, the primary literature is the textbook.