Software moved its cost inside the product. Almost no one is measuring it.
Margin is the economic control layer for AI-agent spend. We price every agent operation against the outcome it produced — a merged PR, a passed test, a resolved ticket — and, once a person approves it, hold the cheaper path only while it keeps quality.
This page is the thesis, not a pitch. Every stat on it is sourced, and none of it is ours; the numbers Margin itself reports trace to committed runs you could reproduce.
the shift
For twenty years, serving one more customer cost almost nothing. That assumption broke.
SaaS earned 75–85% gross margins because another user cost almost nothing. AI-native companies run a model on every request, and that cost lands in cost of goods sold. Fig. 01 draws the analysts’ reading of the shift.
Analysts named the mechanism the OpEx-to-COGS inversion; the bars plot their figures, not ours. The more AI-native a company becomes, the more its profit is simply how efficiently it runs its agents.
- Classic SaaS
- ~80% gross margin · ~20% cost
- AI-native
- ~52% gross margin · ~48% compute, now in COGS
Source: ICONIQ, 2026 State of AI (Bessemer puts the margin band 50–60%); Cloud Capital, Cost of Compute 2026; PwC, 2026 Global CEO Survey. The backdrop, not our meter; every figure Margin itself reports traces to a committed artifact.
why now
AI has a unit for how much you spend. It has none for whether the spend paid off.
Measurement is where this category starts. The defensible ground is acting on the measurement, and undoing the action the moment it stops paying.
Tokens measure how much you used. The task, a second unit Benchmark’s Everett Randle laid out in July 2026 (a marketplace past a $2B run-rate), measures how much you improved a model. The next unit counts what the spend produced: cost per outcome, the number we meter, neutrally, for the lean teams the enterprise players are not built to serve.
One unit sits past all three, and nobody owns it. Cost per outcome is a counter: a 99%-accurate extraction and an 85% one score as one outcome each, so it can never tell you a costlier route earns its bill. That takes what an outcome is worth, a judgment only you can make.
Cost per outcome is the honest default of a value-per-outcome model. No customer has priced an outcome yet, so every figure Margin computes is cost per outcome, and it is labelled as one. Price yours and the meter records it beside the cost, every outcome counted once and unweighted. A value you did not supply is never invented.
what we believe
Three commitments the product is built to keep.
Neutral by construction.
We sell no tokens and resell no inference, so “cheaper” is never an argument we have a reason to lose. Most other tools that touch this spend make money on the tokens flowing through them.
A number you can’t reproduce isn’t a number.
Every figure in the console and on this site traces to a committed run you could rerun. Sample data is labelled as sample data, everywhere, with no exception for a better-looking screenshot.
A person ratifies the loop.
The system proposes a change and proves it first; a human approves it; it reverts itself the moment quality drops below the bar. Autonomy is a dial you set, not a default we assume.
the harder path
We could have shipped a router that guesses. We built the gate that checks.
Measuring is a weekend. Standing behind the measurement enough to act on it, and to reverse the act, is the year. We spent the year.
The quick version of this product sends every call to a cheaper model and reports the saving; nobody checks the cheaper model still did the job. We made the opposite bet, and it lives in three checks:
- Five tests fail the build if a claimed saving can’t be reproduced from a committed run, and a cost win that dropped the pass-rate fails the same way.
- Provenance reads fail-closed. A metric is real or it does not render; no hand-typed number reaches you wearing a
provenlabel. NOT COMPUTEDis a shipped answer. When the evidence isn’t there, the output is the absence of a number, not a plausible one drawn to fill the space.
why the name
Margin is a trading-floor word, and the gate works like one.
On Chicago’s futures exchanges, margin is the collateral that keeps a position open. Fall below the maintenance level and you get a margin call; fail to answer it and the position is closed for you. The clearinghouse that enforces this sits between buyer and seller and trades for neither, which is why both sides trust its books.
The parity gate is the same arrangement for a model route. A cheaper route stays on while its quality holds at 0.97 of the baseline’s. Below that is the margin call, and the automatic revert is the forced close. Margin does not run your inference, so like the clearinghouse it has no side in the trade it reports.
who, and how far along
One founder, pre-revenue, looking for its first three design partners.
Nothing on this site is a customer’s number yet. When one is, it will say so.
Margin is built in Chicago by Subh Mukherjee. There are no paying customers today and no revenue. What exists is the product, the tests that keep its numbers honest, and real metered runs of open-source agents we operate ourselves, which is where every live figure on the console comes from.
The next step is three teams running agents in production who will meter them through Margin for a pilot, so the first numbers from outside our own estate are theirs. If that is you, write to subh@trymargin.io, or find me on X at @subh_mukherjee.
Watch the loop run on real spend.
The console shows the loop on open-source agents we metered ourselves: measured spend, recorded face-offs and the gate’s verdicts, every number tagged real or sample.