Proxy Station
Sign Up
All guides

We audited our own list prices against the providers, and two were wrong

The number a discount is measured against is supplied by the party claiming the discount. We read ours against the source and published what we found.

Proxy StationPublished 10 min read

Every discount is a ratio. The numerator is what we charge, which we know exactly. The denominator is what the provider charges, which we copy from somewhere — and until we checked, we could not have told you where from or when.

On 2026-08-27 we read our catalogue’s list-price column against the providers’ own documentation, model by model. It was wrong in both directions. This is what we found, how we found it, and what we changed.

Why the denominator is the soft number

The upstream software that runs this station ships an official_input_price and official_output_price for each model. It is convenient: render the comparison straight from the catalogue and every model gets a discount badge automatically.

The problem is provenance. Those numbers arrive with the catalogue. Nobody at our end typed them, nobody dated them, and nothing expires them. They are presented with exactly the same confidence as our own prices, which we do know are correct, and the visual result is a single table where half the numbers are audited and half are inherited.

The method

Deliberately unautomated, because the failure being looked for is precisely the kind an automated fetch reproduces:

  1. Take each model with a list-price column.
  2. Open the provider’s own pricing page — not an aggregator, not a comparison site, not a cached copy.
  3. Read the input and output price for that exact model name and variant.
  4. Compare with the catalogue value, and record the date it was read.

The four pages: OpenAI, Anthropic, Google, xAI. All four publish per-million-token prices openly, so this is work anyone can repeat.

What we found

Two errors, in opposite directions, and they are instructive precisely because one flattered us and one did not.

A row that overstated our discount

gpt-5.4-mini carried a list price of 2.50 in / 15.00 out per million tokens — byte-identical to the row for gpt-5.4, its larger sibling. OpenAI’s own page gives 0.75 / 4.50 for the mini.

Our own sell price for that model was correctly lower than the sibling’s, so nothing looked wrong on screen. The comparison simply ran against the wrong reference. The row advertised a 98% discount where the real figure is 93%. Both are large. One of them is true.

A row that understated it

deepseek-v4-flash carried 0.14 / 0.28. DeepSeek publishes 0.22 off-peak and 0.44 peak for input. The catalogue was quoting a figure below the provider’s own cheapest rate, which made our discount look smaller than it is.

That direction matters for the argument. An error that only ever flattered us would suggest the column had been tuned. Errors in both directions suggest something duller and more useful to know: nobody was checking it at all.

What we changed

We stopped using the inherited column for public claims and built a separate, hand-read reference table. Its rules:

  • A price enters only after a person has read it on the provider’s own page.
  • Every entry carries the date it was read and a link to the page it came from.
  • Entries expire after 60 days. Past that, the row drops out of the comparison by itself rather than appearing with a number nobody has re-checked.
  • A model with no verified entry is not shown in the comparison at all, even though its inherited number would have produced a perfectly plausible badge.

The expiry is the part that costs us something, and it is the part we would keep if we could only keep one. A comparison that quietly ages is worse than one that visibly shrinks.

What that costs, in numbers

Models in the catalogue
53 (2026-08-30)
Carrying an inherited list price
44
Hand-verified, dated, sourced
10
Verified entries expire
2026-10-26

Forty-four rows could carry a discount badge. Ten do, where we make a public claim. That gap is the honest cost of the rule, and it is visible on the page rather than described here: the comparison is shorter than it could be.

How to audit any relay’s comparison

The same method works from the outside, and it takes about ten minutes:

  1. Pick three models from the relay’s comparison — one flagship, one mid-tier, one small.
  2. Look up each provider’s own published price for that exact variant.
  3. Recompute the advertised discount yourself.
  4. Check whether the relay says when its reference was read. If there is no date, the number has no shelf life and no accountability.

The small model is where errors concentrate. Mini and flash variants are the ones most often given their larger sibling’s list price, because that is the row a careless import duplicates — which is exactly the shape of the error we found in our own.

Two further things worth knowing before you compare anything: the units differ between sources and mixing them is a thousand-fold error, and list prices carry conditions that a single number hides.

Why we did not automate it

The obvious improvement is a scraper that reads each provider’s pricing page nightly. We considered it and decided against it, for reasons worth stating because they generalise.

  • The failure being prevented is unattended data. Replacing an unread column with an unread scraper reproduces the problem with more moving parts.
  • Pricing pages are prose, not tables. Conditions live in footnotes, tier thresholds in sentences, peak windows in a paragraph. A scraper reads the headline number and drops precisely the conditions that make it correct.
  • Silent breakage. A layout change turns a scraper into a source of confidently wrong figures, and the failure looks exactly like success.
  • The volume does not justify it. Ten models, four pages, sixty days. That is under an hour of reading per cycle.

A scheduled reminder to re-read ten rows beats a nightly job that nobody checks. Where automation does help is in enforcing the rule rather than gathering the data: the expiry is code, so a row that nobody re-reads removes itself.

What we would want from anyone else’s comparison

The same four properties, and they are cheap to check on any relay’s page:

  1. A date beside every reference price. Without one the figure has no shelf life.
  2. A link to the provider’s own page, not to an aggregator and not to the relay’s own table.
  3. The exact model variant named. Mini and flash siblings are where errors concentrate — including ours.
  4. An expiry, or at least a review cadence. Something that makes an unread number disappear rather than persist.

A comparison with all four is checkable in ten minutes. One with none of them is a marketing claim wearing a table’s clothing, and the distinction is not visible from the design.

This is the pricing half of the guide to how relays work and how to audit one. The routing half — whether the model you paid for is the one that answered — is a separate check, and it is the one that matters more.

Figures in this guide were read on the dates shown beside them. Prices change; where a claim depends on a provider’s published price, the link goes to that provider’s own page so you can check it rather than take ours. This guide is reviewed by 2026-10-26.

Check the numbers yourself

Every model on this station, its per-token price and the provider’s published list price are on the pricing page, with no account required to read them.