The maintenance log
Changes
This site's entire claim is that it's maintained and shows you when. That's easy to say and easy to fake, so here is the log: everything corrected, added or re-checked, dated, including the things we got wrong.
Last updated · Every page also carries its own pair of dates at the top · Atom feed of everything below
Why this page exists
A date stamp on a page tells you when it was checked. It doesn't tell you whether anything actually happens when a fact goes stale. This log is the difference between a promise and a record, and it's the one page here you can use to judge whether the rest is worth trusting.
How to read the two dates#
Every page here carries Published and Last verified. They are different claims, and the gap between them is the only part that tells you anything.
- Published is when the page first went up. It never changes.
- Last verified is the most recent date someone opened the sources and checked that the page is still true. It moves only when that actually happens.
- When the two dates match, as they do on most of this site today, it means the page has not been re-checked since it was written. That is worth knowing rather than hiding: a brand-new page is unproven, not proven.
Editing a page does not move its verification date. Fixing a typo, cutting a repeated phrase or rewording a heading changes nothing about whether the facts still hold, and a site that bumped its dates for cosmetic edits would be doing the thing this site warns about: an updated stamp with no substance behind it. Verification dates move when sources get re-read, and the re-reading gets logged below.
Checks that are due#
Scheduled, not aspirational, and with the dates written down. A commitment to check something "monthly" cannot be audited from outside, which makes it the same unfalsifiable freshness signal this site warns readers about. So each check below carries the date it is next due. If one of those dates passes and nothing appears in the log, the site has broken its own promise and you should discount it accordingly.
- Model facts, monthly. Next due 11 October 2026. Prices, model IDs and limits, checked against each vendor's own documentation. The fastest-rotting page on the site by a wide margin, and the only thing here checked on a clock. Since 26 August a script also diffs every figure on that page against all seven vendor pages daily, so a change that used to sit undetected for up to a month now surfaces in a day. It only reports; nothing on this site is updated automatically.
- The home page, monthly. Next due 11 October 2026. It opens with a claim about what search returns for “learn AI”, which is the one thing on that page nobody can check from the page itself, and it summarises every other page here, so it goes wrong whenever the site changes shape rather than when a fact does. It carried a review date of its own from the day it was written and was never listed here, which made that promise the unauditable kind this section exists to avoid. Listed on 30 August 2026, keeping the date it had already promised rather than taking a later one.
- The concept pages and the guides, quarterly. Next due 11 November 2026. They're written to avoid per-model figures precisely so they age slowly, so this is a lighter pass: does the explanation still describe how these systems actually work, and is each page still citing the right document? What decides it is the Sources list rather than the page count — the check is re-reading those documents and seeing which pages move, and every document a guide depends on is already on that list.
- The resource lists on Start here, quarterly. Next due 11 November 2026. Courses get retired, reorganised, and quietly moved behind a payment page. Every link gets opened, not just pinged.
- Whenever a page is added or substantially changed, no date attached. A new page can make an old page wrong without touching a single fact, by turning "this is planned" into a false statement or leaving a cross-reference pointing at nothing. That check is triggered by the event, not the calendar, because that is when the rot happens.
-
End of August 2026 — introductory pricing.Closed early, 11 August 2026, and it found a correction. See the log below. The promotion that was scheduled to expire was instead made permanent, so there is nothing left to re-check at the end of the month.
The log#
23 September 2026 — a sentence on Sources that no longer matched the table under it#
The date line at the top of Sources said that seven of the documents on it had been re-read in full for the Model facts check on 1 September 2026. No row in the table below it says 1 September any more; ten now say 23 September. A reader checking that sentence against the table, which is what the page is for, would have found nothing it described. It now says what the last column means, the day each document was last read in full, instead of counting rows as of one date, so it cannot fall behind the table again. No source or figure changed, and the page's verified date stays where it is, because nothing was re-checked.
Live. Deployed deliberately on 23 September 2026, the same day.
23 September 2026 — one guide re-read early, because the model its caveat is about has been replaced#
Re-read early, for a reason. Getting better results quotes one exception from Anthropic's prompting guidance: that Claude Opus 5 checks its own work without being asked, so telling it to can cost tokens and latency for nothing. The monthly check below replaced Opus 5 on Model facts today, which made this the page on the site most likely to have gone stale. It belongs to the quarterly check due on 11 November, and there was no reason to leave it seven weeks.
Every line it quotes is still there, word for word. The guidance was opened today and read in full. The line calling examples one of the most reliable ways to steer output, the “brilliant but new employee” framing, the colleague test, the advice to have the model check its answer against named criteria, and the Opus 5 exception are all unchanged. The exception still names Opus 5 and no other model.
So the caveat stands as written. It was dated from the start, quoting “the version fetched on 12 August 2026” and calling Opus 5 a model “current at that date”, so that a newer model would not make it wrong, and it hasn't. The page's last verified date moves to today, and so does the guidance's entry on Sources. The other concept pages and guides are still due on 11 November, and that date stays where it is under Checks that are due.
Live. Prepared on a branch, reviewed, and deployed deliberately on 23 September 2026, the same day.
23 September 2026 — the monthly check, run early, because three rows had stopped being current#
Run early, and for a reason. The 15 September change below was still waiting to be deployed, and before it went out the price watcher was run by hand to make sure nothing had moved underneath it. It could not verify three rows: Claude Opus 5, GPT-5.6 Sol and GPT-5.6 Terra. None of their figures had changed. Their vendors had stopped listing them as current. Deciding what replaces a row is the part of the check a script cannot do, so the whole monthly check ran today rather than on 11 October. The next one is still due on 11 October.
Claude Opus 5.5 replaces Claude Opus 5, and costs less. Anthropic's models overview now puts Opus 5.5 in its headline table and lists Opus 5 among legacy models that are still available. Opus 5.5 is $4 input and $20 output against Opus 5's $5 and $25, with the same 1M context and 128k output, and both its cutoffs are June 2026 against May. Opus 5 is still active, with retirement not sooner than 24 July 2027, at the price this table gave for it.
The OpenAI rows are now GPT-6 Astra, Sol and Luna, and the name Sol
now means a different tier. OpenAI's model catalogue lists those
three as its flagship models, and GPT-5.6 Sol and Terra have gone from it
and from the standard price table. GPT-5.6 Sol was a flagship at $4 and $20.
GPT-6 Sol is the model OpenAI says to choose “to balance intelligence
and cost”, which is how it described Terra, at $2 and $10. Anyone
switching from gpt-5.6-sol to gpt-6-sol because
the names match is moving to the middle of the new lineup, not its top; the
top is GPT-6 Astra. GPT-6 Luna, at $0.10 and $0.50, is now the cheapest row
on the table. All three have the same 1.05M context and 128k output, and
knowledge cutoffs of 30 April, 20 April and 18 May 2026. Neither GPT-5.6
model has a shutdown date, and both still carry the prices this table gave
them.
The contradiction recorded on 15 September has settled. That entry found OpenAI calling two different models its flagship, and left the word off the page. The catalogue now calls GPT-6 Astra “our flagship model” and no longer lists GPT-5.6 Sol at all. Sol's own model page still calls it a flagship, but that page no longer backs a row here.
Correction: a note under the table said two different numbers were the same. The note on context figures, written on 11 August, said Google's 1,048,576 and OpenAI's 1.05M were the same number written two ways. OpenAI's model pages give 1,050,000 exactly. The table still prints both as 1.05M, which is fair as rounding, and the note now says they are close, not equal. Nothing anyone could budget from depends on the difference. It is logged because it was wrong, and because it sat directly under the figures it was there to explain.
The watcher had been passing a row it could not see. When
GPT-5.6 Sol left OpenAI's standard price table, prices.py
cross-checked it against the pricing page, found the same ID further down in
a table of cyber-security models, read $4 and $20 there, and reported the
price unchanged. Its other OpenAI probe did flag the row, so the run still
failed, but one line of it said all was well about a model the vendor no
longer listed as a standard offering. It now reads only the standard table,
and a model missing from it is reported as unverified.
Everything else was re-read and holds. Fable 5.1, Sonnet 5 and Haiku 4.5; GPT-6 Astra on both OpenAI pages; Gemini 3.8 Flash, with the rate change on 1 January 2027 in the same words; Muse Spark 1.2 and the contributor tier; and Meta still publishing no Llama price. The long-context note now carries the GPT-6 tiers. The note on end dates has one left, Gemini's, because GPT-5.6 Sol's promotional rate went with its row. Anthropic's and OpenAI's deprecations pages are now listed on Sources, because Model facts now tells you, of four models it no longer lists, that they are still on sale, and that claim should be as checkable as a price.
The daily check had stopped, and nothing said so. The price watcher described under Checks that are due ran on a schedule kept on one machine. When work on this site moved to another machine in September, the schedule did not move with it. So for part of this month, the daily check that section promises was not happening. Exactly how long is not known, because the record of its last run was on the old machine. It is scheduled again from today, with the same rules as before: a missed day is reported, not quietly made up. The setup is now written down beside the code, so the next move cannot lose it the same way without someone skipping a documented step.
Live. Prepared on a branch together with the 15 September change below, reviewed, and deployed deliberately on 23 September 2026.
15 September 2026 — two rows on Model facts, one of which changed no numbers at all#
The monthly check earlier today noted four models published by vendors that
Model facts does not list, and left them for a
human rather than deciding on its own. Two have been acted on. Claude Mythos
5.1 stays out, because Anthropic lists it as limited availability. Meta's
muse-spark-1.3 stays out too: it shares the standard pricing
already printed for 1.2, so adding it would grow the table without telling a
reader anything.
Gemini 3.6 Flash became Gemini 3.8 Flash, and not one figure moved. Same 1,048,576 context limit, same 65,536 output limit, same $0.75 and $3.75, same increase on 1 January 2027 worded the same way, and still no published knowledge cutoff. Only the name and the API ID differ. The reason to swap it has nothing to do with the numbers: Google's pricing page now calls 3.6 Flash “our previous generation Flash model”, and this table says it lists current ones. A row can go stale on that claim while every figure in it stays correct, and the daily watcher cannot see it, because the watcher compares numbers and this is a claim about which model matters.
GPT-6 Astra is new, and replaces nothing. OpenAI documents it as the model to start with, above the GPT-5.6 family, which stays where it is. Its limits are identical to Sol's — 1.05M context, 128k max output — so the whole difference is money and recency: $10 and $50 against Sol's $4 and $20, and a cutoff of 30 April 2026 against 16 February. It carries no promotional wording, so unlike Sol's rate it has no stated horizon. Its long-context tier, $20 and $75, is in the tiering note with the others.
OpenAI calls two different models its flagship, and this table prints neither claim. The model catalogue labels GPT-5.6 Sol “flagship model for complex professional work”; the choosing-a-model section on the same page says to use GPT-6 Astra, “our flagship model”. Both sentences are live right now. Picking one would be settling something the vendor has not settled, and a tidier story than the source supports is the exact failure this site is arranged against. So the contradiction is recorded and the word is left off the page.
The watcher was moved with the table, which is the part that could
have rotted quietly. prices.py holds one Google model
page in its source list and derives the model it expects from that URL, so a
table row and a source that disagree make the row report
UNVERIFIED rather than passing unchecked. Both the URL and the
rate-expiry wording it watches now name 3.8 Flash.
Sources records the same move and why the two are
tied together.
Overtaken before it went out. This entry was written on 15 September and not deployed that day, and before it could be, OpenAI moved the GPT-5.6 models off its catalogue. So the GPT-5.6 family no longer “stays where it is”, and the flagship contradiction above has settled. The 23 September entry above says what changed. This one is left as it was written, because the log records what was done and known on the day.
Live. Prepared on a branch on 15 September, and deployed deliberately on 23 September 2026, eight days later, together with the entry above.
15 September 2026 — the home page check, four days late, and the first sentence is drifting again#
Due 11 September, run on the 15th. The machine this site is maintained from failed and went in for repair; the check was run on its replacement. Four days late is a broken promise by the standard set above, so it is recorded rather than absorbed.
Correction: “None of it is disinterested” is no longer true, and has been removed. On 30 August this page's opening claim was rewritten after a search found that everything on the first page for “learn AI” was first-party — a lab's academy, two platform catalogues, a nonprofit curriculum, two vendor training hubs, a platform blog. Not one was disinterested, and the sentence said so. Running the same search today returns something different: among the first twelve results are a Reddit thread, a Quora answer and a Microsoft Q&A question. Nobody is selling anything in those. They are thin and unsourced, which is a fair thing to say about them, but “none of it is disinterested” is not, and this site does not get to keep a sentence that is merely rhetorically convenient.
The complaint survives the correction, which is again why the sentence was
rewritten rather than cut. Almost every result is still a course catalogue or
a vendor academy — Google's own learning hub ranks first, with
grow.google, Coursera, Codecademy, Microsoft's curriculum and an
Intuit blog behind it. What changed is the absolute. The line now names the
forum threads and moves the weight onto the claim that holds for every result
on that page, commercial or not: none of them tell you whether what they say
still holds. That is the thing this site actually does differently, and it is
a stronger argument than the one it replaces because it is true of all of
them rather than most.
This is the second time in six weeks that the first sentence on the site has needed correcting, and both times for the same reason: it asserts something about the outside world that cannot be checked from the page, and it was written to land rather than to survive. The 30 August entry called that “the same species of error this site spends its time catching in other people”. The species has not changed; only the specimen. A claim about search results has no business being stated as an absolute by a site whose whole argument is that unverified claims rot.
A caveat on the method, which the 30 August check should also have carried. Search results differ by person, place and day. This was one query, run once, from a United States location, through a search API rather than a browser with a search history attached. It is evidence about what that page looks like; it is not a census. The honest version of this check is that a sentence stated as universal cannot be verified by a method that samples, which is an argument for the sentence being less absolute rather than for the check being more elaborate.
Everything else the page claims about this site was checked and holds. Three tracks, all resolving. Ten concept pages listed and ten on the site, with no drift either way. Every page the site has — all twenty-seven, including the eight guides named in the prose — is linked from the home page, cross-checked against Guides, which links nothing the home page misses. No affiliate links anywhere, no sponsors, no newsletter box. The two dates, the orange flagging and the single-page rule for figures are all still accurate descriptions of how the site works.
The last-verified date moves to today, because the search was actually run and the site was actually walked. The next review is 11 October 2026.
Live. Prepared on a branch, reviewed, and deployed deliberately on 15 September 2026.
15 September 2026 — the monthly check, four days late, and the watcher stopping a correction that would have been wrong#
This check was due on 11 September and ran on the 15th. The machine it runs on stopped working and went in for repair, taking the daily watcher with it; the check happened on a replacement once the site was cloned back onto it. Four days is four days, and the section above says plainly that a missed date means the promise was broken, so it is recorded here rather than quietly absorbed. The daily watcher was also dark for that stretch. It was run before this check and reported 52 figures unchanged, so nothing moved unseen while it was off — but a watcher nobody is running is not a watcher, and its silence proved nothing at the time.
The check tried to correct a figure that was already right, and the daily watcher stopped it. This is the most useful thing that happened today, so it goes first and in full.
Anthropic's models overview page renders a comparison table showing Reliable knowledge cutoff and folds the rest away behind a “Show all details” control. Read as a web page, it appears to publish one cutoff per model and no training cutoff at all. Working from that, and from the Transparency Hub — which gives Claude Haiku 4.5 a single “knowledge cutoff date of February 2025” and no separate training date — this check concluded that Haiku's training cutoff of July 2025 had no source behind it, that July 2025 was probably Claude Sonnet 4.5's date misread off a neighbouring row, and that the cell should read February 2025. It was edited to say so.
That was wrong. Appending .md to the same URL returns the
complete table, including a Training data cutoff row, and
that row gives Haiku 4.5 Jul 2025 — exactly what this
table has said all along. The figure was correct, the five-month gap is
real, and the reasoning that replaced it was confident, plausible and
entirely mistaken. prices.py reads the markdown variant, so
when it was run after the edit it immediately reported
CHANGED — Claude Haiku 4.5 training cutoff Feb 2025 -> Jul
2025: the table had moved away from the vendor, not towards it. The
edit has been reverted and the cell reads Jul 2025.
Three things are worth taking from that. A rendered page and its
markdown source are different documents, and a figure behind a
disclosure control is invisible to anyone who does not know to open it; this
page is now checked against overview.md and says so. An
absence of evidence read as evidence of absence is how a right figure gets
overwritten — two sources failing to show a number is not the
same as a source contradicting it, and the rule in
MONTHLY-CHECK.md that says an unverified cell is fine exists
precisely so that gap does not have to be filled by inference. And
the watcher earned its keep in the direction nobody designed it
for: it was built to catch the vendor moving under a stale table,
and today it caught the table moving away from an accurate vendor page.
GPT-5.6 Sol's $4 / $20 is promotional, and this page never said so. OpenAI's pricing page carries a line that was not there when this table took the $4 / $20 figure on 26 August: “GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026.” The figure is unchanged and correct. What was missing is that it has a stated horizon, which is exactly the kind of thing this page exists to carry. It now sits in an open block beside Gemini 3.6 Flash's 1 January 2027 increase, the only other dated price on the page. Note the difference in wording: Google states when its rate ends, OpenAI only when its rate is guaranteed until. Those are not the same promise and the block says so.
A sentence went wrong without any figure moving. The API IDs
note said Claude Haiku 4.5 was “the only row where the distinction
” between a pinned ID and an alias shows. OpenAI's GPT-5.6 Sol page
states that “the gpt-5.6 alias routes requests to GPT-5.6
Sol”, with gpt-5.6-sol as the default snapshot. That
is not a change on OpenAI's side that this page failed to catch — it
was simply never true of the OpenAI rows, and the daily watcher cannot see it
because it compares figures and this is prose. Corrected, and left visible
rather than tidied away.
Sources now says how to read the Anthropic
page. The row for the models overview carries the instruction that
cost this check an hour: read it as overview.md, because the
rendered table hides the training cutoff. Model
facts carries the same warning beside its verification note. Neither is
a new source — it is the same document, read the way the watcher has
always read it, which is the part that was not written down anywhere until
today.
Seen on the vendor pages, not added to the table. Four
models are now published that this table does not list: OpenAI's
gpt-6-astra, which its own docs call the flagship and recommend
as the default starting point; Google's Gemini 3.8 Flash, alongside which the
pricing page now describes Gemini 3.6 Flash as “our previous generation
Flash model”; Anthropic's Claude Mythos 5.1, which is limited
availability; and Meta's muse-spark-1.3, which shares the
standard pricing already printed here. The Gemini one matters most: this
table's stated selection is current general-purpose models, and 3.6 Flash is
now described by Google as the previous generation. Adding or moving a row is
a decision about what this table is for, so all four are noted here and left
for a human rather than taken by a check.
Every figure was unchanged. All 52 were read against the vendor's own page today, and none of them had moved. Anthropic's four rows, OpenAI's two, Google's one and Meta's two match, including both long-context tiers, the Gemini expiry wording, Google's exact 1,048,576 and 65,536 token limits, Meta's contributor tier, and the cells that read “not published” on purpose. One caveat could not be re-verified in its own words: the 11 August correction quotes Anthropic's pricing page on the Sonnet 5 rate becoming permanent, and that language is no longer on the page. The substance holds — it is 15 September and Sonnet 5 is still $2 / $10, so the increase did not occur — but a reader following that link today will not find the sentence in quotation marks. The note is dated and carries its fetch date, so it has been left standing rather than rewritten.
Still outstanding: the home page check, also due 11 September, has not been run and its date above has deliberately not been moved. Moving it would claim a check that did not happen, which is the exact failure this section is built to make visible.
Live. This check ran unattended under the rule in
MONTHLY-CHECK.md that says to prepare everything and stop, so it
stopped. The corrections above were reviewed afterwards and deployed
deliberately on 15 September 2026.
1 September 2026 — the price watcher went blind on one row, which is the thing it was built to do#
The daily watcher reported UNVERIFIED against Claude Fable 5
this morning and exited non-zero: model no longer on the vendor's
table. That is not a price moving. Anthropic has released Claude Fable
5.1 and moved Fable 5 off the headline table it publishes, and the watcher
reads that table by column heading, so the row it had been checking every
day was suddenly not there to check. It could have said “unchanged
” — every figure it had last seen was, after all, still the
figure on file — and that is exactly the quiet blindness it was
written to refuse. It refused, and said which row and why.
Model facts now lists Claude Fable
5.1 where it listed Claude Fable 5: claude-fable-5-1,
1M context, 128k max output, $10 and $50 per million tokens, and both
knowledge cutoffs at June 2026 rather than January. The prices are
identical to the row it replaces. The cutoffs are five months further
forward, which is the whole of the difference this table can show you.
Nothing here was wrong, and that is worth saying plainly.
Every figure the old row carried still matches Anthropic's own page for
Fable 5 today. The model is not retired — Anthropic lists it as
active, with a retirement date no sooner than 9 June 2027 — it has
become a legacy model, and this table has never carried the other six
legacy Claude models either. The row moved because the table follows the
vendor's current lineup, not because a number was ever false. If you are
still calling claude-fable-5, it works and it costs what it
cost. The reasoning is kept on the page itself rather than only here.
Because a figure on that page changed, the sources were re-read rather than assumed: all seven vendor documents behind Model facts were opened today, and every other figure in both tables is unchanged. Anthropic's Opus 5, Sonnet 5 and Haiku 4.5 rows, OpenAI's two GPT-5.6 rows and their 16 February 2026 cutoff, Gemini 3.6 Flash at $0.75 and $3.75 with its stated expiry on 1 January 2027, and Meta's Muse Spark 1.2 all hold. So do the caveats around them, which rot faster than the numbers and were checked as text: the Sonnet 5 introductory rate is still documented as having become permanent, and OpenAI's long-context tier is still double on input. Last verified on Model facts moves to today, and the per-document dates on Sources move for those seven documents and no others.
One thing found and deliberately not acted on: cache reads on Fable 5.1 cost a quarter of what they cost on Fable 5, which makes Anthropic's general “a cache hit costs 10% of the input price” rule untrue for this one model. This site quotes no caching figure — it says caching changes the numbers substantially and sends you to the vendor's page — so there is nothing here to correct. It is recorded because a check that only reports what it changed is not a check anyone can audit.
The monthly Model facts check stays due on 11 September 2026. Doing this early does not buy a later deadline, for the same reason the home page kept its September date on 30 August. All of it is live.
30 August 2026 — three pages that were missing, two of them because they were hard to write honestly#
How to choose an assistant closes a real dead end. Start here has been telling the curious reader to use one deliberately for a week and then naming no assistant at all, which is advice with a hole in the middle of it. The page gives six criteria you can check yourself in a few minutes, each one traceable to a mechanism explained elsewhere here, and no product names whatsoever. The reason for that is on the page: a recommendation is a product name attached to a price, this site's method for anything of that shape is to cite a vendor page and diff against it every morning, and there is no vendor page for “which one is best”. It would be the one data-bearing thing here with nothing to check it against.
What happens to what you type answers a top-three beginner question that was entirely absent. It is written as mechanism — your text goes to somebody else's computer, and everything else follows — and it separates the one question people ask into the three that actually have answers: retention, training, and deletion. It gives no per-vendor figures at all, and says why in a section of its own. Those settings change without announcement and differ by account type and country, so a page here stating them would be confidently wrong in a way no reader could detect. That is the same reasoning that keeps every other figure on Model facts and nowhere else.
Common myths, and where each one is answered was the one worth being most careful about. A myth-busting page is a page of confident claims about what these systems cannot do, and this site's own page on staleness argues that such a sentence has the shortest shelf life in the field — so a page built from them would rot fastest here while still reading as authoritative. It is therefore a router and not an argument: sixteen beliefs, each a link to the page that explains the mechanism, and nothing asserted that is not argued elsewhere with its sources attached. The rule that keeps it that way is printed on the page, because the failure mode is not writing it wrong today, it is letting claims accumulate on it later.
All three are on Guides and linked from the home page, which is the condition the structural check holds the site to. All of it is live.
30 August 2026 — the rules in the tables were too faint to be doing their job#
The hairline used everywhere on this site sits at 1.25:1 against the page. That is right for the edge of a card, where the line is decoration. It is wrong for the row separator in an eight-column table, where the line is the thing that lets you carry a figure back to the model it belongs to — that is content, and the accessibility guideline asks 3:1 of it. Darkening the hairline everywhere would have put that weight on every divider on the site, so there is now a second, stronger rule used only inside tables. It is close to what the print stylesheet has always used, so the printed table has had a separator this strong all along.
In the same pass, three text styles that were under fourteen pixels went up to fourteen. Most of the small type here is labels — a kicker, an eyebrow, a tag — which are short and conventional at that size. Three were not labels: the date stamp, which carries the two dates this whole site turns on; the notes under the tables, which carry the corrections and the pricing warning that costs money; and the column headings, which are read as often as the cells under them. Setting the site's own load-bearing element in its second-smallest size was not a defensible choice.
Behind the scenes, three things that only matter when something goes wrong. The structural check used to raise an error rather than report a problem if a page it depends on were renamed, which is the worst way for a checker to fail: it stops at the first fault and every check after it never runs. It only ever examined links beginning with a slash, so the first relative link anybody wrote would have been the one link on the site nothing was watching. And this log, now past 140KB and growing every time it does its job, will eventually have to be split by year; the feed builder and the check that guards it both used to reach for this page by name, which made that split a change to two programs as well as to a page. They now find log pages by their markings, so an archive is picked up by both without either being edited. That was rehearsed by actually performing the split, which is how a bug in the new arrangement was found: filenames sorted the wrong way and the feed came out oldest-first. All of it is live.
30 August 2026 — correction: the first sentence on this site was wrong about the thing it complains about#
The home page opened by saying that almost everything ranking for “learn AI” is an affiliate page pointing at a course. That was the justification for this site existing, it sat in the first sentence a reader ever saw, and nobody had checked it since the day it was written on 11 August. It was checked today, by running the search and opening the results, and it does not hold.
What actually ranks is first-party: a lab's own free academy, two course platforms' own catalogues, a nonprofit's free curriculum, two more vendors' own training hubs, and one platform's own blog. Not one of them is an affiliate page. The nearest thing to the original claim was a course platform's guide to learning AI, and reading it, roughly two thirds is substantive teaching with the platform's own courses promoted around it — first-party marketing, which is a different and more ordinary thing than earning a commission on somebody else's course. Several of the top results are free, and some are good.
The honest version of the complaint survives, which is why the sentence is rewritten rather than deleted: everything on that first page is published by somebody with a course to sell or a tool to promote, and none of it tells you when it was last checked. That is a fair thing to say and it is what the page now says. The original was not fair. It named a specific commercial arrangement that was not there, which is the same species of error this site spends its time catching in other people, printed at the top of its own front door for nineteen days.
Two smaller things came out of the same pass. Everything else the home page asserts about this site — three tracks, eight guides all named, eight habits, ten concept pages, no affiliate links anywhere, unverifiable figures flagged — was checked against the site itself and holds. And the home page had been carrying a review date of its own since it was written while never appearing in Checks that are due, so the one promise on it that a reader might have audited was not listed anywhere they could audit it. It is listed now, keeping the September date it had already promised rather than quietly taking a later one for having done the work early.
The last-verified date on the home page moves to today. It is the first time it has moved since the page went up, and it moves because the sources were actually opened, which is the only thing that is allowed to move it. All of it is live.
30 August 2026 — a drawing of the two dates, on the page that asks you to trust them#
The home page explained the two dates in a single dense paragraph and then asked you to take the rest of the site on the strength of it. The claim that does the most work here had the least room. There is now a diagram beside that paragraph: two timelines, one where the published and last-verified dates sit apart with the stretch between them marked, and one where they sit on the same point and there is no stretch at all.
The second timeline is the one worth having. A page whose two dates match has not been re-checked since the day it was written, which is the honest position of most of this site, including the home page the drawing sits on. Saying that in a picture rather than a clause is the point: it is the fact most easily skimmed past, and it is the one that decides how much of this site you should believe.
It follows the same rules as the four drawings already here. No JavaScript and no image file — it is markup, drawn in the same small vocabulary of shapes and colours as the others, taking its colours from the stylesheet so it follows light and dark without a second copy of the palette. The caption states the whole thing in words, because a reader who cannot see the picture must not be reading a page with a hole in it. The date on the page has not moved: adding a drawing re-verifies nothing. All of it is live.
29 August 2026 — the price watcher could have been reading the wrong model#
No figure on Model facts was wrong, and nothing on this site changed. This is a repair to the thing that checks the figures, not a correction to the figures, and it is logged because a check nobody can see is only worth what its faults are known to be.
The watcher finds a model on a vendor's page by its API identifier, then
reads the prices and limits printed next to it. It marked the end of that
identifier with a word boundary, which sounds right and is not: an
identifier ends in a letter or a digit, and the next character in a name
like gpt-5.6-sol-mini is a hyphen, so the boundary matches
perfectly well in the middle of a different model's name. Vendors list the
small variant of a model directly beside it, frequently first. On a
constructed page that does exactly that, the old pattern read the cheaper
variant's prices and compared them against the full model's row.
The loud version of that failure is a change reported that never happened, which wastes a morning. The quiet version is worse and is the reason this is written down: if the two models ever agreed on a figure, the watcher would have reported the row unchanged while looking at the wrong model entirely — a clean result from a check that had stopped checking. That is the failure this site has already had once, in a different disguise.
The second fault was plainer. For two of the four vendors the watcher took the first matching row and returned, so a second model from either would simply not have been looked at. A backstop did catch it — every row has to have been reached by some probe or the run fails — but it could only say that a row went unchecked, not why. Both now work through every row and say what they could not verify and what is missing. Where a source page covers one model and the table gains another, the new row reports itself unverified and names the page that needs adding, rather than borrowing the answer from its neighbour.
The watcher's own limit is now stated in it as well: one vendor publishes a price table that is not labelled by model, so reading a price out of it is only unambiguous while the table here carries a single priced row from that vendor. A second one would be looked up by name, and reported rather than guessed at if it could not be found. Before and after the repair, the run against all seven vendor pages is the same: fifty-two figures, none changed, none unverified. All of it is live.
29 August 2026 — the deploy now refuses to publish a page that is lying about itself#
Entries on this page say whether the change they describe is live or only prepared. It is a small promise and this site broke it earlier today: a deploy went out while an entry still read “Committed, not yet published”, so for a while this page was telling readers that something was unpublished at the same moment they were reading the published version of it. On the one page whose entire job is being the honest record.
The structural check cannot catch that, and it is worth being precise about why, because the reason is more interesting than the bug. Before a deploy, “not yet published” is true. Nothing in the file distinguishes an entry that is honestly waiting from one that was left behind, and no amount of reading the file more carefully will separate them. The only thing in the world that knows the difference is the act of deploying. So that is where the guard had to go, and until today there was no guard there at all — only the intention to remember.
There is now a deploy script, and it refuses to upload anything while a single entry still claims to be unpublished. It also re-runs the two generators, requires the structural check to pass, and requires the working tree to be committed, so that what gets served is what is in the history rather than an edit that exists nowhere else. All of that runs before anything touches the network, because a deploy cannot be taken back: the wrong page is being served the instant it uploads.
Afterwards it checks the site rather than the folder it just sent. It fetches this page back and confirms that no entry claims to be unpublished, and it rebuilds model-facts.json from the live Model facts page and compares it to the live JSON. Verifying the files you uploaded tells you what you meant to publish, which is the thing you already knew.
What it will not do is fix the wording for you. Changing “Committed, not yet published” to say the opposite is writing a sentence in the log, and no script here writes this site's sentences — the price watcher has refused to edit a figure since the day it was written, for the same reason. It stops and tells you. And it is worth saying what this does not do: it does not stop anything being deployed that should not be. It stops exactly one lie, which happens to be the one that was told.
One thing it got wrong on the first try, since this page is the place for that. The guard originally looked for the phrase “not yet published” anywhere on this page, and the first entry it refused to publish was this one — which quotes that phrase three times while explaining the guard. A check that fires on writing about the thing is a check somebody eventually switches off, so it now matches the full marker sentence including its full stop, which prose quoting the phrase does not contain. Where the two could still be confused it errs toward refusing and prints the line numbers: a minute of your attention is the cheaper mistake. All of it is live.
29 August 2026 — the table a machine can read, and the second parser that is now gone#
The two tables on Model facts are now published as model-facts.json as well as HTML. Nobody asked for it, but the alternative was worse than doing nothing, and that is the actual story here.
prices.py, the script that checks these figures against each
vendor's own page every morning, had its own parser for the tables. That
parser is the worst failure this site has had: in August it was quietly
dropping rows it could not read and finishing by reporting that the table
matched every vendor page. It was patched, and the patch was sound, but the
shape of the problem survived it — one hand-edited table, two separate
programs reading it, either of which could go blind on its own. There is now
one parser. It lives in modelfacts.py, it writes the JSON, and
prices.py reads the JSON rather than the page.
What that does not do is make the parsing safe. It is still a regular expression over HTML somebody edits by hand, and it can still be wrong. What changed is that it can only be wrong in one place, and a new structural check rebuilds the JSON from the table on every run and fails if the two disagree — so a row lost in parsing now stops a deploy instead of turning into a clean price report. That is a smaller claim than deleting the failure, and it is the true one.
The figures in the table carry their own values as attributes now, so the
JSON can be built without anything having to interpret “1.05M”.
Those attributes say what the printed cell says and not one thing more:
data-tokens="64000" on a cell reading 64k is the number 64k
means, not a claim that Google's limit is exactly 64,000 — it is
65,536, and the notes under the table have always said so. Two copies of a
figure is the arrangement this site refuses everywhere else, so it is allowed
here only because the check compares each attribute against the text of its
own cell. If that ever stops being true the attributes should be deleted, not
trusted: a reader given 4 from a cell that now reads $5 is worse
off than one who had to read the cell.
The blanks needed saying properly too. Four cells on that table are empty
because Meta does not publish the figure, and one is empty because Llama
models have no vendor-set limit to publish. In JSON all of that would arrive
as null, which reads as “we don't know” — the
exact impression this site spends its time trying not to give. Each one
carries the reason instead, in the file, in words.
Separately and much more simply: not one table on this site had a
caption, and not one header cell said whether it headed a row or
a column. Six tables, thirty-nine rows. Someone reading the dense
eight-column price table with a screen reader got cells with nothing to
attach them to. Every table now names itself, and every figure now belongs to
a named row and a named column. Nothing moved on screen: the two pages that
gained only headers and captions render pixel for pixel as they did before.
All of it is live.
29 August 2026 — a link that read out fifty words, and four repairs behind it#
Every card on this site was one enormous link. The whole box —
heading, description, and the small line underneath — sat inside a
single a element, which is convenient with a mouse and
hostile with a screen reader: the link announced itself by reading its
entire contents, close to fifty words, before the reader could tell
whether they wanted it. The card headings sat inside those links too, so
moving through a page by heading landed on a link every time. The
structure the page was written with was not the structure it offered.
Each link now wraps its title and nothing else. The clickable area is put back over the whole card with a stretched pseudo-element, so nothing changes for a mouse, and the card border lights when the title inside it takes keyboard focus. Nothing about the page moved: the card boxes are the same size to the pixel, which was checked by rendering both versions and comparing them rather than by reading the stylesheet. The one real difference is eight pixels of slack under the track cards that came from a list rule leaking in, and was never meant to be there — the concept cards had already been told to ignore it.
Behind that, four smaller repairs. No page had an article or
a section anywhere in it, so nothing in the markup told a
reader's software which part of the page was the page and which part was
the furniture around it. Every page is now one article, and every run
under a heading is a section named by that heading — free, because
the headings already carry the ids. Nothing moved on screen. A structural
check now requires the article, so the next page written cannot quietly
go without one.
The stylesheet never declared color-scheme. Dark mode
repainted everything this site draws and nothing the browser draws, most
visibly the scrollbar down the side of a long page, which stayed pale
against a dark one. That is one line. Separately, the skip link moved the
scroll position but not the keyboard focus in Safari, so the first Tab
after using it went back into the navigation it had just skipped;
main is now focusable, so the link lands where it says it
does. The response headers gained HSTS, a
Permissions-Policy turning off a list of device features
this site has no use for, and object-src 'none'.
One item on the same list was dropped rather than done, which is worth
saying plainly. The plan was a noindex header on the 404 page,
so that a crawler cannot index an unbounded number of addresses that all
render it. Header rules match the address that was requested, not the file
that gets served, so the rule would have covered a literal request for
/404 and missed every mistyped address — which is the
whole of the problem. The page already carries a noindex tag
in its own markup, and a crawler reads that whichever address it arrived
on. The header would have read as a second guard while guarding nothing,
so there is no second guard. All of it is live.
29 August 2026 — two dead ends, and a nav decision that can no longer quietly rot#
Model facts is the most linked-to page on this site and the one most likely to be arrived at cold from a search, and it offered exactly one way onward: a single link to Context windows, which is probably where the reader had just come from. Someone who lands there for a price, gets it, and wants to know how far to trust it was being shown the door. It now offers the concepts behind the numbers and, separately, the three pages for checking any of it yourself — where the figures came from, when they last moved, and how to tell whether something has gone stale.
Thinking and reasoning was reachable from none of the three tracks on Start here. It is listed as a foundational concept and every other foundational concept is routed to; this one was reachable only sideways, from its sibling pages. A reader following the path as written never arrived. It is now the fourth step on the practical track, where it belongs: knowing when a slower answer is the better one is a using-it question, not a building-it one.
The third item was a question rather than a fault: should Guides go back in the navigation? It was taken out on 14 August for a stated reason — that everything on it was reachable from the home page directly. Two guides were added on 26 August and the home page was not updated, so that reason quietly stopped being true while the decision it justified stayed in place. Nothing noticed for three days, and it was the same paragraph that was also claiming six guides where there were eight.
Which makes the answer straightforward. The reason is true again as of yesterday, so Guides stays out of the nav — and the structural check now reads the list of guides off the Guides page itself and fails if any one of them is not linked from the home page. The decision has not changed. What changed is that it can no longer stop being justified without anyone hearing about it, which was the actual problem. If that check ever fails, the two honest answers are to link the guide from the home page or to put Guides back in the nav. Deleting the check is not one of them. Both fixes and the check are live.
29 August 2026 — the page with the prices on it did not print the prices#
There were no print styles at all. On screen the Model facts table scrolls sideways inside its own box, which is the right answer for a phone. Paper has no sideways. Printed or saved as a PDF on A4, the table ran off the edge and took four of its eight columns with it — including both price columns — with nothing on the page to say anything was missing. The one page on this site that exists to carry prices printed without them, and had done since it was written.
The fix is one load-bearing line, that table cells may wrap on paper, plus tidying: no navigation, no permalink marks, headings repeated across pages, and an external link printing the address it points at, since a reader holding paper cannot follow it otherwise. This was checked by actually printing the page to a PDF and looking at it, rather than by reasoning about the stylesheet. Two earlier attempts to check it another way both gave confident wrong answers.
That test caught a second problem, made yesterday. The eight notes under the table were folded up into disclosures, and a folded disclosure prints as a heading with nothing underneath it — so seven of the eight notes, including both corrections, printed as empty titles. On paper a reader cannot click, which makes folded the same as deleted, and this site does not get to hide its corrections. The obvious fix, telling the contents to display anyway, does not work: a closed disclosure hides its content somewhere a stylesheet cannot ordinarily reach. There is a newer selector that can, and it is now used, so this is fixed in current browsers and not in older ones. The warning about long-context pricing is deliberately not folded in the first place, which is why the one note that costs money never depended on any of this.
Separately, fourteen page descriptions were long enough to be cut off mid-sentence in search results, and each was written in four places at once — the description itself, two social-preview copies and the structured data — so all four moved together. They are shorter now, with the part that distinguishes the page moved to the front. All of it is live.
29 August 2026 — eight words defined, nine sentences rewritten, one paragraph folded up#
A pass for the reader this site claims to be written for. Eight terms were being used on these pages and defined nowhere: attention, calibration, confabulation, corpus, prompt engineering, semantic search, streaming and workflow. They are in the glossary now. The worst of them was attention, which was the heading of a section on Context windows — the page this site calls its highest-leverage one — and appeared nowhere else on the site, defined or otherwise. Two of the others carried whole arguments: hallucination's case for why some invention is unavoidable turns on a model being well calibrated, and the site recommends confabulation as the better word without ever saying what it means.
Nine sentences were rewritten for the same reason: they assumed a reader who already knew something. A heading in Latin on a page for beginners (“whole-corpus questions”), an unexplained HTTP status code, a reference to “the ratio everyone quotes” on the home page written for someone who has never met it, a research finding compressed to “the shape is a U” before the shape had been described. Section ids did not change, so every link into these pages still lands where it did; only the words a reader sees are different. The glossary count on Guides went the same way as the count on the home page yesterday: it said forty-four, it is now fifty-two, and rather than change the number it has been removed. A figure that has to be remembered every time a page grows is a figure that will be wrong.
And the note under the Model facts table, which had grown to 608 words in a single paragraph holding eight unrelated topics, is now eight foldable notes. Two rules about how. Each correction keeps its date and its headline visible while it is folded, because a site whose whole claim is that it publishes its mistakes cannot hide them behind a disclosure widget. And the note that actually costs money — that the prices in the table are short-context rates, and long-context work is roughly double — is not folded at all. It had been sitting at word 590. All of it is live.
29 August 2026 — three corrections, and the rule that should have caught one of them#
The home page said six guides where there were eight. It had been wrong since 26 August, when two guides were added and the sentence describing them was not. The fix is not to change six to eight. A count on the home page is a figure that has to be remembered every time a page is added, which means it will be wrong again; it has been removed, and the sentence now names the pages instead. In the same sentence, two of them were named without being linked — which is precisely the finding logged here on 12 August about two tracks naming pages they never pointed at. It came back because nothing was watching for it.
How to check an AI's answer quoted a document twice and never cited it. It was the only guide on the site with no Sources block. The document was listed on Sources the whole time, so the citation existed, just not anywhere the reader doing the checking would find it. This site's own sentence, from 12 August: a citation you cannot follow to the claim is worse than no citation at all.
Getting better results names a current model. That is against the rule below, and it is staying, because it is a dated quotation from a named vendor document and the whole point of the passage is that model-specific advice inverts on the next model. What was missing was the date inside the sentence. It now says when the guidance was fetched, and says outright that the exception will stop being true of the newest model well before it stops being printed here. A claim that carries its own expiry is a different thing from one that does not.
And the rule itself now has a check behind it. Every hard figure and every model name is supposed to live on Model facts and nowhere else — that is what makes one page rot instead of nine, and it is the reason this site is maintainable by one person at all. Nothing enforced it. The third correction above was found by a person reading, which is not a system. The structural check now reads the model names out of the table itself, so it cannot fall out of step with the thing it checks, and reports any that turn up elsewhere.
Four pages are exempt and each says why in the code: this log, which quotes wrong figures on purpose; the two concept pages granted the arithmetic exemption on Model facts; and Sources, which records where each figure came from. An exemption that stops being needed is reported too, because an allowlist nobody prunes stops being a list of exceptions and becomes permission for anything. What the check does not catch is worth stating plainly rather than being discovered later: word counts, window sizes written out in prose, and dates, none of which can be told from ordinary numbers without raising more false alarms than findings. All of it is live.
29 August 2026 — four diagrams, in the four places the prose was worst#
Until today there was not a single image anywhere on this site — no drawing, no photograph, no diagram on any page. That was mostly a good instinct badly applied. It kept the site fast and kept it honest, and it also meant five mechanisms that are fundamentally spatial were being taught entirely in sentences. The worst case was on Context windows, which described a research finding with the words “the shape is a U” to a reader who has, by construction, never seen the graph.
There are now four drawings. The desk, shown across three turns, so the thing that is genuinely surprising — that the whole conversation is laid out again from scratch every time, and the earliest part goes over the edge when it stops fitting — is visible rather than asserted. The U itself. The line between the model and the software around it, on Tool use, because the single most useful fact on that page is which side of the line each thing happens on. And the loop on Agents, where both of the hard problems live in the return arrow.
They are inline drawings using the same handful of colours defined once for the whole site, so they follow light and dark mode without a second copy of the palette, and they cost no extra request. No JavaScript, no change to the content security policy, and no image files.
Two rules they are held to. Every figure carries a caption saying the same thing in words, which is not a courtesy to anyone who cannot see the picture but a test of the picture: if the caption cannot be written, the drawing was decoration and does not belong. And none of them carry a number. The U has no figures on its axes on purpose — the shape is the finding, the numbers behind it come from a 2023 paper whose models are long retired, and a diagram with a price or a limit drawn on it would simply be a second page with numbers on it.
Two things were wrong in the first version and were caught by looking at them rather than by thinking about them. The message falling off the third desk had been drawn on top of that desk's own label. And on Tool use the arrow carrying the result back started from the wrong box, so the drawing showed the result arriving directly from the outside world — quietly contradicting the one point the diagram exists to make. Both are the kind of error that survives any amount of re-reading the source.
One honest limitation. The drawings scale with the column they sit in, so on a phone their labels end up smaller than the body text even after being sized up for narrow screens. The caption is what carries the content there, which is the main reason the caption rule is a rule. Separately, no last verified date moved for any of this: drawing a picture of an explanation is not the same as checking the explanation is still true, and this site does not move that date for work that did not do that. The drawings are live.
28 August 2026 — dates a machine can read#
The robots file asks every crawler that takes a
fact from this site to carry the last-verified date along with it. That was
an odd thing to ask, because until today the date was only prose —
bolded text in a sentence, in a format a parser has to guess at, with
nothing marking which of the dates on a page was the one being asked about.
Every date on the site now sits in a time element with an
ISO value beside the words a reader sees. The request is the same; it is now
one a machine can actually comply with.
Each page also carries a block of structured data naming what the page is, when it was published, when it was last verified, and the address it is served at. Model facts declares itself a dataset and names the four quantities it measures. Nothing about the pages themselves has changed.
This needed no change to the content security policy, which is worth saying
because the source of every page now contains the word
script and this site has made a point of serving none. The
policy is still script-src 'none', the strictest setting
available, and the position taken on 25 August — that saving a reader
one keystroke is not worth giving that up — is unchanged. A structured
data block is data, not code: the browser never executes it, and the policy
never applies to it. That was tested rather than assumed, by serving the
site under its own real policy and confirming the browser reported no
violation and the block still parsed.
The uncomfortable part is that the last-verified date is now written twice on every page, once for a reader and once for a machine. Two copies of the same fact is the exact arrangement this site refuses everywhere else, and it is the reason there is only one page with numbers on it. It is allowed here on one condition: the structural check now fails if the two dates disagree, or if either goes missing. That was confirmed by changing each of them independently and watching the check refuse both. If that guard is ever removed, the right move is to delete the structured data, not to trust it.
What the structured data deliberately does not do is repeat the pages. The glossary declares itself a set of defined terms without listing its forty-four entries, and the log does not restate its entries as posts, because the Atom feed is already exactly that and keeping two machine-readable copies of the same log is the same mistake in a different file. Structured data here describes a page. It does not become a second copy of one. Committed and live the same day.
28 August 2026 — the structural check now enforces the rules#
The structural check resolved every link, every heading permalink and every id on the site, and enforced not one of the four rules this site says it runs on. A page could lose its last-verified stamp — the element the whole premise rests on — and the check printed clean. So could a page that lost its navigation, its language attribute, or its skip link. Each of those was confirmed by removing it and watching nothing happen.
One of them was worse than a miss. The check that finds pages nobody links
to reads links out of each page's main element, so a page that
lost that element contributed no links at all — which meant losing it
went unnoticed and quietly weakened the orphan check for every page
it pointed at. A check that gets less useful as the site breaks is worse
than a check that fails, because it fails in the direction of reassurance.
Every page must now carry a stamp, one h1, the shared
navigation, a skip link, a language attribute, and a canonical URL matching
the address it is actually served at, with og:url agreeing with
it. The sitemap is compared against the files on disk in both directions, so
a page missing from it is a finding rather than a page that silently stops
being indexed. All twenty-seven pages passed on the first run, which is the
point of writing it this way: it goes green today and only ever speaks up
about a regression. Ten deliberate breakages confirmed it speaks up about
each one. Committed and live the same day.
28 August 2026 — the watcher could report success while blind#
The entry below this one, from 26 August, says of the price watcher that it fails loudly and never quietly, and it names the two ways that was tested: a figure was altered to confirm a change is caught, and a source URL was pointed at nothing to confirm the run fails rather than passes. Both tests still pass. Neither of them touched the shape of the table itself, and the shape of the table is where the script was broken.
The watcher reads Model facts by finding rows and recognising them by how
many cells they hold. A row it could not recognise was discarded without a
record — not reported, not counted, simply gone — and a row that no
probe happened to select produced no result either. Neither case is a change,
and neither is a failure to verify, so a run could end by printing that Model
facts matches every vendor page while having quietly stopped watching part of
it. Promoting one cell from td to th was enough to
do it, which is exactly the edit someone makes to improve a table for screen
readers.
Two other checks had the same shape of fault. feed.py builds the
feed from this log and check.py confirms the feed still matches
the log, which is the arrangement described on 25 August and the only reason a
generated feed is safe to rely on. Both found entries with the same pattern,
and that pattern required the id to be the first attribute in the heading. An
entry written any other way was invisible to both at once: the feed shipped
without it and the check agreed the short feed was correct. Two checks that
fail together are one check.
The third is smaller and worse. This site's fourth rule is that a figure it could not verify is left visibly blank rather than filled in, and the table already carries cells reading “no first-party price”. The price comparison turned its input into a number with nothing to catch a cell that is not one, so the first time a priced figure was honestly marked unverifiable, the run would have stopped there and skipped every check after it. The rule and the tool disagreed, and the tool would have won.
All three are fixed. Rows are matched as markup rather than by counting on a
pattern; a row that does not parse is reported instead of dropped; the number
of rows is asserted, so one going missing is itself a finding; and every row
must now be claimed by a probe that actually compared it, because parsing a
row and checking a row are not the same thing. feed.py refuses to
write a feed shorter than the log, and check.py counts the
headings itself rather than sharing the pattern it is supposed to be checking.
Each fix was tested the way the 26 August ones were, by breaking the page on
purpose and confirming the tools complain.
What this still does not do is worth stating, since the entry it corrects made its promise a little too broadly. The watcher now knows that every row reached a probe. It does not know that the probe compared the right figure, or that the vendor page it read was the page it should have read. And none of these faults were visible to the tools themselves: finding them meant reading each script against the claims it makes, which is something a person has to decide to do and no schedule enforces. No figure on Model facts was wrong as a result of any of this — the first run after the fix checked 52 figures against all seven vendor pages and found nothing had moved. The changes are live.
26 August 2026 — a daily watcher on the price table#
The check logged below found a price that had moved, and the more useful finding was how long it had been wrong: up to twelve days, because the table is checked once a month. Monthly checking means a page about money can be wrong for a month, which is a poor answer from a site whose whole claim is that it is maintained. There is now a script that fetches all seven vendor pages and diffs every figure against what the table says. Run daily, the window becomes a day.
It reports and never edits. That is a deliberate limit rather than a missing feature: deciding what to do about a change is judgement, and this site's own procedure has said since 14 August that a new model on a vendor page is not automatically a new row. A script that quietly rewrote the table would also be a script that could quietly rewrite it wrongly, at three in the morning, with nobody reading.
The rule it is actually built around is the second one: it fails loudly and never quietly. If a figure cannot be found because a vendor restructured a page, or a URL moved, or the fetch failed, it says so and exits with an error. It must never report "unchanged" when what happened is that it could not look. A watcher that goes blind and keeps saying everything is fine is a false freshness signal, which is the exact thing this site tells readers to distrust. Both behaviours were tested by breaking them on purpose: a figure was altered to confirm it is caught, and a source URL was pointed at nothing to confirm the run fails rather than passes.
What it cannot do is worth stating publicly, because otherwise this entry would be claiming more than it should. It compares figures the table already carries. It cannot notice a model that ought to be added, a pricing tier that did not exist last month, a cell that should stop reading "not published", or a caveat still present in the same words that has stopped meaning what it did. All of those have happened here at least once. The monthly read by a person stays exactly as it was, still due 11 September. A clean run from the watcher means nothing moved underneath the table, not that the table is right.
26 August 2026 — the monthly check, run early, and OpenAI has cut a price#
The monthly Model facts check is due on 11 September. It was run today instead, sixteen days early, and it found something, which is the argument for running these off-schedule occasionally: a table that is only checked when the calendar says so is wrong for up to a month at a time by design.
GPT-5.6 Sol has come down from $5 / $30 per million tokens to $4 / $20. The long-context tier moved with it, from $10 / $45 to $8 / $30. Both of OpenAI's own pages agree, the pricing page and the model catalogue, which is worth saying because a single page can be mid-update. The 14 August entry records OpenAI as re-read and unchanged, so the likely account is a cut in the twelve days since. That cannot be proved from outside, and it is worth being exact about why: a vendor publishes today's price, not the date it last moved, so a cut on the 20th and a bad reading on the 14th are indistinguishable from here. The honest claim is the narrow one: the table was wrong today, and it was checked and right on the 14th. Anyone who budgeted a Sol job from here in that time over-estimated by a quarter on input and half on output. Over-estimating is the better direction to be wrong in and it is still being wrong.
Everything else was unchanged. All four Claude rows match the vendor's table cell for cell, including the two cutoff columns and the pinned Haiku API ID. GPT-5.6 Terra is still $2 / $12 with a $4 / $18 long-context tier. Gemini 3.6 Flash is still $0.75 / $3.75, still with the 1 January 2027 increase to $1.50 / $7.50 that this table already states. Meta's Muse Spark 1.2 is still $1.25 / $4.25, still with a contributor tier, and Meta still publishes no first-party Llama price. The Sonnet 5 note that was the subject of the 11 August correction is still on Anthropic's pricing page in the same words, so that correction still stands.
Three things were found and deliberately not acted on, because a verification pass is the wrong moment to redesign a table. OpenAI's third GPT-5.6 model is still listed and still not here, as it was on 14 August. OpenAI now documents a shorter alias alongside the pinned Sol ID, which is the same alias-versus-pinned distinction the Haiku row already carries a note about. And Anthropic's pricing page lists a model marked limited availability that this table does not carry. All three are selection decisions, and the selection is deliberately narrow: current general-purpose text models, not a catalogue.
Verification dates move on Model facts and on the seven rows of Sources that were actually re-read. Nothing else on the site moved. The next Model facts check is still due 11 September: running one early does not buy the site a later deadline, which would be the obvious way to turn a good habit into a loophole.
26 August 2026 — an editing pass on the three newest pages#
Three pages written in two days were measured against the pages written before them, looking for repeated phrasing rather than repeated facts. The site did this once before, in the pass logged on 12 August that found the same rule stated on six pages, and the same method found four things.
Em dashes had crept in. The four guides written earlier carry none at all in their body text. The three new ones had arrived with twenty-one between them. That is not a rule anybody wrote down, which is exactly why it was worth catching: a house voice is mostly habits nobody names, and drifting out of one is how a site stops reading as though one person wrote it. All twenty-one are gone, replaced by the colons, commas and full stops the older pages use for the same job.
One passage was another page rewritten. How to tell if it actually got better closed its section on model-graded evaluation with a paragraph that restated the argument on How to check an AI's answer nearly sentence for sentence. It now points at that page instead and keeps only what belongs to measuring: a grader sharing the blind spots of what it grades does not record them as failures but agrees with them, so the mistake arrives as a passing score and the metric moves the way you were hoping.
A phrase and two habits. "The one almost nobody does" had been used on two pages. One page said "the one that" three times. Two pages repeated a phrase inside a single paragraph. All reworded.
No verification date moved on any page, and no claim changed. This was wording only, which is the case the dates on this site exist to distinguish: re-reading a source is a different act from rereading a sentence, and only one of them says anything about whether the page is still true. What the pass did not touch is the wording shared between a page's own sources block and its row on Sources. That overlap is the index doing its job, and rewording it for the sake of variety would make a citation harder to trace, not easier.
26 August 2026 — the quarterly check now names the guides#
The scheduled quarterly pass above said "the concept pages". The entry logged yesterday for How to tell if it actually got better said that page joins the quarterly check, and there are now eight guides in the same position: no figures on them, and a citation each that can quietly stop being the right one. A reader comparing the two would have found a page claiming cover the schedule did not offer. The block now names both.
This reads like the site promising more work, and it is close to the opposite, which is why it was worth writing down rather than just editing. The quarterly check has never been a walk through a list of pages. It is a re-reading of the documents on Sources, followed by seeing which pages have to move as a result — and every document the guides depend on was already on that list, because a page here does not get published citing something the list does not carry. Naming the guides adds no documents to re-read. It corrects a description of the check that had quietly stopped matching the check.
The 25 August entry has not been edited. It was accurate about the intent and the schedule was wrong; a log that gets tidied after the fact is not a record, and the correction belongs here where it is dated.
26 August 2026 — new page: Designing for the ways it fails#
The Builder track on Start here has four steps. Three linked somewhere. The fourth — assume it will fail and design for it, naming rate limits, refusals, truncated output and malformed responses — linked nowhere, because there was no page to link to. That is the same finding as 12 August, when two tracks were found naming pages they never pointed at; the difference is that this time the page genuinely did not exist. It does now, and the step points at it.
The argument the page is built on: the failures most likely to reach your users do not arrive as errors. A response cut off at the output ceiling and a response the model declined to give both come back with a successful status, recorded in a field naming why generation stopped. Nothing throws, no retry fires, and monitoring shows a healthy service. Reading that field and branching on it is a few lines, and it converts a class of silent corruption into something you can see.
Also on it: why rate limits do not reset the way people assume, since capacity replenishes continuously rather than at the top of a minute, so there is no moment to wait for; why a quoted per-minute limit can be enforced over much shorter intervals, meaning a burst fails while your average sits comfortably under; the rate-limit refusal that never clears and that a naive client will retry forever; that the official SDKs already retry with backoff, so a hand-written loop sits on top of one that exists; and that streaming moves an error to after the successful status, where the standard handling does not reach it.
Three vendor documents read today — the API errors page, the rate limits page and the Messages API reference — all now on Sources. The page carries no limits, tiers or figures, and says plainly that it describes one vendor's API because that is the documentation that was read, rather than implying the shape is universal. It closes by conceding its own limits: everything on it is ordinary engineering, and the failure that is actually new to this technology is the confident wrong answer, which no retry policy touches.
26 August 2026 — new page: Why it does things nobody asked for#
Why it does things nobody asked for began as a request for a page about AI systems escaping their sandboxes during tests, and is deliberately not that page. An incident list is the fastest-rotting thing this site could publish after prices, and unlike prices it only grows; worse, the genre is mostly misreported, so collecting the stories under that heading would have meant a confident register over thin secondary coverage, on the site that names that failure. What the page does instead is explain the mechanism, which does not rot, and then teach the reading skill, which is the part that survives the next headline.
The mechanism is specification gaming: satisfying the letter of an objective while missing its point, because the measurable version of the goal had a loophole and optimisation found it. No intent is required and none should be imagined. The examples are the vendor's own — a boat-racing system that scored higher by circling checkpoints forever instead of finishing the race, and a model making a test harness report success without running anything, which the write-up compares to a student writing "A+" on their own essay.
The finding that deserves the alarm is not the cheating and does not make headlines, because it is about generalisation. In one study a model that learned to cheat on programming tasks began showing unrelated misaligned behaviour it was never trained for, appearing at the same point in training. In an earlier one, models taken through a curriculum of escalating opportunities to cheat went on to tamper with their own reward function, which they were never trained to do, while a model that skipped the curriculum never did. Cheating turned out not to stay in its box.
And the caveats are the page's spine rather than a footnote, because they are what the coverage drops: environments deliberately built to be hackable, models told they were in training, a hidden scratchpad to plan in, rarity even then. The 2024 write-up states outright that its authors make no claims about how likely production models are to behave this way in realistic use. When the people who ran the experiment refuse the extrapolation the headline rests on, that is the most useful signal available, and it is always in the source rather than the article about it.
No measurements are reproduced. They are rates for specific models in specific setups, they would rot like a price, and a percentage quoted away from its method is how a careful result becomes a scary statistic — the same rule this site applies to the four research papers it cites. Both write-ups are by the developer of a commercial model, which the page says is worth knowing in both directions: closest to the evidence, and not disinterested. The glossary gains reward hacking and specification gaming, bringing it to forty-four.
26 August 2026 — Tokens gains the cost shape it was missing#
Tokens explained that you are billed per token and that output costs more than input, and stopped there. It never said why a long conversation gets expensive out of proportion to its length: these systems hold nothing between requests, so every message resends the whole conversation, and the tenth message pays to read the first nine again. Cost grows with roughly the square of the length. Doubling the turns quadruples the input you are billed for.
It matters more for agents, which resend an accumulated trace at every step, so a long run spends most of its input budget re-reading its own history rather than working. Both mitigations were already on the site and unlinked from the place the problem is described: caching the unchanged prefix, and not carrying history you no longer need.
Added as a section rather than a page, on the grounds that a page for it would have been padding. It is arithmetic rather than a per-model figure, so it belongs on the one page whose thesis is arithmetic, and it needed no source: it follows from how the request works. No verification date moved — nothing was re-read, and nothing that was already there changed.
25 August 2026 — the log gets a feed#
There is now an Atom feed of this page, so a correction can reach you without you remembering to come back and look. It is linked in the head of every page here, which is how feed readers find one, and it carries all thirty-two entries in the log with their dates and their permalinks.
This is the only subscribable thing this site will offer, and it is worth saying what it is not. It is not AI news, which is covered by thousands of people with staff and which this site would do badly and then abandon. It is not a newsletter, which is refused on the record along with ads and affiliate links. What arrives is: a page you may have read has been corrected, or re-checked, or was wrong. Nobody else in this subject publishes a corrections feed, largely because publishing one requires issuing corrections.
It is generated rather than written. A small script reads the log and rewrites the feed from it, so an entry cannot exist in one and not the other, and the structural check now fails if the two have drifted apart. That mattered enough to build properly: a feed is a promise that a correction will reach the people who subscribed, and the ordinary way that promise breaks is not malice but a month where the log moved and the feed was forgotten. A rule in the procedure would have been the weaker fix. The check is the stronger one, and it was tested by breaking the feed on purpose and confirming it failed.
One claim on the site had to change to stay true. The README said the site has no build step, and now one file is generated, so it now says what is actually the case: every page still ships as the committed HTML a reader receives, and the single derived file is the feed. Fourth self-referential correction in two days, and the first one caught before it was published rather than after.
25 August 2026 — new page: How to tell if it actually got better#
How to tell if it actually got better closes a gap this site created for itself on 14 August. The glossary gained an entry for evals that day, described as the conspicuous absence because it names the discipline the whole site argues for — checking output against a standard instead of trusting that it looked right. Then no page taught it. The site told readers to change their prompts and offered no way to find out whether the change helped.
It is the missing third of the practical track. Getting better results is about changing what you send. How to check an AI's answer is about judging one output. Neither addresses the question underneath both, which is whether a change helped across the inputs you did not personally re-read, and that question cannot be answered by reading output at all.
The argument the page turns on: one run is not a sample, the case you re-ran is the one that just failed, and you are grading with the same brain that wrote the change sixty seconds ago. A prompt change is also not local — an instruction added to stop one irritating behaviour applies to every output, including all the ones that were already fine, and nothing will tell you when it damages those.
The most useful thing in the source is the one that runs against instinct, so it is quoted rather than paraphrased: prioritize volume over quality. More cases with rough automatic scoring beats fewer cases graded carefully by hand, because careful grading is what caps the number of cases and the number of cases is what decides whether an observed difference is real. The page adds the reason that is about people rather than statistics: a check needing judgement gets run once, when you are enthusiastic.
Cited to Anthropic's "Define success criteria and build evaluations", read 25 August 2026 and now listed on Sources. That document is written for people building software on a model, which the page says out loud rather than borrowing the authority, and it does not reproduce the code, the machine-learning metric names, or the example targets expressed as numbers. Those belong to the context they were written in. What transfers is the discipline, and the page gives the version that fits in a spreadsheet: one prompt, ten real inputs, a column per version, read against each other rather than against your memory.
Carries no figures and no model names, so there is nothing on it to go stale. It joins the quarterly check with the concept pages. The structural check caught the one thing a new page always breaks here — it was reachable only from the navigation — and it is now linked from Guides, from the glossary entry that named it, and from both of the pages it sits between.
25 August 2026 — correction: the homepage was claiming a complete list#
The homepage offered three pages under the tracks and called them "the ones worth having read". That was accurate on 14 August, when three was all there was. Two guides have landed since, so the sentence had quietly become a claim that two pages on this site are not worth reading — which is not what it was written to mean and not something anyone here believes.
It now says the three are where to start, names what follows them, and points at Guides for the full six. The trio itself is unchanged. Growing it to five would undo the fix logged below, which was about a homepage that offered too much and buried the page a reader actually wanted; the answer to a longer list is a hub, not a longer homepage.
Third self-referential error found in one afternoon, alongside the dependency count above, and they share a shape worth naming: a sentence that counts or ranks the site's own pages is true on the day it is written and is broken by the next thing that gets added. Nothing about that looks like rot, because the sentence still reads perfectly well. No verification date moved on the homepage. This was a claim about the site, corrected, not a fact re-checked.
25 August 2026 — correction: the dependency count was one short#
Sources said it listed twenty documents: eleven carrying figures, five explaining a mechanism, four research papers. The tables underneath held twenty-one, and the mechanism section held six. The count was wrong from 14 August, when the four papers were added in the same pass that added a sixth mechanism document, and the sentence was updated for the papers and not for the other one.
Small, and worth logging anyway, because of where it happened. That sentence exists to make the maintenance promise auditable — it tells a reader how much work re-verification actually is so they can hold the site to it. A number that undercounts the list it introduces is the same species of error this site keeps finding in other people's writing: a claim about itself that was true when written and quietly stopped being true one edit later. It now reads twenty-two, which includes the document added today.
Every URL on that page was requested today and all twenty-seven returned a working page, including the six learning resources, so the page's verified date moves. The per-document "last read" dates do not, apart from the new row: requesting a URL confirms a link, which is a weaker claim than reading the document, and the page now says which of the two it is making.
14 August 2026 — new page: How to check an AI's answer#
How to check an AI's answer is the companion to Is what you're reading out of date? and exists because those are not the same job. A published page stands still: it has a date, a byline, and other readers who might have caught the error before you arrived. An answer written for you has none of that, and the error in it was shaped to your question, which is exactly what makes it read like the work of someone who understood the problem.
It covers which claims are worth checking at all, since verifying everything costs more than not using the tool; the failure modes worth knowing by name, including the citation that exists and is real and simply does not contain the claim; why a confident negative is the least reliable sentence in any answer; and where to stop, on the argument that these systems are most useful precisely where checking is cheapest.
One technique here is worth the price of the page: ask the same question again in a fresh session and compare the specifics. A recalled fact tends to come back the same way; an invented one tends to be invented differently. Anthropic's developer guidance calls this Best-of-N verification. The page also states the limit, because the asymmetry matters: disagreement is evidence, agreement is not. A mistake absorbed from training data is held perfectly steadily.
Cited to Anthropic's "Reduce hallucinations" guidance, read 14 August 2026 and now listed on Sources with the pages that depend on it. That document is written for developers building systems rather than for people using an assistant, which the page says rather than quietly borrowing the authority.
Carries no figures, so there is nothing on it to go stale. It also absorbs one page that had been proposed and will now not be written separately: a guide to what happens to what you type. What it deliberately does not do is include that as a section. Data retention is a question about privacy, not about whether an answer is true, and bolting it onto a page about verification would have made both arguments worse. Better left unwritten than written in the wrong place.
14 August 2026 — saying plainly who this site is not for#
Start here now opens by saying that this is a reference rather than a course, that nothing here finishes, and that a reader who wants a linear beginner's course they can get to the end of should go and use one, because several exist and some are cheap.
Sending people away seems like an odd thing to add, but the alternative is worse. A reader who wants a course and finds a reference concludes the reference is badly organised, which is a fair complaint about a thing they should not have been reading. Saying so in one paragraph costs less than being quietly wrong for them.
It appears in exactly one place on purpose. An earlier pass through this site found the same rule restated six times across six pages, and the reader had understood it the first time.
No verification date moved.
14 August 2026 — every section now has its own link#
Hover any heading on a concept page, a reference page or an entry in this log and a appears beside it. Clicking it puts that section's address in the address bar. 122 sections across 16 pages now have one.
The reasoning is about how a page gets used rather than how it reads. Someone who wants to settle an argument about, say, why fine-tuning does not add knowledge should be able to link to the paragraph that says so rather than to a page and an instruction to scroll. Pages get cited at the level they can be addressed at, and until today that level was the whole page.
There is deliberately no copy button. That would need JavaScript, and this
site serves none at all: its content security policy is
script-src 'none', the strictest setting available, which is
only possible because there is no script to allow. Saving one keystroke is
not worth giving that up. The address bar is the copyable thing.
One consequence worth stating publicly, because it is a promise: a section's link will not change once it exists. The identifiers were generated from the headings as they read today and are now fixed. A heading can be reworded later, but its link stays as it is, because by then it may be pointed at from somewhere this site cannot see. A link that quietly stops working is the same failure as a fact that quietly stops being true.
No verification date moved. Nothing here changed a claim.
14 August 2026 — the homepage was sending people to the wrong places#
The homepage opened well and then argued with itself. After the three tracks it spent 326 words re-explaining how the site works, listed all ten concept pages, and only then, in a comma-separated line at the very bottom, mentioned Getting better results — the one page written for the track the same homepage calls "where most people are, and least served." The page a reader most likely wanted was the last thing offered, behind an essay about the site's own principles.
Three pages are now offered directly under the tracks: Getting better results, What AI is actually bad at, and Is what you're reading out of date? The three rules are compressed to 123 words in one paragraph, kept rather than cut because they are why the rest is worth reading, and pointed at About for anyone who wants the longer answer. Guides and Sources now appear on the homepage at all, which they did not before.
Guides has also come out of the navigation, and About has gone in. Guides was added on 12 August to rescue three pages that were reachable from almost nowhere, and it did that job. But all three are now one click from the homepage, so what is left is a hub page occupying a navigation slot on all 23 pages in order to forward people to pages they can already see. About earns that slot better: it carries the byline, the funding position and the one commercial connection this site has, and on a site whose only product is trust, that is not footer material. Guides still exists and nothing linking to it has broken.
Nothing was verified and no verification date moved. This was a change to where the links point, not to a single claim.
14 August 2026 — correction: the Gemini prices here were the future ones#
Model facts listed Gemini 3.6 Flash at $1.50 input / $7.50 output per million tokens. The current price is $0.75 and $3.75. Google's pricing page gives those rates "through December 31, 2026", with the higher pair "starting January 1, 2027". This site had printed the price that takes effect next year as though it applied today, so anyone budgeting a Gemini Flash job from this table would have doubled their estimate. It was wrong from 11 August, when that row was first filled, until today.
The table now carries the current rate and the date it expires, stated in the note rather than left implicit. That is the lesson from the Sonnet correction below, applied in the opposite direction: there, a figure was right and its caveat had gone stale; here, the caveat was missing entirely and took the figure with it. A dated price without its date is half a fact.
This was the first run of the monthly check, fired early as a test. It found the error, and it also found that the check cannot run where it was first put: a cloud sandbox that blocks Google's, OpenAI's and Meta's documentation at the network level, which would have produced a confident-looking pass over one vendor of four. The check now runs somewhere it can actually reach the sources. Every figure on the page was re-read today against all seven primary documents; Anthropic, OpenAI and Meta were unchanged.
One thing found and deliberately not acted on: OpenAI now lists a third model in the GPT-5.6 family, cheaper than the two here. This table is a selection rather than a catalogue, and quietly growing it during a verification pass is how a maintained page turns into an unmaintained one. It is noted here instead.
14 August 2026 — the glossary becomes navigable, and gains four terms#
Glossary had 38 terms, each with an anchor built so it could be linked directly, and nothing on the page that surfaced them. One heading, no index, no way through it but scrolling. The anchors existed to be used and the page hid them. It now opens with an A–Z index, and the terms are grouped under letter headings, so a specific word is two clicks away rather than a scroll and a squint. Letters with no terms are shown greyed rather than dropped, so the row keeps a stable shape.
Four terms were added, all of them things people now say routinely and none of them explained anywhere else on this site. Evals was the conspicuous absence: it names the discipline this entire site argues for, which is checking output against a standard instead of trusting that it looked right. Context engineering, context rot and compaction all belong to the same shift, which is that managing what goes into the window has become a bigger job than writing the prompt. That brings the glossary to 42.
No verification date moved. Adding an entry is not re-checking one, and the definitions here carry no figures to re-check.
14 August 2026 — the two research gaps are closed#
Sources has said since the day it was built that two pages describe findings from the research literature while citing none of it, and that the gap would stay flagged until someone read the papers rather than cited whatever a search turned up. Four papers were read today, and both pages changed as a result. Sources now has a Research papers section, and the count of documents this site depends on goes from sixteen to twenty.
Context windows was making one claim where there are really two. The page said attention thins out over a long context, citing nothing. The paper that established the effect measured position: the same passage is found reliably at the start or end of a context and much less reliably in the middle. But that work is from 2023 and the models it tested are retired, so cited alone it would be precisely the stale-authority move this site warns about. What keeps it standing is separate and newer: a 2025 benchmark that rebuilds the needle-in-a-haystack test so the question and the buried answer share almost no wording, and finds accuracy still falling hard as context grows, on models sold on the size of their windows. The page now separates the two, says which is older, and says what makes the newer one a stronger test.
Hallucination gained a section it should have had from the start. Its explanation of the mechanism was right and stays as written. What it lacked was that part of this is provable: there is a statistical lower bound on invention for a well-calibrated model, tied to facts that appear exactly once in training, independent of architecture or data quality. That is a stronger statement than the page's "it isn't malfunctioning", because it says a certain amount of invention is what working correctly looks like. The second addition is more hopeful and more contestable: a 2025 argument that invention persists partly because benchmarks score a guess above an admission of uncertainty, so the fix is partly a scoring convention rather than a technical barrier. The page presents that as an argument its authors are making, not as a settled result, because that is what it is.
Verification dates move on both pages, and this time on Sources too: four of its documents were read today, along with a re-read of the vendor page both concept pages already depended on, which still says what they say it says. No figures from the papers were copied onto the concept pages. They are measurements of specific models from specific years, they would rot exactly like a price, and the papers are one click away for anyone who wants them in their proper context.
14 August 2026 — corrections: two overstatements, and dates on the promises#
A mechanism was stated as an impossibility. What AI is actually bad at said a model "cannot have been trained on documents describing itself, because it did not exist when the data was collected," and Training cutoff said much the same more briefly. The conclusion is right and the mechanism was not: information about a model does reach it after the fact, through later stages of training and through the standing instructions a product sends with every message, which is why an assistant reliably knows its own name. Both passages now say that, and explain why a relayed answer still isn't self-knowledge. This was the site committing its own documented failure — a confident register on thin coverage — on the page that names it.
A rule was stated more absolutely than it is kept. Three pages claimed every hard figure on the site lives on Model facts. Two concept pages carry arithmetic of their own: Tokens, where comparing two vendor ratios is the entire argument, and Context windows, which quotes a word count to make a million tokens concrete. The rule is really about per-model figures, so it now says that. The alternative was stripping the arithmetic out of the one page whose thesis is arithmetic.
The scheduled checks now carry dates. The list above committed to "monthly" and "quarterly" while naming no date, and told readers to discount the site if a date passed with nothing logged — a test that could not actually be run. Each check now says when it is next due. A fourth was added that has no date on purpose: the check that runs when a page lands, because a new page can make an old page wrong without touching a single fact, and that has now happened here twice.
Also: the site has a share card at last, so a link to it no longer arrives as a blank rectangle. It carries a rubber stamp reading "last verified" with the date field deliberately left blank, because a date printed into a static image is exactly the kind of freshness signal this site tells you not to trust. The date belongs on the page, where it can be moved. No verification dates moved for any of the above: these are corrections and additions, not re-readings.
14 August 2026 — re-check: a source that rotted by standing still#
Thinking and reasoning cited the vendor's extended thinking documentation. That document is still published and still correct. It was the wrong citation anyway, which is the case Sources was built to catch and did: extended thinking has become one mode among several rather than the way thinking works, so a page that describes it accurately had stopped describing the subject.
Both of that page's sources were re-opened and re-read today, and the citation now points at the vendor's current thinking documentation. Every claim the page draws from it survived the move: thinking tokens billed as output, thinking counting toward the ceiling on a single response, the shift from a fixed token budget to the model deciding per request with a coarse effort setting, thinking recurring between tool calls, and a change of thinking configuration invalidating prompt caching. The faithfulness research it cites was re-read too and is unchanged.
Nothing on the page needed rewriting, which is the point worth keeping. The page was written to avoid the settings and model names that were going to move, so when they moved it survived and only its footnote was wrong. The verification date on that page moves to today because the sources were genuinely re-read. The date on Sources itself does not: two of its sixteen documents were re-read, not sixteen, and the per-document dates in its table say which two.
12 August 2026 — new page: Sources#
Sources lists every outside document this site depends on: what each one is cited for, which pages fall over if it changes, and when it was last actually read. Sixteen documents, plus the six learning resources on Start here.
It exists because of the citation failure logged below. A figure was cited to two pages that did not contain it, and nothing in the site's structure could have caught that — each page carried its own sources block, and no page carried the whole list. Now one page does, which is where a citation that leads nowhere becomes visible.
It also makes the maintenance promise checkable. "Re-verify quarterly" is easy to say and impossible to audit from outside. "Re-read these sixteen documents and see which pages move" is a specific afternoon, and you can now hold this site to it. The page also flags something a per-page sources block would never surface: the extended thinking document it cites has not changed a word and is quietly becoming the wrong citation, because the mode it describes is being superseded. A source can rot by standing still.
Four citations elsewhere on the site were pointing at URLs that redirected,
two of them because of a doubled /docs/ in the path. They now
point where they actually resolve. One of those, on
RAG and retrieval, turned out to redirect
somewhere different from what the review that flagged it claimed —
which is its own small argument for opening the link rather than trusting
the report.
12 August 2026 — the tracks now lead somewhere#
Start here sorts you into one of three tracks and then, until today, sent two of them off the site. The Curious track told readers to "learn the three failure modes: confident wrongness, silent staleness, and losing the thread" and linked none of them, while this site has a page on each. The Builder track said "tool use, then retrieval, then agents, in that order" and linked none of those either, though all three were written and sitting one click away.
That is a worse problem than a missing page, because nothing about it looks broken. Every link on the page worked. The pages existed. They simply never met, and a reader following the site's own advice would have gone looking for material that was already here. Both tracks now link to the concept pages they name, at the step where they name them, and the Curious track points at the limits page and the glossary that were written for exactly that reader.
Concepts was a flat list of ten, which quietly implied that prompt caching and hallucination are equally relevant to someone who has never called an API. They are not. The ten are now grouped as Foundational (six) and Building (four), using the labels each page already carried at the top of itself — the distinction was recorded in the pages and shown nowhere. The Building group says plainly that a reader who wants to use AI well rather than build with it can stop at the six.
Also removed: two counts of other people's course catalogues ("12 courses", "103 short courses, 13 full, 10 certificates"). They were accurate on 11 August and they are the fastest-rotting numbers this site had, sitting on a page reviewed quarterly, and knowing the size of a catalogue has never helped anyone choose a course. No verification dates moved for any of this: nothing here was a fact re-checked, only pages finally pointed at each other.
12 August 2026 — correction: a citation that didn't lead to the claim#
What this page used to say.
Tokens stated that using OpenAI's
tiktoken library undercounts Claude tokens by 15–20%, and
attributed the figure to Anthropic's models overview and token counting
documentation.
What's true. The figure is real and it is Anthropic's. It
is not on either of those pages. Both were re-read on 12 August 2026 and
neither mentions tiktoken, third-party tokenisers, or any
undercounting percentage. The number is published in Anthropic's token
counting guidance for developers, which is now what the page cites and links
to.
This is the worst kind of error this site can make, and worth being plain about why. A wrong number is a wrong number. A right number behind a citation that doesn't contain it is a trap: a reader who does the sixty-second check this site teaches will follow the link, fail to find the claim, and correctly conclude the figure was invented. On a site whose entire product is checkable sourcing, an uncheckable citation does more damage than an uncited assertion would have.
12 August 2026 — three corrections from re-reading the vendor pages#
Verification dates moved on Tokens, Model facts and Getting better results, because the sources were actually re-read rather than the pages reworded.
- The headline claim on Tokens was too broad. It said the familiar "one token ≈ 4 characters, ≈ 0.75 words" rule is "now materially wrong for current models." Anthropic's own model table still gives 0.75 words per token for Claude Haiku 4.5 — a current model, listed on Model facts — and for the previous generation of million-token models. The rule changed with a tokeniser, introduced at Claude Opus 4.7, not with time. The page now says that, which is a narrower claim and a more useful one: the ratio is a property of a model version, and one vendor's documentation correctly carries both sets of numbers at once.
- Model facts told you the wrong vendor does threshold pricing. The note said Google prices some models differently above and below a prompt-size threshold. OpenAI does it too, and steeply: GPT-5.6 Sol is $10 input / $45 output above the threshold against the $5 / $30 in the table, and GPT-5.6 Terra is $4 / $18 against $2 / $12. The table quoted only the short-context rate with no asterisk, so anyone budgeting a long-context job from it would have been out by roughly double on input.
- Getting better results dropped a caveat from a quote. It recommended asking a model to verify its own answer against named criteria, which the source does recommend. The same paragraph of that source names an exception — Claude Opus 5 verifies its own work unprompted, and the instruction carried over from older models costs tokens and latency for nothing. Trimming a caveat to make advice land cleaner is precisely what this site accuses other people of.
One smaller fix in the same pass: the Model facts column headed "API ID"
showed claude-haiku-4-5, which Anthropic documents as the
alias. The API ID is claude-haiku-4-5-20251001. Both work; the
column claimed to show one specific thing and showed the other.
Worth recording what did not change. Every other figure on Model facts was re-checked against the vendor's own page and still holds, including the Claude Sonnet 5 rate of $2 / $10 that the 11 August correction below established — it is listed with no introductory marker and no end date, so that correction stands.
12 August 2026 — five stale claims, four of them created by this site's own edits#
A review of every page found statements that had been true when written and were no longer true a day later. All of them describe the site itself, which is the failure mode this site is least entitled to have.
- The 404 page was describing a site that no longer exists. It said five concept pages were live, with more coming, when ten are live and the queue was deliberately abolished. Its section list predated the Guides reorganisation, so it omitted Guides, Getting better results and About while listing three pages that had moved inside Guides. Its meta description still used the wording corrected on the page body below.
- About described the vendor rows on Model facts as empty. They were filled the day before. The disclosure it supported is still needed, since the coverage really is uneven, so it now names the actual asymmetry: two cutoff columns Anthropic's table has and the other one doesn't, and Google represented by a fast-tier model rather than its flagship.
- Tool use said the Agents page was planned, twelve lines above a box linking to it. The editing pass logged below caught the box and missed the sentence.
- The out-of-date guide promised four habits and listed five. The fifth was added when the two-dates change landed; the count wasn't. That page's stamp notes it carries no figures on purpose, which made the one number on it the only thing that could be wrong, and it was.
- The sitemap told crawlers that none of the 12 August work happened, carrying 11 August against every page edited since.
The pattern is worth naming, because it isn't the pattern this site was built to catch. None of these were facts about AI that rotted; they were cross-references that broke the moment something else changed. Nothing about them looks like a claim that needs re-checking, which is exactly why they survived a pass that was looking for stale figures. Verification dates did not move: no sources were re-read, and rewording proves nothing.
12 August 2026 — the corrections this site asked for had nowhere to go#
Four pages invited readers to report anything wrong or stale, and this one called corrections the most important kind of entry in the log. None of them said where to send one. There was no address, no form and no link anywhere on the site — an invitation with no door attached.
That is a worse failure than any single wrong figure, because the entire argument for trusting a self-maintained reference is that someone outside it can push back. Without a channel, this log was a record of the site auditing itself, which is exactly the kind of unfalsifiable freshness claim the method page teaches people to distrust.
Every page footer now carries a Report a problem link, and the corrections notes here and on About name it directly. It goes to the public issue tracker rather than a private inbox, so a report is visible whether or not it gets acted on. The site still has no contact form, no services page and no call to action; being reachable about the facts is a different thing from being sold to.
12 August 2026 — Getting better results#
Getting better results fills the hole this site had been carrying since launch. The homepage claims the practical middle is where most people are and where the internet serves them worst, and until now that group got a reading list.
It is prompting written as applied mechanism: eight habits, each traced to the concept page that explains why it works, on the theory that knowing why beats memorising a list. It also states plainly what doesn't work, and is careful about the difference between "this is unsupported" and "this is false", because most writing in this genre is not.
No prompt library, no copyable templates. Those are what most prompting guides sell and the fastest-rotting thing in the field: written for particular models, quickly out of step with the products, and teaching nothing that transfers.
12 August 2026 — a Guides section, and an accessibility pass#
Three pages had quietly become unreachable except from the homepage and
the footer: the out-of-date guide, the limits page and the glossary. They
now have a home at Guides, which is in the
navigation — Guides was taken back out of the navigation on
14 August, once all three were reachable from the homepage directly. See
the entry above. "Home" came out of the nav to make room, since the
wordmark in the corner already does that job.
The accessibility work matters more than it sounds. This site is plain HTML with one stylesheet and no JavaScript, which makes it nearly free to get right, and there was no excuse for the gaps:
- The date stamps failed contrast requirements. The grey used for them, and for table notes, sat at about 3.8:1 against the page in light mode, below the 4.5:1 minimum for text that size. It is now around 5:1. The stamp is the most load-bearing element on the site, and it was the hardest thing on it to read.
- Keyboard users can now skip the navigation and land straight on the content, and every focused link shows a visible outline.
- The wide tables on Model facts can be scrolled from the keyboard and announce themselves to screen readers, rather than being a mouse-only region.
- Motion is disabled for anyone whose system asks for reduced motion.
A reference site that is awkward to read is failing the people most likely to be looking something up in a hurry.
12 August 2026 — every page now carries two dates#
Pages used to show a single last verified date. On a site where everything was written at once, that date did no work: it was also the day the page was written, so it proved nothing except that the site existed. A single date cannot tell you whether a page was published last week or re-read last week, and those are very different claims.
So every page now shows Published and Last verified separately, and there is a short note on how to read them. Right now they match almost everywhere, which is the honest state of affairs: most of this site has not been re-checked since it was written, and saying so is more useful than a stamp that implies otherwise.
The rule that makes this mean something: editing a page does not move its verification date. The editing pass below reworded headings across a dozen pages and moved none of them, because rewording proves nothing about whether the facts still hold. Verification dates move when sources are actually re-read, and that gets logged here.
12 August 2026 — an editing pass, and an admission#
Every page on this site was written in a single sitting, and it showed. Read three of them in a row and the formula surfaced: five pages closed with a box titled "Where this page stops," four more with one titled "What this page deliberately avoids," and several opened with "The one-sentence version." One metaphor appeared on three separate pages. That is what mass-produced writing feels like from the inside, and it is a fair thing to hold against a site that criticises content farms.
So the repeated headings are gone, replaced with ones that say something specific to the page they sit on, and the recurring phrases have been thinned. No arguments changed and no facts moved; this was about the site reading like it was written on purpose rather than produced.
The pass also caught something more serious than style. Several of those closing boxes still said pages were "planned" that have since been written, and one pointed readers at a follow-on that had already shipped. Those are now correct and linked. Stale internal cross-references are a quiet form of the exact rot this site is about, and they are easy to miss precisely because nothing about them looks like a fact.
11 August 2026 — correction: the Sonnet 5 price warning was wrong#
What this page used to say. Model facts carried an asterisk on Claude Sonnet 5's $2 / $10 rate, describing it as introductory, giving $3 / $15 as the standard rate, and telling readers to check whether the promotion had ended before budgeting off it. It was presented as the site's worked example of why undated pricing pages are dangerous.
What is actually true. Anthropic's pricing documentation now states that the $2 / $10 pricing, "announced at launch as introductory pricing through August 31, 2026, is now the standard price," and that "the previously scheduled increase to $3 / $15 per million input/output tokens on September 1, 2026 will not occur." The warning and the asterisks have been removed, and the correction is stated on the page rather than quietly swept up.
This check was scheduled for the end of August. Running it early cost nothing and caught a page that had been wrong since the vendor's decision changed. Worth being blunt about the lesson: the error was not in the number, which was right the whole time, but in the story attached to it. A stale caveat is as misleading as a stale figure, and it is far easier to miss, because nothing about it looks like a number that needs re-checking.
The rest of the Anthropic table was re-verified at the same time. Context windows, output limits, prices and both cutoff dates for all four models match the vendor documentation as of today, unchanged.
11 August 2026 — two reference pages: the limits, and the vocabulary#
What AI is actually bad at covers the weaknesses that follow from how these systems are built rather than from this year's models being immature: no reliable sense of its own ignorance, no clock or continuity, unreliable self-report, trouble with exactness and character-level work, whole-corpus questions that fail quietly, reliability decaying over long tasks, inconsistency between identical requests, and thin coverage delivered in the same confident register as thorough coverage. It deliberately lists no tasks, because task lists are what make this genre of writing stale within months. It also says what these systems are genuinely good at, since a limits page that only lists failures gives a false picture.
Glossary defines the vocabulary in a sentence or two per term, linking onward where a full explainer exists. Terms only, no product names: products change faster than definitions, and this page is meant to stay useful.
Small correction while adding them: the not-found page claimed to list "everything on the site, in full," which stopped being true several pages ago. It now says "the main sections" and points at the new pages too.
11 August 2026 — the blank vendor rows on Model facts are filled#
Model facts carried empty OpenAI, Google and Meta rows from launch, flagged as unverified rather than guessed. Those are now checked against each vendor's own documentation, with every source linked and dated on the page.
Two things came out of the check that were worth more than the numbers:
- Meta doesn't fit the table, and that is the finding. Llama models are downloaded and run wherever you choose, so Meta publishes no per-token price for them; cost depends entirely on the host. Meta's paid API sells different models, and its pricing page states rates without publishing context or output limits, so those cells now read "not published" rather than borrowing a figure from a third party.
- Cross-vendor price comparison is weaker than it looks. Identical per-token rates don't mean identical cost for the same job, because tokenisers differ and the same text is a different number of tokens for each vendor. The table now says so, rather than inviting the comparison it can't support.
Also recorded: OpenAI publishes a knowledge cutoff for its current family, while the Google model page checked does not, which is why that table has no cutoff column. Google's tiered, batch, flex and priority rates and Meta's contributor tier are noted but not reproduced, with a pointer to read the vendor page before budgeting.
11 August 2026 — Thinking and reasoning, and the roadmap closes#
Thinking and reasoning is live. It explains the mechanism without the mystique (the model writes out working, then answers with that working in front of it), separates the problems where it genuinely helps from the ones where it only produces a longer route to a guess, and covers a cost trap that catches people: changing your thinking settings mid-conversation invalidates prompt caching.
It also carries the finding most articles on this subject leave out. Anthropic tested whether stated reasoning reflects actual reasoning by slipping models hints, and found they often used the hint without mentioning it. Visible reasoning is working shown, not a confession, and it should not be treated as a way to verify an answer.
That completes the ten concept pages planned at launch. The concepts hub no longer lists a queue, for the reason given there: a list of pages that don't exist yet helps nobody, and keeping the existing ones correct is the harder and more useful half of the job.
11 August 2026 — two more concept pages: Fine-tuning and Prompt caching#
Fine-tuning argues the distinction the previous two pages kept deferring: it changes how a model behaves, not what it knows. Facts trained in this way cannot be cited, cannot be corrected without retraining, and carry a standing upgrade cost, since a tuned model is frozen against a base model that will be superseded. Use cases and the "prompt engineering may be all you need" line are quoted from OpenAI's own documentation, along with the observation that knowledge acquisition is absent from its list of benefits.
Prompt caching covers the cost lever that rarely makes it into tutorials, and the accident that quietly ruins it: putting something that changes every request inside the section you meant to reuse, so the system writes a cache entry every time and never reads one. That failure produces no error and no visible defect, only a bill. Mechanics quoted and cited from Anthropic's documentation; the rates, lifetimes and minimum sizes deliberately left out, because they are model-specific and change.
That closes the original concept roadmap except for Thinking and reasoning. Nine live.
11 August 2026 — new concept page: RAG and retrieval#
RAG and retrieval is live, the seventh concept page. Its argument: retrieval is mostly a search problem, which is why most of it fails at the search rather than at the model, and the single most useful habit is to test the retrieval step separately from the answer. Also covers the case for skipping the machinery entirely and simply pasting your documents in, which is now the right first move for small collections. Chunk-context and hybrid-search points are quoted and cited from Anthropic's engineering write-up.
No recommended chunk size, no vector database comparison, no embedding rankings, no accuracy figures. Same reasoning as every other concept page: those rot fastest and read as authoritative while doing it.
11 August 2026 — new concept page: Agents#
Agents is live, the sixth concept page and the first written after launch. It takes the position that an agent is a model in a loop with tools, that the loop is where every hard problem lives, and that the workflow/agent distinction is the useful one to hold onto. Definitions are quoted from Anthropic's own engineering guidance and cited on the page.
Deliberately contains no benchmark scores, no framework comparison and no reliability percentages, for the reason given on the page: those are exactly the figures that would make it wrong within a season while still sounding authoritative.
11 August 2026 — three pages added, including this one#
Added the same evening as launch, and all three are about making the site's claims checkable rather than adding more surface to maintain:
- This log. The reasoning is in the box at the top of the page: a date stamp says when something was checked, and says nothing about whether anything happens when a fact rots.
- Is what you're reading out of date? The verification method used here, written up so it works on anything else you read. Contains no figures at all, on purpose.
- About. Who writes this, who maintains it, how it's funded, and the one commercial connection, disclosed rather than tucked away. For a site whose product is trust, having no named human anywhere was a real gap.
11 August 2026 — the site went live#
Published at plainlyai.org with nine pages: a home page, a placement guide, five concept explainers and the model facts table. Everything dated on the day it was written.
Three corrections came out of verifying the Start here resource lists, and they're recorded here rather than quietly patched, because a site about trusting sources doesn't get to hide its own edits:
- A course was wrongly tagged "free." One recommended catalogue publishes no pricing on its index and mixes free short courses with paid certificates. Retagged "mixed, verify per course."
- The beginner track pointed at developer documentation. That documentation says, on its own front page, that it's written for developers. Beginners now get a beginner-facing course, and the API docs moved to the builder track where they belong.
- "Free" was softened to "price not stated" on two entries. Both are openly accessible, but neither publishes pricing on the page checked, and openly accessible today is not the same as documented as free.
One finding from writing the Tokens page is worth repeating here, because it's the clearest example of the rot this site exists to counter: the near-universal rule of thumb for converting words to tokens is now materially wrong, and dividing out a vendor's own published figures shows it. It has been repeated unchanged across the web for years, including by models that learned it from the web. The arithmetic is on the page.
Same day — the plumbing#
Not content, but it affects what you get served: clean URLs throughout, a real not-found page instead of a blank one, and a strict content security policy. The site ships no JavaScript at all, so the tightest possible policy costs nothing. Nothing here tracks you, because there is nothing here to track you with.
What counts as a change#
Three kinds of entry appear in this log, and they're not equally important.
Corrections are the ones that matter. Something here was wrong, someone noticed, it's fixed, and the log says what it used to say. A site that only ever adds and never corrects is either perfect or not checking.
Re-checks are the invisible work: a figure confirmed still accurate, a link opened and found alive. Nothing on the page changes except its date. This is most of the actual maintenance, and it's the part that content farms cannot fake, because it costs exactly as much as being wrong costs nothing.
Additions are new pages. Least interesting from a trust standpoint. Anyone can add.
Found something stale?
A wrong or out-of-date fact here isn't a nitpick, it's a bug, and it's the one kind of bug this site can't tolerate. Corrections to figures are especially welcome, most of all if you can point at a primary source. The correction gets made and it gets logged on this page with what it replaced.
Report one in the issue tracker. It is public on purpose. A correction log kept by the same person who writes the pages is only worth as much as the record of what was reported, and that record shouldn't live where it can be quietly edited.
Learn the method → Is what you're reading out of date? · The numbers → Model facts