7 min read Rocky Elsalaymeh
Our docs described a product that did not exist
Everyone audits the chatbot that invents a fact. Nobody audits the manual. Team-X's docs described a hosted platform that was never built, and strangers could read every line of it. Here is what was wrong and what the repair missed.
Everyone is worried about the chatbot that invents a fact. The quieter failure ships in the manual: a document that describes a product more complete than the one you built, published where strangers act on it.
Team-X, an open-source, local-first desktop app for running AI-agent organizations, shipped v3.2.0 on 2026-05-11 with a docs honesty pass. Its changelog says the docs had described a hosted API, a teamx CLI, freemium tiers, account signup, and a plugin marketplace. None of those existed. The pass stripped them.
This is the account of what was wrong, how a page gets that way, what the repair looked like, and where the repair fell short. Every Team-X claim below is checked against the git tags, not against memory.
What did the docs promise that the product never shipped?
A hosted platform. Team-X is a desktop app with no server of ours in the loop.
The v3.1.0 developer reference opened its API section with one sentence: “Team-X exposes a REST API for workspace automation.” Then it showed how to call it.
# API Key in header
curl -H "Authorization: Bearer YOUR_API_KEY" \
https://api.teamflow-x.com/v1/workspace
That excerpt is copied verbatim from the v3.1.0 file. The page was not a typo or a stale paragraph. Its table of contents listed Workspace API, Webhooks, and Plugin Development as sections five, six, and seven.
Here is the gap, surface by surface, as the changelog and the v3.2.1 docs record it.
| Surface | What the old docs described | What the v3.2.1 repo says |
|---|---|---|
| API | A hosted REST API with Bearer keys | A local-first desktop app; no hosted REST API |
| Extension | Webhooks, OAuth, and a public plugin marketplace | MCP servers and role packs are the two extension points that ship |
| Connectors | GitHub, GitLab, Slack, Discord, Jira, and Notion integrations | No first-party connectors; build them as MCP servers |
| CLI | A teamx binary with ticket, employee, budget, run, and workspace subcommands | No published binary; a developer CLI named team-x-ai with five commands |
| Pricing | Credit packages and Free, Basic, Pro, and Enterprise tiers | Team-X is free; the cost is your LLM provider’s bill |
| Accounts | Sign up and log in with email or social login | No account; the app opens to the Workspace Setup Wizard |
The CLI was the largest single fiction. The changelog counts 530 lines documenting a binary that does not exist. The cleanup took cli-reference.md from 717 lines to 236.
Who drafted it, and does that matter?
The repo does not say. I am not going to guess, because a guess in a post about unprovable claims would be the same mistake.
What the repo does say is how it describes the result. The changelog says the new denial section exists to preempt any drift back toward “the v3.1.x hallucinated content.” The repaired CLI page says the old subcommands “were aspirational and never built”. The changelog says the docs sync would have published roadmap content “as shipped product.”
Whoever or whatever drafted those pages, the failure has a known shape. A developer guide is expected to have an API section, so one appears. A product page is expected to have pricing, so a pricing table appears. The prose is fluent, the structure is familiar, and nothing in it is checked against a compiler.
Models are one source of plausible specifics. In a study of code-generating models, researchers reported that “the average percentage of hallucinated packages is at least 5.2% for commercial models and 21.7% for open-source models.” That finding covers package names in generated code, not documentation. The mechanism it shows is the relevant part: a plausible name that nobody verified.
My working theory, labeled as a theory: the fiction survived because every page read like a real one. A reviewer skimming a table of contents sees a complete product. Only someone who holds the code in their head sees the hole.
Why does a wrong doc cost more than a wrong chat answer?
Because a reader acts on documentation with no conversation to push back in. A chat answer gets questioned. A manual gets obeyed.
Two public cases show the price. In 2024 the British Columbia Civil Resolution Tribunal held Air Canada liable after its chatbot gave a passenger bad advice. BBC reporting quotes the tribunal member: “It should be obvious to Air Canada that it is responsible for all the information on its website.”
In April 2025 a support bot at Cursor told a user that being logged out when switching machines was expected behavior under a new policy. Ars Technica reported it plainly: “But no such policy existed, and Sam was a bot.” The Register quotes the company’s own correction: “Unfortunately, this is an incorrect response from a front-line AI support bot.”
I am not offering a legal opinion about what a docs page binds you to. I am pointing at the direction of travel. Readers and tribunals both treat what you publish as what you said.
The stakes in our case were concrete. Per the changelog, a docs-sync script was mirroring the placeholder brand to the live website, and the Strategia-X site-docs sync would have published roadmap content as shipped product. The host in that curl example belonged to a brand the changelog calls a placeholder that “crept into the docs.”
Developers already discount AI output. The 2025 Stack Overflow survey found that “More developers actively distrust the accuracy of AI tools (46%) than trust it (33%)”. Wrong docs do not have to be written by a model to spend that trust. They only have to be wrong.
A model drafts what is plausible. Only the code knows what shipped.
How do you repair docs that lie?
Rewrite from the code, then write down what is absent. That was the repair, and it had three parts.
First, rebuild each page from a source that can be checked. The restored integration guide was, per the changelog, sourced directly from API_ENDPOINTS.md. The changelog also states the direction of authority outright: “the upstream is now the source of truth.” When the website and the app repo disagree, the app repo wins.
Second, add a denial list. The integration guide now ends with a section titled What is intentionally not here. It opens with “No native GitHub, GitLab, Slack, Discord, Jira, or Notion integration.” A reader who wants to know whether a feature exists gets a direct no, which is better than silence that a model-written neighbor page might fill.
Third, grep for the leftovers. The changelog records the check: “Verification: zero matches remain for freemium, credit-based,” plus a list of fictional command names, outside the intentional denial lists.
Even the repair needed repair. An earlier pass softened an answer about sharing a workspace so it would not invent button names. The changelog says that was wrong too: “the earlier softening was wrong.” The label Collaborators had been invented in earlier drafts. The final answer was written by reading the operator-access code and the portability settings, not by hedging.
That is the lesson inside the lesson. Caution in prose is not verification. The fix that held was reading the implementation.
Did the honesty pass catch everything?
No. I read the v3.2.1 tag to check.
At v3.2.1, line 41 of the FAQ says “there is no paid tier, no employee quota, no workspace quota.” Roughly 130 lines later, under the heading “How many employees can I hire?”, the same file carries a table headed “Employee Quota.” Free gets 3 employees, Basic 10, Pro 25, and Enterprise 50+.
Two lines in one file contradict each other. The verification grep did not catch it because the grep list was made of terms, and this table used none of them.
The repo’s one mechanical claim gate points elsewhere. The script check-claim-evidence.mjs “Parses CLAUDE.md structured claims (IPC channel table + bus events table)” and checks each against the code. Its pre-commit hook has a line that reads “Skip if CLAUDE.md is not in the staged diff”. That is a good gate for the file it reads. It does not read docs/user-guide/.
So the state at v3.2.1 is plain. The worst fictions were removed, one residual table remained in the FAQ, and the user-guide pages had no automated claim check of their own. I would rather report that than round it up.
What should you do before your next release?
Treat every drafted page, by a person or a model, as a list of claims. Then do these five things.
- Grep for capability nouns. Search your docs for API, webhook, SDK, CLI, plan, tier, account, and marketplace. Every hit needs an implementing file or a deletion.
- Write the denial list. Add a “what is intentionally not here” section to your integration docs. State absences as flatly as features.
- Anchor each feature sentence. For every sentence that names a capability, record the
path:linethat implements it. No anchor, no sentence. - Gate docs like code. Write the Docs describes “a philosophy that you should be writing documentation with the same tools as code”, and its list includes code reviews and automated tests. Add a check whose input is your docs folder, not only your config file.
- Find what mirrors you. Sync scripts copy fiction at machine speed. List every site, wiki, and index that republishes your docs, and decide which side wins a disagreement.
The corrected integration guide, including its list of what Team-X will not do, lives at /docs/integration-guide/. The weekly note on how this project is run is at Field Notes.
Frequently asked questions
What was fictional in the Team-X docs before v3.2.0?
Team-X's changelog says its docs described a hosted REST API, a teamx command-line tool, freemium pricing with credit packages, a hosted account signup flow, and a plugin marketplace. Team-X is a local-first desktop app and ships none of them. The v3.2.0 docs honesty pass stripped them out.
Can AI-drafted documentation describe features that do not exist?
Yes, and the Team-X changelog calls its own case hallucinated content. The Team-X repo does not record who or what drafted those pages, so the cause is unproven. What worked was rewriting each page from the code and adding a plain list of what Team-X intentionally does not do.
How do you stop documentation from describing features that were never built?
Team-X's v3.2.0 pass rewrote pages from the code, added a what-is-intentionally-not-here section, and grepped for leftover terms. That grep still missed an employee-quota tier table in the FAQ, so the lesson is to gate docs against the code with a real check, not keywords alone.
Docs are a claims surface. Strangers act on every line.