UKWELI. Integrify your document
Integrity Watch · Folio 05 Figures frozen

About half of new English-language articles online are now mainly written by AI, by one count. The company behind ChatGPT says its models still make things up.

A growth agency counted the articles. OpenAI, which builds ChatGPT, has described the flaw in its own words. South African courts have been handed citations to cases that could not be found. The question that follows is not only who wrote a document. It is whether anyone checked it.

We keep Integrity Watch, a public register of documents that were corrected, withdrawn, refunded or punished after someone found made-up material in them.

This folio is about a trend, and what follows from it.

Half #

Graphite drew a random sample of 55,400 English-language articles from Common Crawl, a large public archive of the web. Each article was dated between January 2020 and March 2026. Graphite ran each one through three commercial AI detectors, Pangram, GPTZero and Copyleaks, and averaged what they found.

Its headline: "The number of articles published on the internet that are primarily AI-generated (50%) is equal to the number written by humans (50%)."

The shape of the curve matters more than the headline number. Within 12 months of ChatGPT's launch, the share rose to about 36%. Graphite puts it at 48% by 24 months. Since the first quarter of 2025 it has stayed between about 45% and 51%: 49.6% in that quarter, a dip to 44.6% in the next, 50.9% at the end of 2025, and 49.9% in the first quarter of 2026.

Share of new English-language articles that are mainly AI-written

Graphite · by quarter of publication · average of three AI detectors · January 2020 to March 2026

This measures who wrote the articles, not whether they are accurate. On articles from before ChatGPT, Graphite measured the detectors wrongly flagging about 1.4% to 1.8% as AI, so small values here may not mean AI use. The line starts rising during 2022, before ChatGPT's launch in November. Redrawn by Ukweli from Graphite's published source data; the figures are Graphite's.
Show the numbers
QuarterMainly AI-written
2020-Q10.97%
2020-Q20.90%
2020-Q30.95%
2020-Q40.64%
2021-Q11.04%
2021-Q21.23%
2021-Q30.89%
2021-Q41.18%
2022-Q11.54%
2022-Q22.37%
2022-Q33.85%
2022-Q44.60%
2023-Q114.02%
2023-Q225.14%
2023-Q333.69%
2023-Q435.92%
2024-Q138.23%
2024-Q237.22%
2024-Q341.81%
2024-Q447.04%
2025-Q149.61%
2025-Q244.58%
2025-Q349.22%
2025-Q450.90%
2026-Q149.94%

Three things about this count are worth holding on to.

It measures authorship, not accuracy. Nothing in the study asks whether any article is right or wrong. A machine-written article can be accurate. A human-written one can be false.

It is a count made by detectors. An article is "primarily AI-generated" when the detectors say so. Graphite tested them before trusting them. On 15,700 articles published before ChatGPT existed, which it treats as human-written, the detectors wrongly flagged between about 1.4% and 1.8% as AI.

It leaves out the hardest case. Graphite says it did not test the detectors on drafts that AI wrote and a person then edited: "We did not evaluate the accuracy of AI detectors using this strategy." So the count cannot tell us how much of the other half was started by a machine and finished by a person.

Graphite also reports, from a separate 2025 study, that AI-written articles appear far less often in Google results and in ChatGPT and Perplexity citations than they do on the web. That study used a different detector. Its figures and this one are not measured the same way, and we do not join them.

What the count does show is the trend. By the detectors' measure, a large share of new articles is now mainly machine-written. What those machines are known to do is the next question.

The builders say so #

On 5 September 2025 OpenAI, the company behind ChatGPT, published research on why its models make things up. It defines the problem plainly: "Hallucinations are plausible but false statements generated by language models."

And it does not exempt its own product: "ChatGPT also hallucinates. GPT‑5 has significantly fewer hallucinations especially when reasoning, but they still occur. Hallucinations remain a fundamental challenge for all large language models, but we are working hard to further reduce them."

Its explanation is about incentives. The researchers argue that models "hallucinate because standard training and evaluation procedures reward guessing over acknowledging uncertainty." A model graded only on right answers learns that a confident guess scores better than "I don't know."

They give an example from their own team. Asked for the title of one author's doctoral dissertation, a widely used chatbot "confidently produced three different answers—none of them correct."

To be fair to the paper, it also rejects the idea that hallucinations cannot be avoided. Its answer to the claim that "Hallucinations are inevitable" is: "They are not, because language models can abstain when uncertain." The fix it proposes is to reward models for saying when they do not know.

That is a reason for hope about future models. It is not a reason to trust the text already being published. The people who build these tools describe false, confident statements as a challenge they have not yet solved.

Read next Folio 04: Two speeches told South Africa's lawyers to check everything. Nobody checked the speeches.

The Deputy Chief Justice and the Legal Practice Council chairperson warned the profession about unchecked AI. The record of their own speeches carried errors of its own.

Measured, not assumed #

A builder's admission tells you the flaw exists. Two large studies show how often invented or inaccurate material turns up: in AI answers about the news, and in the published scientific record.

In the news. Working with the European Broadcasting Union, and building on earlier BBC research, journalists from 22 public service media organisations in 18 countries, working in 14 languages, assessed how ChatGPT, Copilot, Gemini and Perplexity answered questions about the news. They evaluated more than 3,000 answers. The EBU calls it "one of the largest cross-market evaluations of its kind", and summarised the results in October 2025:

How AI assistants handled questions about the news

EBU study with 22 public service media organisations · 18 countries · 14 languages · more than 3,000 answers from ChatGPT, Copilot, Gemini and Perplexity · October 2025

Almost half

of all AI answers had at least one significant issue

A third

of responses showed serious sourcing problems

A fifth

contained major accuracy issues, such as hallucinated and/or outdated information

The EBU's own words, quoted as published. We show them as words, not bars, because the EBU page states them that way and we have not read the exact percentages at source.

Its conclusion was that AI "routinely misrepresents news content, no matter which language, territory, or AI platform is tested."

In science. In May 2026 The Lancet published a letter reporting an audit by researchers led by Maxim Topaz of Columbia University. They checked 97.1 million references in 2.5 million biomedical papers against PubMed, Crossref, OpenAlex and Google Scholar. They found 4,046 references they classed as fabricated, because the publication could not be found in any of those sources. They were spread across 2,810 papers.

The rate climbed steeply. About 1 in 2,828 papers carried at least one fabricated reference in 2023. By 2025 it was 1 in 458. In the first seven weeks of 2026 it was 1 in 277.

Biomedical papers citing at least one fabricated reference

Topaz et al., The Lancet, 2026 · 2.5 million papers in PubMed Central, January 2023 to February 2026 · papers per 10,000

A reference counts as fabricated when the authors could not find it in PubMed, Crossref, OpenAlex or Google Scholar. Bar heights are the letter's "1 in" figures converted to papers per 10,000. * Early 2026 means the first seven weeks of the year; that bar is hatched. The authors say their method finds fabricated references, not their cause. Drawn by Ukweli from the figures published in the letter.
Show the numbers
PeriodPapers affectedPer 10,000
20231 in 2,8283.5
2024no rate given
20251 in 45821.8
Early 2026*1 in 27736.1
Affected papers with no publisher action, at the time of the audit98.4% of 2,810

Two findings in that audit matter most for anyone who publishes. First, the fake references were hard to spot: "topically specific, correctly formatted, attributed to real researchers", with plausible dates. Second, almost nobody had acted on them: "Of the 2810 affected papers, 98·4% had received no publisher action at the time of our audit."

The authors are careful about cause. They write that "Our system identifies the problem, not its cause." Fake references can come from paper mills, from deliberate misconduct, or from "uncritical use of artificial intelligence (AI) writing tools." The sharp rise from mid-2024 matches the time it takes AI-assisted papers to reach print. But they note that paper mills and changes in indexing "might also have contributed."

What the audit shows is that invented sources are entering the published record faster, and that at the time of the audit almost none had been acted on. Why each one got there is a separate question.

Ukweli reads a document before it goes out and tells you which of its claims, quotes and sources cannot be confirmed. Put a document through it →

In South African courts #

Our own register records what happens when invented material is found. Most entries come from courts, and South Africa has its own.

See the South African entries in the register

Around the world. Damien Charlotin, a legal researcher, keeps a public database of these court decisions, and says it has been cited in several court decisions. It "tracks legal decisions in cases where generative AI produced hallucinated content – typically fake citations, but also other types of AI-generated arguments." When we read it on 18 September 2026, it listed 2,041 cases and gave its last update as 14 September.

Counted from his published data by decision date, 16 date from 2023 and 61 from 2024. In 2025 there were 852, and in 2026 so far, to 15 September, 1,112. Comparing like with like, 1 January to 14 September, 2026 had 1,111 against 384 in 2025.

By quarter, the rise was steepest through 2025. Since the end of 2025 it has held at roughly 400 to 450 decisions a quarter.

Court decisions on AI-hallucinated content, as tracked by one database

Charlotin AI Hallucination Cases Database · by quarter of decision, April 2023 to 15 September 2026 · 2,041 dated cases, read 18 September 2026

* The last bar is a part quarter, to 15 September 2026, and is hatched. Part of the rise reflects more people looking and more courts writing the problem down. Recent quarters may still rise as decisions are added. The database counts decisions it has found, so these numbers are a floor. Drawn by Ukweli from the database's published data (CC BY 4.0).
Show the numbers
QuarterDecisions
2023-Q26
2023-Q33
2023-Q47
2024-Q110
2024-Q27
2024-Q319
2024-Q425
2025-Q156
2025-Q2123
2025-Q3268
2025-Q4405
2026-Q1449
2026-Q2411
2026-Q3252 (to 15 Sep)

More than half involve people who went to court without a lawyer. Most of the rest involve lawyers. In 33, the database lists a judge or arbitrator among those who used AI. Six are South African.

Who used the AI, in the court decisions tracked

Charlotin AI Hallucination Cases Database · 2,041 dated decisions to 15 September 2026 · read 18 September 2026

Where a decision lists more than one party, it is counted once: first as a judge or arbitrator if one is listed, then as lawyers, then as self-represented. In 20 decisions the database marks the AI use as alleged. Drawn by Ukweli from the database's published data (CC BY 4.0).
Show the numbers
Party using AIDecisionsShare
Self-represented litigants1,17257.4%
Lawyers and legal staff81539.9%
Judges or arbitrators331.6%
Experts or not stated211.0%

Two cautions travel with those figures. Some of the rise reflects more people looking, and more courts writing the problem down. And the database is clear about its own limit: it "does not track the (necessarily wider) universe of all fake citations or use of AI in court filings." Its numbers are a floor, not a ceiling.

The lawyers. In Mavundla v MEC: Department of Co-Operative Government and Traditional Affairs KwaZulu-Natal, the Pietermaritzburg High Court was given a list of authorities in support of an application for leave to appeal. Judge E Bezuidenhout asked the court's two law researchers to find them. Her judgment of 8 January 2025 records the result: "Of the nine cases referred to and cited, only two could be found to exist, albeit that the citation of one was incorrect."

The judge asked the candidate legal practitioner who drafted the document whether she had used an AI application such as ChatGPT. She denied having done so. The court did not find where the cases came from. Separately, it ran a brief experiment of its own. The citation of one missing case was entered into ChatGPT. "The system responded that the case did indeed exist and revolved around the powers of the Public Protector." Asked whether the case dealt with the point it had been cited for, ChatGPT said it did. The judgment records that this "immediately illustrated the unreliability of it as a source of information and legal research."

The court ordered the law firm to pay the costs of the extra hearings and sent the judgment to the Legal Practice Council.

The bench. In a judgment of 31 July 2026, three judges of the Johannesburg High Court dismissed an appeal against an acting judge's ruling in a family dispute. The ruling's outcome stood. But in a separate judgment, Judge Ingrid Opperman set out 11 discrepancies in the citations in that ruling. They had been confirmed by the senior librarian of the Johannesburg Society of Advocates.

One case the ruling relied on, Lubbe v Volkswagon SA, appears in Judge Opperman's table with a single line: "This case does not exist."

Judge Opperman wrote that "The most plausible explanation, certainly for the fictitious Lubbe reference, is that it is the product of the use of Artificial Intelligence (AI) and what has been dubbed ‘hallucinations’." She was equally clear about the limit of that view: "I make no finding on whether AI was used." She said she would forward her judgment to the chairperson of the Legal Practice Council for an investigation. What it records, she wrote, are not findings but "observations to be investigated."

Her ruling puts the stakes plainly: "Names of non-existent cases are not law." And it says why checking matters: "An essential quality of law is its verifiability; others must be able to find it and check it."

Our own register shows the same spread from a different angle. Its earliest entries come mostly from the United States and Australia, with one from Germany. Its recent ones include Kenya, India, Brazil, Colombia, the Philippines and South Africa. They reach beyond courts, into newsrooms, government policy, academic journals and consulting reports. And they reach the people who decide cases: judges' chambers in the United States, a trial court in India and an arbitrator in Quebec, in each case with AI use found or admitted. The register confirms cases one by one. It shows how far the problem has spread. It does not measure how often it happens.

Neither South African court found that AI was used. Both found that a document placed before a court cited cases that could not be found. That is the failure that matters to the person on the other side, whatever caused it.

Two different questions #

When people first meet the Graphite figure, a natural response is to reach for a detector. A detector and a check answer two different questions, and each has its uses.

A detector asks who wrote it. It estimates whether a machine or a person produced a text. That is useful for measuring a trend, and it is how Graphite's figure was made.

A check asks whether it holds up. It tests whether each claim, quote and source can be confirmed. A fabricated case citation is just as fabricated if a person typed it. An accurate paragraph is just as accurate if a machine drafted it.

Neither answer stands in for the other. Graphite measured false-positive rates below 2% on the articles it tested, on web articles from before 2022 and on AI text from a single plain instruction. A 2023 study in Patterns found that seven older detectors "incorrectly labeled more than half of the TOEFL essays as 'AI-generated'". Those essays were written for TOEFL, an English test taken by people whose first language is not English. Nobody should be accused on a detector's score alone.

People make these errors too. Folio 04 found a misquoted judgment, a wrong date and a misread figure in two speeches to South African lawyers. None of it was fabricated, and none of it was blamed on AI. It was simply not checked.

The Lancet audit points the same way from the other direction. Its authors used an AI model as one of several filters in their checking system. They then had three independent reviewers check a sample of 500 of its results. AI was part of the check, and people were accountable for it.

So for anyone about to publish, the question that protects the reader is the second one: were the claims checked before the document went out?

A claim travels #

The Graphite research is itself a small example of how a claim changes as it moves.

In October 2025 Graphite reported that AI-written articles had "surpassed" human-written ones in November 2024. The reporting that followed diverged. Axios headlined it "Exclusive: AI writing hasn't overwhelmed the web yet". TechRadar ran "The internet is now mostly written by machines, study finds". Futurism ran "Over 50 Percent of the Internet Is Now AI Slop, New Data Finds". Graphite had measured articles. In two of the headlines, the finding became a claim about the internet as a whole.

In May 2026 Graphite revised its work with three detectors instead of one. The new figures were, on average, 3.3 percentage points lower. It marked the October study as "superseded" and explained why.

But the correction did not reach everything. Graphite's separate study on search results, linked from the new one, still says that "the number of AI-generated articles published online in November 2024 exceeded the number of human-written articles." Graphite's revised study puts the share at 48% two years after ChatGPT's launch, and its quarterly figures do not pass half until the end of 2025.

We also found that the superseded study now gives its sample as 43,000 articles. The same page as archived on 18 October 2025 gave 65,000. The notice on the page explains the lower percentages. It does not mention the change in sample size.

None of this is fabrication. It is ordinary drift, in careful work, by a firm that corrects itself in public. It shows how easily a claim outlives its correction, whoever wrote it.

What this means for South African publishers #

The Press Council of South Africa issued a guidance note on AI in November 2023. It is not binding, and it does not extend the Press Code. But its first two points go straight to this problem. On accountability: "Member publications retain editorial responsibility for everything that is published, no matter which tools are used in production. To ensure compliance, any AI-generated material must be checked by human eyes and hands." On accuracy: "Generative AI is known to be prone to the invention of facts (known as ‘hallucinations’)."

That standard applies well beyond newsrooms. Communications teams, government departments, law firms and researchers all publish documents that others rely on. Four things follow.

Check the claims, not the author. Every quotation, figure and source in a document either can be confirmed or cannot. That is true whoever drafted it. And check against the source itself, not against another chatbot. In Mavundla, ChatGPT confirmed a case that does not exist.

Treat the finished-looking source as the risky one. The Lancet audit found that fake references were correctly formatted and credited to real researchers. The Johannesburg judgment cited a case with a full law-report reference. Invented material does not look invented.

Say which tools you used. The Press Council note asks publications to indicate clearly when tools were used to produce an item. A reader can weigh a disclosed method.

When you correct, correct everywhere. A correction that reaches one page and not the next leaves the old claim in circulation.

The trend in the Graphite data has not reversed. About half of new articles is a lot of text, and the company behind the best-known tool says its models can still state false things with confidence. The answer is not to stop using those tools. It is to check what they produce before it carries your name.

How we checked #

We read every source named here at its origin: the Graphite studies, OpenAI's research page, the EBU's report page, the Lancet correspondence, the Charlotin database and its published data, and both South African judgments. Each quotation was copied from the source as published. The archived Graphite page was read through the Internet Archive's Wayback Machine.

We did not run this piece, or any of these documents, through Ukweli.

Our standard. Every quotation and figure here was confirmed by reading its original source, not a summary of it. Where we could not read a source ourselves, we say so, or we leave the claim out. This folio was researched and drafted with the help of Claude, an AI model made by Anthropic, which also built the charts. Every quotation and figure was then checked against its source. The folio was signed off for publication by Ukweli's editor. To report an error, write to hello@ukweli.io. Corrections are dated and shown on the page.

Where a source limits its own finding, we have carried the limit with it. OpenAI says hallucinations can be avoided when models abstain. The Lancet authors say their audit finds fabricated references, not their cause. Neither South African court found that AI was used. Graphite did not test drafts that people edited.

Figures are as published on the dates given. The register changes weekly, which is why this folio links to it rather than counting it.

Sources #

  1. AI Now Writes as Many Online Articles as Humans — Graphite, Five Percent research, May 2026 — graphite.io/five-percent/research/ai-now-writes-as-many-online-articles-as-humans-do
  2. How Does AI-Generated Content Perform in Search and Answer Engines? — Graphite, October 2025 — graphite.io/five-percent/research/ai-content-in-search-and-llms
  3. More Articles Are Now Created by AI Than Humans — Graphite, October 2025, marked superseded — graphite.io/five-percent/research/more-articles-are-now-created-by-ai-than-humans · as archived 18 October 2025, web.archive.org/web/20251018053600/https://graphite.io/five-percent/more-articles-are-now-created-by-ai-than-humans
  4. Why language models hallucinate — OpenAI, 5 September 2025 — openai.com/index/why-language-models-hallucinate · paper: Kalai, Nachum, Vempala and Zhang, arxiv.org/abs/2509.04664
  5. News Integrity in AI Assistants — European Broadcasting Union, 21 October 2025 — ebu.ch/research/open/report/news-integrity-in-ai-assistants
  6. Topaz M, Roguin N, Gupta P, Zhang Z, Peltonen L-M. Fabricated citations: an audit across 2·5 million biomedical papers. The Lancet 2026; 407: 1779–81 (Correspondence), online 7 May 2026, issue of 9 May 2026; corrected version first online 16 July 2026 — thelancet.com. Figures here are from the corrected version as served on 18 September 2026
  7. Mavundla v MEC: Department of Co-Operative Government and Traditional Affairs KwaZulu-Natal and Others (7940/2024P) [2025] ZAKZPHC 2; 2025 (3) SA 534 (KZP), Bezuidenhout J, 8 January 2025, paras [20], [21] and [50] — saflii.org
  8. F.J.L v T.G.O (2025/220239) [2026] ZAGPJHC 875, High Court of South Africa, Gauteng Division, Johannesburg, Wright, Mahosi and Opperman JJ, 31 July 2026, paras [32], [35], [37], [62] and [122] — saflii.org
  9. Acting judge told to explain possible AI 'hallucinations' in judgment — Tania Broughton, GroundUp, 4 August 2026 — groundup.org.za
  10. Liang W, Yuksekgonul M, Mao Y, Wu E, Zou J. GPT detectors are biased against non-native English writers. Patterns 4(7), July 2023 — cell.com/patterns
  11. Guidance note on Artificial Intelligence — Press Council of South Africa, 28 November 2023 — presscouncilsa.org.za
  12. Exclusive: AI writing hasn't overwhelmed the web yet — Megan Morrone, Axios, 14 October 2025 — axios.com
  13. The internet is now mostly written by machines, study finds — Eric Hal Schwartz, TechRadar, 17 October 2025 — techradar.com
  14. Over 50 Percent of the Internet Is Now AI Slop, New Data Finds — Frank Landymore, Futurism, 14 October 2025 — futurism.com
  15. AI Hallucination Cases Database — Damien Charlotin, last updated 14 September 2026, read 18 September 2026; by-year figures counted by decision date from the database's published CSV (CC BY 4.0) — damiencharlotin.com/hallucinations

The standing block #

Read the register and download the data →

Integrity Watch is a verified register, neutral by design. Every entry is confirmed against a reputable source, and if it cannot be confirmed, it is not listed. Ukweli is a content integrity platform built by The Orkestra.

Get the next folio

One email a month from Integrity Watch: the new folio, and what changed in the register since the last one. Nothing else, and you can leave at any time.

The register itself stays free and public, and the full dataset is yours to download whether you subscribe or not.

Know the truth before you publish.

Ukweli checks documents for fabricated and unverifiable citations, sources, quotes and claims, before they leave your hands. Integrify a document · Get the next folio