For most of the past three years, the question of how much artificial intelligence quietly writes science has been answered in careful, modest percentages. A preprint posted in the middle of August offers a very different number, and it is large enough to change the shape of the conversation.
What Happened
On 20 August, Nature reported on a study estimating that almost nine out of ten English-language papers published in December 2025 and archived in PubMed Central show signs of large language model assistance. Across the whole of 2025 the estimate falls to 77 per cent. For 2024 it is 52 per cent. The work, by Linus Holzwarth, Rita González-Márquez and Dmitry Kobak, was posted to arXiv on 12 August and has not yet been peer reviewed.
The underlying technique is familiar. Since 2023, researchers have tracked model usage by counting how often certain words spike in the literature — the vocabulary that language models reach for more readily than people do. What changed here is the sampling and the estimator. The team analysed full texts rather than abstracts alone, and used a method that returns direct estimates of usage rather than lower bounds.
That second change accounts for most of the gap with earlier work. A 2025 paper in Science Advances, co-authored by some of the same researchers, put the figure for 2024 abstracts at at least 13.5 per cent. Re-run with the newer, more sensitive method, that same 2024 abstract figure rises to 31 per cent. The old number was never wrong; it was a floor, and the floor was low.
The study also broke its estimates down by section, and this is where the results become more interesting than any single headline percentage. An estimated 78 per cent of discussion sections in December 2025 papers carried signs of model assistance, against 58 per cent of results sections. Abstracts, introductions and discussions ran consistently ahead of methods and results — the parts of a paper where prose does the most work sit at the top of the range.
Kobak told Nature that his first response to the numbers was disbelief. “I was sure that we did something wrong,” he said. Repeated checks left the estimates standing.
Why It Matters
A figure this high stops being a story about misconduct and becomes a story about infrastructure. At 13 per cent, model-assisted writing is a behaviour that policy can single out and manage. At something approaching 90 per cent, it is the default condition of a literature — closer to spellcheck or reference managers than to a practice that can be flagged, disclosed and contained. Rules written for the first situation do not transfer cleanly to the second.
The section-level split is the part worth dwelling on, because it separates two activities that the word “AI-assisted” flattens together. Smoothing the English of a discussion section is a labour question. Generating a results section is an evidentiary one. Kobak singled out the results figure as the concerning number, given how readily models fabricate specifics that look plausible in context. Fifty-eight per cent is not proof that data are being invented, but it is a large surface area on which invention would be hard to see.
Introductions carry a quieter risk. They set up what a field considers settled, contested and worth doing next — and they now appear to be among the most heavily model-touched parts of a paper. “Whatever bias the LLM may have will just suddenly permeate the literature,” Kobak warned. If a few models write the framing for most of biomedicine, the field’s sense of its own open questions starts to converge on whatever those models find most probable.
The caveats run in both directions. Detection rests on stylistic fingerprints, so a heavily edited model draft can pass as human, while a researcher who has absorbed the cadence of these tools may be counted as one of the nine. The estimate is best read as a measure of stylistic influence rather than a census of authorship — which makes it, if anything, a harder thing to write a rule about.
Reaction
Researchers who spoke to Nature found the scale plausible while urging caution about how far it generalises. PubMed Central is a large and important slice of the literature, but it is a slice; the study covered English-language papers only, and more analysis will be needed before the figure can be extended across disciplines.
Kyle Siler, a social scientist at the University of Toronto whose own 2026 study in PNAS estimated that 57 per cent of 2025 papers across academic disciplines were probably AI-influenced, was blunt about where this leaves the field. “The toothpaste is out of the tube, and it’s not going back,” he said.
What makes the measured rates striking is how far ahead of self-reporting they run. A 2025 Wiley survey found that 71 per cent of researchers said they use AI for writing assistance — already a majority, and still well short of the December estimate. The preprint’s authors take the consistency between the two as supporting evidence, and suspect the true rate exceeds what people are willing to state on a form.
That gap is the reaction worth watching. It is not evidence of concealment so much as of a norm still being negotiated in public. Disclosure policies differ sharply between journals, and a researcher with no clear guidance and no obvious harm in view has little reason to volunteer that a model tightened their prose.
What’s Next
The most immediate question is what disclosure is even for. A checkbox that nearly every author ticks conveys almost nothing. If the useful distinction is between editing and generating — and the section-level data suggest it is — then policies will need to ask which parts of a paper a model touched, not whether one was opened at all.
The second question is verification rather than detection. Detectors are in an arms race they are structurally unlikely to win, and the results-section finding points somewhere more tractable: checking whether the numbers in a paper match the data behind it. That is slow, unglamorous work, and it is closer to what peer review was always supposed to do.
The study itself needs extending before its headline figure can carry much weight. Peer review is pending. Non-English literatures are unexamined, which matters especially because model assistance is most valuable to researchers writing in a second language — arguably the clearest good in this whole picture. And a single month, December 2025, is a thin basis for a trend line, even a steep one.
Closing Thoughts
There is a temptation to read a number like this as decline, and it is worth resisting. Scientific writing has never been the part of science that carried the truth; the data, the methods and the willingness of other people to check them do that. Prose is the delivery mechanism, and it has always been shaped by whatever tools were to hand — the typewriter, the word processor, the reference manager, the non-native speaker’s dictionary.
What is genuinely new is uniformity. A tool that helps everyone write the same way does not just save time; it narrows the range of what gets said, and it does so invisibly, in the sections where fields decide what they think they know. That is the finding worth carrying forward from this preprint — not that nine in ten papers had help, but that the help is starting to sound the same.
한글 요약
8월 20일 네이청는 8월 12일 arXiv에 게재된 노문 한 편을 소개했습니다. 급트 대학교의 Dmitry Kobak 연구팀 등이 PubMed Central에 색인된 영문 녕문을 분석한 결과, 2025년 12월에 출파된 녕문의 약 90%가 대형 언어 모델(LLM)의 도움을 받은 헬적을 보왜습니다. 2025년 전체는 77%, 2024년은 52%엠습니다. 이전 연구가 제시한 13.5%와의 차이는 이번 연구가 초록이 아닌 직접추정치를 내는 방법을 써고, 초록이 아닌 전문을 분석했기 때문입니다. 해당 노문은 아직 동료 평가를 거치지 않았습니다.
팅 미세한 결과는 생산별 분포업니다. 2025년 12월 녕문의 녤의(discussion) 삹션 78%가 LLM 헬적을 보왜으면, 결과(results) 삹션은 58%엠습니다. 초록로가나 서론처럼 산문의 비중이 높은 부분이 상위권에 있어서, 문장을 다도는 작업과 내용을 생성하는 작업으로 나낔어 봄 필요가 있다는 점을 보여중니다. Kobak은 결과 삹션 수치를 가장 우려스렕게 보약습니다. 한편, 서론 부분의 녕은 비죲은 학계가 무엇을 미해결 문제로 보는가에 편향이 스밨 수 있다는 우려로 이어직니다.
토론토 대학의 Kyle Siler는 이미 된돌린 변화라고 평가했습니다. 다만 이 추정은 문체의 스타일을 데이터로 한 것이어서, 저자권 자체보다는 문체 영향력을 측정한 검과로 읽는 것이 정확합니다. 영어권 밝 녕문과 비영어권 연구자, 다른 학문 분야로의 확장은 아직 남아 있습니다. 참고 출처: Nature(2026.8.20), Holzwarth 외 arXiv 프리프린트(2026.8.12), Kobak 외 Science Advances(2025), Siler PNAS(2026).