The Machine, the Learning Illusion and the 30,000-Paper Pile

When I sat down to write a book on GLP-1 receptor agonists, the goal was a single book: a referenced text written from a pharmacologist's perspective, running from molecular mechanism to clinical endpoints. What came out was a 14-chapter, 48-pages book with 71 references fully typesetted book. The whole thing is here:
📕 GLP-1 Receptor Agonists — the full book (PDF, 48 pages, in Turkish)
But this post isn't about the book. On one side, evidence is accumulating that AI use degrades learning. On the other, I used the same tools to write a book and came away understanding the topic better than I would have by the conventional route. That isn't a contradiction: the harm the evidence points to comes from how the tool is used, not from the tool. That's the actual claim of this post — and the book is a concrete instance of it.
First the problem: the GLP-1 literature is no longer a readable pile
While writing this post I checked PubMed (23 July 2026, title/abstract search): there are 30,192 records on GLP-1 receptor agonists. 1,338 of them are systematic reviews or meta-analyses, and 1,011 of those were published in the last five years alone.
Those numbers break the standard advice. "Read the systematic reviews" now means putting 1,338 titles in front of you. Working out which ones are sound is a separate job that comes before learning the subject itself.
And that triage is harder than it looks. A concrete example, straight from this post's own topic: the meta-analysis most widely circulated as evidence that "ChatGPT improves learning" — 51 studies pooled, peer-reviewed, published in a Nature-portfolio journal — was retracted in April 2026, because discrepancies in the meta-analysis undermined the editor's confidence in the validity of the analysis (retraction note). "Peer-reviewed" and "meta-analysis" are not, by themselves, a sufficient filter.
That is the first argument for writing a book for yourself: a text that filters the pile against your question, and leaves the judgement to you, turns out to be more useful than an off-the-shelf review.
So what does AI actually do to learning?
This deserves an evidential answer rather than an emotional one. The best study we have was, as it happens, run in Turkey.
Bastani et al. (PNAS, 2025) ran a field experiment with nearly 1,000 students across 9th, 10th and 11th grade at a large high school in Turkey. Two AI tutors were deployed: "GPT Base", mimicking a standard ChatGPT interface, and "GPT Tutor", prompted to safeguard learning by offering teacher-designed hints instead of answers. The result has two layers. While the assistant is available, performance improves sharply (48% grade improvement with GPT Base, 127% with GPT Tutor). But when access is taken away, the GPT Base group performs 17% worse than students who never had access at all. The safeguards in GPT Tutor largely eliminate that damage. In the authors' words, without guardrails students use GPT-4 as a "crutch".
That single finding is the thesis of this post: the harm comes from receiving finished answers, not from the tool.
A few more studies sit alongside it:
- Gerlich (Societies, 2025), a mixed-method study of 666 participants, found a significant negative correlation between frequent AI use and critical thinking scores, with cognitive offloading as the mediating link. It is cross-sectional, so it shows a pattern rather than causation.
- Lee et al. (CHI 2025) collected 936 real usage examples from 319 knowledge workers: higher confidence in the AI predicts less critical evaluation, while higher confidence in one's own expertise predicts more.
- Kosmyna et al. (2025) tracked 54 participants with EEG across four sessions; the LLM group showed the weakest neural connectivity, struggled to quote correctly from essays they had written minutes earlier, and reported the lowest sense of ownership over their own text. This one is a preprint, with a small sample and published methodological criticism — so it belongs in "preliminary findings" language, not assertive language. That is the evidence-tier rule described later in this post, applied here.
One caveat matters: there is no evidence that "AI reduces your synapses". What has been measured is EEG-level connectivity, recall, and exam performance after access is withdrawn. But the practical conclusion lands in the same place — delegate the mental effort that produces learning, and the learning goes with it.
Two old findings about where learning actually happens
The picture makes more sense next to two classic results from the psychology of learning:
- The generation effect (Slamecka and Graf, 1978): information you produce yourself is retained better than information you merely read.
- The testing effect (Roediger and Karpicke, 2006): retrieving material from memory improves long-term retention substantially more than re-reading it.
What they share is that the effort is the mechanism. A finished answer removes exactly that mechanism. The difference between GPT Base and GPT Tutor is precisely this: one eliminates the effort, the other directs it.
Why writing a book is a different mode of use
When you write your own book, generation and retrieval stay with you. What gets automated is the part that has nothing to do with learning:
| Stays with me (where learning happens) | Automated (mechanical load) |
|---|---|
| Which chapter goes in, how much weight each debate gets | File structure, LaTeX patterns, the build chain |
| Writing the sentences, explaining a mechanism in my own words | Pulling reference metadata, DOI verification, citation auditing |
| Clinical interpretation: how much a signal matters in practice | Style consistency in figures, format auditing |
| Which sources to discard | Sweeping 1,338 reviews down to a candidate pool |
The gain runs both ways: the weeks that would have gone into literature sweeps and formatting come back, while the load of writing, deciding and interpreting — the part that exercises the synapses — stays where it was. And the book bends to my needs: it goes deep where a pharmacologist needs depth and passes over the rest. You cannot tune an off-the-shelf review that way.
Put in the language of the Bastani study: the kitap-yazma skill is my guardrail. Not to withhold answers, but to preserve the order, the rules and the verification while leaving the decisions to me.
The book itself: GLP-1 agonists
The outline was: introduction and definitions, history, incretin physiology, PK/PD, molecular mechanism, structure and chemistry, clinical evidence in type 2 diabetes, obesity, cardiovascular outcomes, neurological effects, adverse effects, drug interactions, special populations, and future directions.
I went deep where it matters for a pharmacologist: the two-step binding mechanism of GLP-1R as a Class B GPCR, biased signalling between the Gs and β-arrestin pathways, what glucose-dependent action actually means when compared with sulfonylureas, design differences across the cardiovascular outcome trials, and directly practical topics such as pre-operative aspiration risk.
The technical side:
\documentclass[10pt,twoside,twocolumn]{book}, XeLaTeX,polyglossiafor Turkish- Palatino as the body font — a textbook feel that holds up over long reading
- Five
tcolorboxcallout types: clinical, mechanism, evidence, warning, and summary table - Citations via
natbib+ BibTeX, 15 tables, 18 figure slots
None of this skeleton was designed from scratch. It settled while writing the Drug Safety in Preganancy and Lactation book, was copied into the PROTAC book, and used a third time for GLP-1. By the third repetition it was clear this wasn't a template — it was a process.
Figures: every figure comes from somewhere
The most time-consuming part of a scientific book isn't the text, it's the figures. There are two legitimate routes, and they should not be blurred together.

The first route is adapting a figure from a paper. In that case the source is cited explicitly and an "adapted from" note goes under the figure. The second route is generating a new figure from text you wrote yourself. For that I used Gemini's image generation model.
The real trick here isn't the choice of model — it's freezing the style template. I wrote a single STYLE block: minimalist medical illustration, pastel palette (cream, light blue, pale green, dusty pink), flat white background, dotted right-angled annotation lines, Turkish labels only, text large enough to read on a phone. Every figure request reused that block unchanged. The only thing that varied was the subject description appended below it.
The result: dozens of separate figures that look like parts of one book. Style consistency is half of a book's identity. (The three figures in this post were produced the same way, with the same style template.)
Then the skill: encoding the repetition
After three books I had an unwritten checklist: gather sources first, then revise the chapter titles against the literature, then build the scaffold, then write chapters, then produce figures, then compile, then audit. Every time, I was trying to remember that list. Instead of remembering it, I wrote it down.
kitap-yazma ("book writing") is a Claude Code skill that splits this flow into phases:

| Phase | Work |
|---|---|
| 0 | Setup and intake: book title, length, audience, tone, part structure |
| 1 | Multi-channel source pool: scite, PubMed, CrossRef, review PDF mining, industry sources, patents |
| 1.5 | Locking the outline |
| 2 | Scaffold: main.tex, front matter, chapter files, references.bib |
| 3 | Chapter writing |
| 4 | Figures: adapted + generated |
| 5–6 | Build (xelatex → bibtex → xelatex → xelatex) and quality audit |
Small Python scripts sit underneath the phases: status tracking, turning intake answers into configuration, scaffolding, source retrieval with DOI verification, figure extraction from PDFs, building, citation auditing, abbreviation auditing. All standard library — so there is no setup burden.
The part that matters most: zero fabrication
The most dangerous capability of a language model in scientific writing isn't being wrong — it's being plausible. A DOI that doesn't exist, a journal name that could plausibly exist, a year in exactly the right format. You cannot spot it by eye.
This isn't an impression; it has been measured. Walters and Wilder (Scientific Reports, 2023) checked 636 citations generated across 42 topics one by one: 55% of GPT-3.5's citations and 18% of GPT-4's were entirely fabricated, and among the real ones, 43% and 24% respectively carried substantive metadata errors. On the medical side, Chelli et al. (JMIR, 2024) ran the same test on systematic review references: hallucination rates of 39.6% for GPT-3.5, 28.6% for GPT-4 and 91.4% for Bard. Asking for sources is not the same as getting sources.
That's why the skill's strictest rule sits on the source side:

The flow is one-way: find → check for retraction and resolve the DOI → pull metadata from CrossRef → deduplicate by DOI → tag the evidence tier → sync to Zotero. No source enters references.bib before its DOI resolves.
The tier tag isn't just a ranking; it determines the wording. A claim resting on a systematic review can be written assertively; a claim resting on a preprint is written as preliminary — as I did above with the Kosmyna study; a phase-status figure from a company blog isn't written at all until it's cross-verified. Anything unverifiable is flagged, not hidden.
Why the retraction check sits at the front of that flow is visible earlier in this post: a paragraph resting on a retracted meta-analysis takes the whole text down with it.
The same logic applies to computational sections: instead of presenting a method I couldn't actually run on my own machine as if I had run it, the method is described, with the fact that it wasn't run stated plainly.
What I did not automate
The skill's job isn't "write the book" — it's "keep the order and the rules while the book is being written". It stops at the end of every phase and waits for approval. The decisions stay with me:
- Chapter titles and ordering are finalised after the source sweep — because the literature usually shows that the first draft went deep in the wrong place.
- Which debates make it into the book, and how much weight each study gets, is an editorial call.
- Clinical interpretation — how much a given signal actually matters in practice — isn't something to automate.
What gets automated is the mechanical part: file structure, LaTeX patterns, reference metadata, citation auditing, the build chain, style consistency. Not the book itself, but everything around the book.
Takeaway
The question about AI in education isn't "should you use it" but "which job are you handing over". The harm in the evidence comes from receiving finished answers: when access is withdrawn, nothing is left behind. Hand over the literature sweep, the formatting and the verification while keeping the writing, the deciding and the interpreting, and the picture inverts — you gain the time and you don't lose the learning.
Writing a book for your own needs is a fairly pure instance of that: you learn the subject without drowning in 1,338 reviews, you fix what you learned in your own sentences, and you are left with a text you can return to in which every sentence has been verified. The defence against hallucination isn't staying away from AI either; it's building a flow that resolves every DOI, drops the retracted papers, and flags whatever cannot be verified.
I've built the same pattern for thesis writing and thesis statistics too. They share one backbone: split into phases, stop at the end of each one, fabricate nothing, verify the source, leave the decision to a human.
References
- Bastani H, Bastani O, Sungu A, Ge H, Kabakcı Ö, Mariman R. Generative AI without guardrails can harm learning: Evidence from high school mathematics. PNAS. 2025;122(26):e2422633122. doi:10.1073/pnas.2422633122
- Gerlich M. AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking. Societies. 2025;15(1):6. doi:10.3390/soc15010006
- Lee HP, Sarkar A, Tankelevitch L, et al. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers. CHI 2025. doi:10.1145/3706598.3713778
- Kosmyna N, Hauptmann E, Yuan YT, et al. Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task. arXiv preprint. 2025. arXiv:2506.08872 — not peer-reviewed.
- Slamecka NJ, Graf P. The generation effect: Delineation of a phenomenon. J Exp Psychol Hum Learn Mem. 1978;4(6):592–604. doi:10.1037/0278-7393.4.6.592
- Roediger HL, Karpicke JD. Test-Enhanced Learning. Psychol Sci. 2006;17(3):249–255. doi:10.1111/j.1467-9280.2006.01693.x
- Walters WH, Wilder EI. Fabrication and errors in the bibliographic citations generated by ChatGPT. Sci Rep. 2023;13:14045. doi:10.1038/s41598-023-41032-5
- Chelli M, Descamps J, Lavoué V, et al. Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis. J Med Internet Res. 2024;26:e53164. doi:10.2196/53164
- Retraction Note: The effect of ChatGPT on students' learning performance, learning perception, and higher-order thinking: insights from a meta-analysis. Humanit Soc Sci Commun. 2026. doi:10.1057/s41599-026-07310-z