We value your privacy

    We use cookies to analyse traffic, improve our website and show relevant content. You choose what we may use. Read our privacy policy.

    Live AI news
    Back to articles// AI Opinie

    Model collapse: is AI leading to a dull average?

    How AI trained on AI output devolves into beige mediocrity - and why reasoning models can partially break that trend.

    Remy Gieling Published 15 mei 2026 Updated 15 juni 2026 8 min read
    Een unieke koraalrode figuur staat tussen vele identieke grijze silhouetten — visualisatie van model collapse en homogenisering door AI.

    Increasingly, the text, images, and code on the internet are created by AI. This same output is subsequently ingested by the next generation of AI models. The obvious question arises: does this create a kind of "regression to the mean" where the sharp edges disappear, leaving only a beige average?

    The short answer: yes, that risk is real. There is an official scientific term for it: model collapse. In the tech world, it is sometimes jokingly referred to as Habsburg AI after the European royal family that developed increasing genetic defects due to extreme inbreeding. Something similar happens with AI: the "genetic variation" of data disappears, making output less accurate and less creative.

    But there is also an important counter-argument: modern reasoning models are partially breaking this spiral. Below I explain both sides, supported by key research.

    What exactly is model collapse?

    Model collapse describes what happens when generative AI models are trained on data partially generated by previous AI models. Shumailov et al. demonstrated in Nature (2024) that models further trained over generations on synthetic data slowly lose the tails of the distribution the rare, strange, and surprising examples. What remains is what is statistically most probable.

    The nuances, the outliers, and the human "vibe" are filtered out. Not all at once, but gradually. It is precisely this pattern convergence toward the mean that is already recognizable in much AI-generated text.

    Three reasons why AI output gravitates toward the middle

    1. Reinforcement Learning from Human Feedback (RLHF)

    Most consumer chatbots ChatGPT, Claude, Gemini are fine-tuned with human feedback. Reviewers give a thumbs up or down, and the model learns which answers are considered "pleasant." This carries two side effects:

    • Standardization. Models develop a preference for a recognizable structure: short introduction, numbered list, neat conclusion. It works but it is the same everywhere.
    • Tone-leveling. Because the AI must not offend anyone and must remain objective, sarcasm, strong opinions, and stylistic outliers are sacrificed. The result is that "beige" tone that can now be recognized within half a paragraph.

    2. Sycophancy: the AI as a yes-man

    A related problem is sycophancy: the model prefers to confirm your assumption rather than contradict you. This is also an unintended consequence of RLHF aligning with the user generally scores better with human reviewers than a critical rebuttal.

    The effect: a digital echo chamber where social desirability triumphs over factual sharpness. For those wanting to use AI to think better, this is a serious issue.

    3. The Collective Novelty Paradox

    Doshi and Hauser published a study in Science Advances (2024) regarding AI in creative writing. Their finding is remarkably dualistic:

    • At the individual level: stories became measurably better smoother, grammatically more correct, more enjoyable to read.
    • At the collective level: the stories looked much more alike than texts written purely by humans.

    As the authors summarize: AI raises the floor but lowers the ceiling. Fewer poor texts but also fewer unique outliers. They call this the decline of social diversity in creativity.

    The Alignment Tax: reliable or creative, not both

    There is a technical parameter that illustrates this well: temperature. This determines how "freely" a model can choose the next word.

    • Low temperature (≈ 0.1). The model consistently chooses the most probable next word. Safe, predictable and boring.
    • High temperature (≈ 0.9+). The model is allowed to take improbable paths. More surprising, more creative but also more frequently factually incorrect. Research (including Shumailov et al., 2024) shows that this is precisely where models "go off the rails."

    Researchers call this the Alignment Tax: the price in creativity and diversity paid to obtain a factually reliable model. A new movement even speaks of good hallucinations the proposition that a certain degree of "invention" is necessary for true creativity, because creativity by definition involves something not literally present in the training data.

    Why reasoning models partially break the spiral

    So much for the "yes." Now for the "no" or at least the nuance.

    Previously, generative models were quite linear: prompt in, transformer model applied, answer out. This resulted in both many hallucinations (output was never cross-checked) and generic texts.

    The current generation of chatbots operates differently. When you ask a question, a reasoning or thinking mode is usually active. The model first reflects:

    • What information do I need?
    • Do I need to look something up to answer this?
    • Which stylistic choice fits here formal, informal, technical?
    • Is text even the right format, or would an interactive overview work better?

    This reasoning step has two effects. One: hallucinations decrease because the model can explicitly say "I am not sure about that, let me look it up." Two: the model does not automatically pull the most standard style out of the bag but considers that a degree of variation might be appropriate.

    In short: an input → output model is bland and degrades quickly. An input → reasoning → output model is a step in the right direction to prevent this where desired. Though it should be noted that more creative models often hallucinate more frequently.

    What does this mean for you as a user?

    Three practical consequences for those using AI daily:

    1. Use AI as a sparring partner, not an author. Let the model sharpen your thinking, not take over your voice. Write both the first and the last version yourself.
    2. Actively ask for contradiction. "Give three reasons why this is a bad idea" works better than "what do you think?" the latter activates sycophancy.
    3. Consciously choose a reasoning model for complex questions. For quick tasks, a fast chatbot suffices. For strategy, analysis, or substantive nuance, a reasoning model is almost always worth the effort.

    Conclusion

    Model collapse is not a conspiracy theory it is a measured effect, supported by a growing list of peer-reviewed research. At the same time, it is not an inescapable fate. Reasoning models, conscious prompting, and the discipline to not abdicate your own judgment are three powerful counter-forces.

    The question is not whether AI gravitates toward the average. It does. The question is whether you as a user, writer, and thinker are willing to do the work to rise above it.

    Key sources

    • Shumailov et al. (2024). AI models collapse when trained on recursively generated data. Nature.
    • Doshi & Hauser (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances.
    • Research on RLHF and the Alignment Tax (various, including Anthropic & OpenAI).
    • Recent literature on good hallucinations and reasoning models (OpenReview, 2025–2026).

    Thinking further about AI risks

    Questions about future-proofing your AI strategy? Read our piece on the biggest concerns about AI, or book a session via AI Consultancy and train your team with a tailored AI Workshop.

    Remy Gieling — Mede-oprichter, AI-expert & bestseller-auteur bij ai.nl

    // About the author

    Remy Gieling

    Mede-oprichter, AI-expert & bestseller-auteur

    Tech-expert (1988) gespecialiseerd in kunstmatige intelligentie en mede-oprichter van ai.nl, The Automation Group, Proxies en eBrain.ai. Oud-hoofdredacteur van diverse zakenmerken en daardoor een geoefend verteller op het podium en in de media. Verzorgt jaarlijks 150+ AI-keynotes in binnen- en buitenland en is gastdocent aan Nyenrode. Co-auteur van zeven boeken, waaronder 'Handboek AI Strategie' en 'AI Agents', en bekend als presentator op radio en RTL Z. Reist langs de labs van OpenAI, Nvidia en Tencent en vertaalt de nieuwste doorbraken naar inzichten die leiders direct kunnen toepassen.

    LinkedIn
    // GET STARTED// How we can help

    Beyond reading — let AI work for you.

    // CONTINUE READINGAll articles

    More from AI Opinie.

    Newsletter

    Always up to date on AI.

    Once a month: cases, frameworks and concrete examples of what works in practice. No noise.

    No spam. Unsubscribe any time.