Do LLMs deserve this much attention in academia?

by 

Friday, Aug 28, 2026 | 6 minute read
part of 2026 ii | #Computation

While the first public release of ChatGPT, which kicked off the AI1-boom, happened less than four years ago, Large Language models (LLMs) now seem to have permeated every part of our digital existence. While this is indubitably harmful for the environment and worrisome for a host of other reasons, this article focuses on the impacts of LLMs on academia. Investigations into the capabilities of LLMs are becoming more numerous, leading to a general waste of time and effort, publications of low impact, and may even hold back more rigorous research.

But first, what is a LLM? The increase in performance on language tasks exhibited by LLMs can mainly be attributed to two factors: the Transformer architecture and the size as measured in number of parameters. Transformers implement attention, which allows the model to process all words in the input in conjunction, instead of one by one. This permits more accurate modeling of the interplay between words in a sentence, paragraph, or text. Moreover, LLMs are trained on corpora of previously unheard of sizes and have billions of weights. It cannot be discounted that there is a direct correlation between model size and performance, which AI companies heavily indulge when spending millions of dollars on compute resources to train a single model. I want to alert you to the fact that this was highly uncommon before the AI-boom.

As is commonly known, LLMs have a number of problems, including their ability to confidently present hallucinated opinions as fact, the immense power and water consumption of data centers, copyright issues when companies nonconsensually scrape protected materials from the internet, and I could keep going on. Of particular note are the ties of AI companies, datacenter providers and hardware manufacturers with each other, as well as in the financial sector and government. While chat bots are generating losses, the companies who provide them gain trillion dollar valuations and make billion dollar deals with each other. This poses a significant risk to the stability of the economy.

Given the drawbacks of LLMs, why are they becoming the subjects of investigation across academic fields? There are many articles asking “Can LLMs solve X?”2 or reviews comparing the abilities of different LLMs. I think that such research categorically misunderstands the capabilities of LLMs and language modeling in general. Further, such publications do not seem a worthwhile investment of grant money and research time.

It is important to mention that LLMs present a generational performance improvement in language modeling tasks including text generation, machine translation, sentiment analysis and named entity recognition. Further, since LLMs are trained on vast datasets including repositories of knowledge (e.g., Wikipedia), it is to be expected that some of this knowledge is absorbed by the model and reproduced when generating text. LLMs are not trained to perform reasoning, however often Sam Altman and Dario Amodei tweet that superintelligent AGI is around the corner. It is, however, more reasonable to view the release of the Transformer architecture as a generational improvement, while interpreting the work of AI-companies as iterative gains since. It is to be expected that real AGI will take coordinated, principled efforts by the academic community around AI for years or decades to come.

I advise researchers to keep realistic expectations with respect to the capabilities of LLMs in mind when formulating hypotheses. As an example, while LLMs are capable of generating output which reads as legal text to the layperson, LLMs are not LawyerAI™︎ and should not be used to generate legal advice. There is now a plethora of investigations asking naively whether LLMs can solve complex tasks in various domains. Frustratingly, this often even involves data from other modalities such as image or video data. Such research commonly disregards established results and methods, while glancing over relevant nuances of the field. The answer is too often “yes” for toy examples, but “no” for cases which experts regard as interesting and complex.

I support the publishing of negative results. However this does not apply, when the result is a forgone conclusion. When such research realizes that LLMs cannot solve a highly complex problem, for which specialized methods already exist or are in active development, they often do not offer meaningful improvements. Finding out that LLMs cannot perform a complex task, such as solving a challenging reasoning problem, is not research, but to be expected due to their nature as language models and does thus not present a result in itself.

A respectable publication should not just note that LLMs can’t solve the issue, but provide a rigorous, principled, and well-researched method addressing the shortcoming. Otherwise we risk peer-reviewed research becoming a review outlet for private, closed-source, closed-weights models; that is to say: product-testing. This will lead to reproducibility problems and publications being outdated quickly, since LLMs are updated frequently and the guardrails which are placed around chat bots are not made public.

In general, the rate of submissions to journals have multiplied since the start of the AI-boom. Amongst these submissions are many written partially or entirely by LLMs, studies investigating the capabilities of AI, and methodologies relying on code written almost entirely by AI. Such submissions do not represent good faith contributions to academia, but aim to play the system to increase measurables. This overloads the unpaid volunteer peer reviewers. This is bad, since it slows down the review process and inhibits the publication of valuable contributions. Unfortunately, it also increases the already high pressure to publish and review faster and faster, creating destructive incentives for good-faith researchers.

It seems that the marketing of large AI companies has convinced not just users, but grant givers and researchers that LLMs are the future of intelligence. This is why we are suddenly seeing LLMs in every aspect of the academic process. This article is meant to point out the problems of using LLMs in academia. I urge you the reader to stay true to your field of expertise and see LLMs as what they are: tools to process language. In this light, I recommend you to continue exploring (hopefully LLM- free), well-founded hypotheses on the basis of prior, principled research. Lastly, I hope to convince you that when the AI-hype settles and realistic views of LLMs become more commonplace in the academic community, it will be the quality of your research that matters, not the quantity of papers you published with “AI” in the title.


  1. Please note that I am deliberately misusing the term here. AI is the well-established field of research into intelligent behavior of machines. LLMs fall under the umbrella term of AI (Machine Learning, Neural Networks, NLP) but are not representative of the field. In a case of semantic broadening it has become common to refer to LLMs as AI. ↩︎

  2. I encourage you to search “Can LLMs” on Google Scholar yourself. ↩︎

© 2025 - 2026 The Illogician

The student magazine of the Master of Logic at Amsterdam's ILLC