We have featured pieces on the “enshittification” of social media platforms with advertising slop degrading the user experience. Now AI is only accelerating it. There is a rapid increase in written content on LinkedIn that most of us can recognise as AI generated. Derek Thompson interviews Max Spero, the founder of Pangram, the most used tool to check for AI content on the internet, to understand how bad the problem is, why we should care to stop it and how Pangram does it.

First, how bad is the problem: “If you look at the scale of AI and bot traffic on the internet, it’s gone from a small fraction of human traffic to recently passing 50 percent. I think it’s not long before 99% of internet traffic is bots. We’re at risk of the “dead internet theory,” where the internet is just this echo chamber of bots talking to bots.

…In the last few weeks, the Wall Street Journal has reported that “AI has plunged the book publishing industry into utter chaos.” The New York Times reported that Spotify, LinkedIn, and others are “trying to dig out of a digital sewage heap full of low-quality content made by artificial intelligence.” The Times cited Pangram, which reported that AI was used to create nearly half of the posts with more than 50 words on X or Twitter.

…LinkedIn had 41% of its long-form content generated by AI. We did a study with Stanford and the Internet Archive, which found that in May of last year, nearly 40% of the internet was AI-generated. I think this number is just going to keep expanding. There are people trying to pollute the internet for different reasons, whether they want their own narratives in the LLM training data or they’re trying to win at SEO. Now it’s easier than ever to produce AI content for search results.”

LinkedIn in particular has been the biggest victim. Many of us got off Facebook several years ago but found LinkedIn still a decent civilised way to stay connected and updated about our social/professional network. Now, it is rife with what is obvious AI slop.

“Why is LinkedIn more full of AI content than Twitter? I think it’s because there’s this incentive people feel: Posting on LinkedIn regularly is going to improve my career prospects. It’s going to make me more likely to get a good job because I’m going to be visible to people who are relevant in my career. If posting gives you an advantage to your career, people are going to find the easiest way to post more, which ends up being using AI.

Today’s algorithms also incentivize volume and quantity over quality. Hopefully we’re going to see that change in the coming years as quantity becomes free, while quality is really the limiting factor.”

Besides the ‘enshittification’ of the platform, what’s wrong with AI generated content? Here’s Thompson ranting about it: “I care about the slopification of the written word. It matters for education, as I think it would send a terrible message to young people (and their teachers!) in the trenches of high-school English classes if the most esteemed newspapers and publishers simply didn’t care if essays published under your name were effectively composed by an electrical exchange taking place inside some faraway data center. The act of writing is an act of thinking, and to automate our paragraphs is to populate our screens with words without filling our minds with the ideas necessary to conceive of them. Finally, I do not enjoy the idea of a technology that severs the connection between composition and knowledge, such that I cannot know if an author actually knows anything about what they have published under their name.”

How does Pangram figure out if something is AI generated?

“Pangram is a classifier model. It’s looking at text and saying, “I believe this is AI-generated, AI-assisted, or human-written.” It’s not doing any generation. It’s not an LLM, like ChatGPT.

We have a training set of writing that we know was written by a human. For example, a five-star Yelp review about Denny’s or a 500-word essay on Moby-Dick. Then in each case, we’re picking an AI model and asking it to write something similar. We’re going to ask ChatGPT to write a five-star review about Denny’s. We’re going to ask Claude to write a 500-word essay on Moby-Dick. Pangram compares the human and AI items, which are similar in topic but different in style, and we’re learning the difference between how AI produces text and how humans do.

If you look at the broad field of AI systems, there are models that discriminate. For example, Waymo has cameras and a pedestrian detection system that takes an image input and says either, “Nope, there are no pedestrians,” or, “Yes, there are pedestrians, here they are.” That’s closer to what Pangram is than ChatGPT. Pangram isn’t producing anything. It’s making a judgment based on an input.”

Spero elaborates on this. He shares some of the hallmarks of AI writing which we can relate to: “The early ChatGPT and AI models were trained on this instruction-tuned data set and would overuse some words and phrases. For example, “delve,” “tapestry,” “intricate.” They really loved these individual words. If you saw the word “delve” in an out-of-place setting, that was an immediate red flag to some people.

Maybe a year later, they were able to hammer out these word-level inconsistencies. But there were still patterns. For example, “It’s not just X, but Y” is called a negative parallelism, and LLMs love it because it has a lot of impact on the reader. Sentence constructions that use em dashes were associated with good writing. Now I think they’ve even hammered some of these out. So I’m relying on longer-context signals. It’s at the paragraph or sentence level. I could look at the shape of the text and tell you that looks like ChatGPT or Claude.

…Something I’ve noticed is that long pieces of AI writing often try to summarize and re-summarize. The tell is at the scale of the paragraph. They’ll be like, “This is genuinely important.” “This is the main course.” “The bottom line is…” Every sentence is trying to outdo the previous one in summarizing what it’s trying to say better.”

If you want to read our other published material, please visit https://marcellus.in/blog/

Note: The above material is neither investment research, nor financial advice. Marcellus does not seek payment for or business from this publication in any shape or form. The information provided is intended for educational purposes only. Marcellus Investment Managers is regulated by the Securities and Exchange Board of India (SEBI) and is also an FME (Non-Retail) with the International Financial Services Centres Authority (IFSCA) as a provider of Portfolio Management Services. Additionally, Marcellus is also registered with US Securities and Exchange Commission (“US SEC”) as an Investment Advisor.