Sumanth Pandugula (@summyfeb12) on X

X (formerly Twitter) ·

7 min read Original article ↗

TL;DR - Pay attention to the idea or concept the AI-generated text is trying to convey, not the syntax around the idea like the language and vocabulary used.

Human Brain

Ideas and thoughts are complex and abstract multidimensional entities in the human brain. An individual can construct multiple such entities, often weaving them many times through several layers in their head. But how do they convey the same rich information to the person next to them? Let’s say I have an idea about how the future of human life could look in a couple hundred million years. It’s very easy to build all the layers in my head. For example, I can imagine intergalactic travel using large ships, watching sunrises on Mars like it’s a day trip, and taking a vacation on the most beautiful planet in the Andromeda galaxy and then heading home to my under-ice, underwater pod on Europa. For you reading this, you would be building a very different image of what those things would look like. But if we both could show how we imagined this in our head, they will vary a lot on multiple levels starting with how many space movies we watched to how imaginative we are to our ability to describe it via a medium such as text, images etc and our expertise in each. When beings of the same or various kinds started to communicate with each other, they used different mechanisms; we can see that across evolutionary history. Smaller fish communicate by moving their whole bodies, and whales communicate using songs. Early humans conveyed their thoughts through cave paintings, signs, etc. Over the last fifty thousand years, we have come a long way from images. Yet the most spoken language by humanity so far relies on 26 letters and a few hundred thousand words.

Lossy Language

I read this article "Neuralink and the Brain's Magical Future" by @waitbutwhy 9 years ago, back in 2017, and the main takeaway for me from that was that ideas and thoughts are very high resolution in the human brain, and when we convey them to others using language, we are losing a lot of that resolution to the medium, such as language - and that’s before factoring in things like that individual’s expertise in the above-said language, the strength of their vocabulary, and skills like reading, writing skills, storytelling, drawing, etc. The goal of @neuralink was to eliminate the use of language and communicate ideas and thoughts without losing any resolution. Maybe we will have that someday, but let’s talk about what we have today and what we are trying to do.

We are adding more words to the English vocabulary every year, and we may keep doing that for a lot longer from now, but the most spoken language in the world is only spoken by 20% of the population; the rest of the world speaks different languages. We have been trying to solve the problem of translation for years, and with AI we are finally getting somewhere. Yet keep in mind, you are already losing a lot of resolution when converting ideas from your mind to your native language, and you lose a good chunk of it again in translation.

Level Playing Field

When generative AI became mainstream after @openai released ChatGPT. A lot of people find it easier to convey their ideas to other people. Because AI can write your ideas elaborately, generate images, and translate without the barriers of expertise in the language. I do understand the obsession with penalizing the AI-generated text. Frontier models trying to sign the text generated using their models, @Pangram detecting the AI-generated text and the whole campaign calling them slop instead of trying to understand the idea in the first place. I think humans are partially responsible for what we are seeing with AI-generated text today. William Shakespeare could have very well conveyed how Brutus was manipulated into backstabbing Caesar in less than 200 words; he used 20,000 and I had to spend good 2 years in high school reading it. I am sure JK Rowling could’ve conveyed her message through Harry Potter by not writing a million words. This could be true for all authors, but they had a specific reason to do it and the context is different in every scenario. We trained our models using text written for print media and that’s not what is expected of it in social media. The writer who uses AI should be solely responsible for limiting the words where necessary and thoroughly proof-reading it to make sure the idea they want to send across with the level of detail they intended before they send it across. I do acknowledge the onus on the reader to go through hundreds of lines of generated text to understand the idea if the writer doesn’t do their job; in that case, I leave it up to the reader to choose between consideration and turning it down. The goal of language is to convey ideas and thoughts across. It feels like a lot of importance lately is given to the ship that carried the cargo instead of the cargo itself.

Cargo Over the Carrier

A very well-thought-out concept or idea articulated using AI is always better than a poorly thought-out concept articulated very well due to their expertise in language. It’s the thought and idea that needs to be rewarded, not the expertise in language. AI makes that possible for 80% of the world. It makes it possible irrespective of the native language they speak, their ability to read, write, or their neurodivergence. You don’t need to be an expert in a language or its grammar rules to convey your ideas across.

Recently, all frontier model makers including @anthropicai , @GoogleDeepMind , @openai , @grok are secretly signing and/or watermarking AI-generated text and media with signatures and using them to detect generated content. I wish we could move towards the detection of the origin of the concept or idea. Whatever watermark they are using to detect should include a breadcrumb trail of the prompts or at least the keywords in the prompts that led to the response. Can the models detect if some text/media is AI-generated? Yes, and what value does it give the reader, except that the language is AI-generated? But can it also give some details about the idea behind it,

Can it detect if the idea is original or plagiarized?

Can it give the actual citations/references for this idea?

Can it show the keywords in the original prompts that led to this response?

That’s where the actual value of these tools lie.

What Next?

We have come a long way from communicating using hand signals and cave paintings to what we have as the most evolved language today. We were able to make incredible progress in science, technology and art, explore space, and build nation-states with the language we developed and tools we had at our disposal. This is what we have achieved with a lossy compression medium like language. AI is breaking down a lot of barriers in grammar, translation, vocabulary and even the accent with the latest voice models. The future is certainly headed towards the ultimate goal of conveying ideas between each other without losing a lot of resolution to the medium. I can only imagine what's in store for humanity. But in the meantime, we should reward the semantics and not the syntax around it.

P.S. This article is fully human (me) generated and pangram can vouch for me.