Are they just Stochastic Parrots? (with Fabio Massimo Zanzotto) | Dario Zanca Podcast Ep. 13 [🇮🇹 video & 🇬🇧 AI-curated transcription]


The following is a summarized and AI-curated English transcript of the podcast episode. While it aims to capture the key points and essence of the conversation, this is not a verbatim transcription. Some sections may be rephrased for clarity and readability.


Host: Dario Zanca
Guest: Fabio Massimo Zanzotto


Dario: Fabio, your research focuses on Natural Language Processing. In particular, you’ve studied large language models—LLMs—which are often seen as a new form of artificial intelligence. A major debate today is whether these models truly understand language, or if they’re just “stochastic parrots,” repeating patterns based on statistics. What’s your take?

Fabio: I initially hoped these models could offer insights into how the human brain works—especially how we learn language. But over time, I’ve become skeptical. I now see these models more as large language memories rather than models that understand. They store immense amounts of human expression and retrieve it in response to input. That doesn’t imply comprehension. They respond well when someone, somewhere, has already expressed something similar. It’s memory, not reasoning. So maybe it’s not artificial intelligence we’re seeing—it’s collective human intelligence, repackaged.

Dario: But what about when we challenge models with novel concepts—new words or made-up problems—and they still produce meaningful, even creative, answers?

Fabio: Good point. Models work with tokens, not whole words. These tokens often resemble parts of words with partial meaning—like mini building blocks. So when we invent a word like opportunisk (opportunity + risk), the model can piece together a plausible interpretation by combining concepts it has memorized. But it’s still recombination, not the creation of new knowledge. The core operation remains: predicting the next token based on context.

Dario: So where’s the line between memorization and generalization?

Fabio: It’s blurry. In traditional machine learning, we distinguish between inductive learning (learn from examples) and deductive learning (apply known rules). LLMs mix these in strange ways. They consume massive amounts of data inductively but appear to reason deductively. But no human learns like that—we don’t need to memorize the internet to solve problems. That’s one reason I think LLMs are doing overfitting on the world.

Dario: Is there a way to test if a model is just memorizing?

Fabio: Yes. We’ve experimented with controlled datasets and “data contamination” tests. For example, we built a SQL translation task with a dataset not available online and found that model performance dropped significantly. That supports the idea that prior exposure (memorization) plays a huge role.

Dario: Do you think these models have structural limits?

Fabio: Absolutely. A key problem is variable binding—the ability to consistently apply logic across references to the same variable. Neural networks aren’t good at this, and it hasn’t been solved since the 90s. Current architectures like transformers use attention, which approximates variable binding, but it’s not the same. Neuro-symbolic AI tries to combine symbolic logic with neural networks, but so far, it’s more deep learning in disguise than a true synthesis.

Dario: Can we reach real “understanding” with different architectures?

Fabio: Maybe. I’m not ruling it out. But the current methods aren’t enough. We need systems that learn with fewer data and combine memory, learning, and inference more like humans do.

Dario: What about the creative power of AI? Should we call these systems creative if they produce stories, art, or music that move us?

Fabio: We should ask: What is creativity? There’s repetitive creativity—recombining known elements—and disruptive creativity, which changes paradigms. LLMs are very good at the former, not the latter. Think of how photography transformed visual art. True creativity is what came after photography—impressionism, abstraction—not what came before.

Dario: If language is our most powerful tool, as Harari argues, what happens when AI can manipulate it?

Fabio: Language shaped human civilization. Now we have machines that can wield language, powered by collective human knowledge. Societies that give individuals access to this “collective intelligence” will likely move faster. But there’s risk. These tools must remain accessible, transparent, and democratic. If they’re controlled by a few, or used irresponsibly, they can mislead, divide, and manipulate.

Dario: Final thought—are you hopeful?

Fabio: Cautiously. We need to reimagine education. People must learn not just facts, but how to analyze, question, and discern. That’s our best defense against misinformation and manipulation. Let’s use this moment to build something better—not just technically, but socially.

Dario: Fabio, thank you so much for sharing your thoughts with us today.

Fabio: Thank you, Dario.

- / 5
Grazie per aver votato!