Esiste un malinteso pittoresco, e abbastanza diffuso, secondo cui i modelli linguistici “capiscono” quello che diciamo loro. Da questo equivoco discende un secondo malinteso, ancora più tenero: che basti scrivere qualcosa — qualsiasi cosa, purché sufficientemente lunga — perché il modello, con la sua aria da oracolo digitale, indovini il resto.

È una fede commovente. Ed è anche il motivo per cui milioni di prompt, ogni giorno, producono risposte che sembrano generate da un cugino svogliato in un giorno di pioggia.

Il malinteso fondativo

Quando parliamo con un essere umano, l’ambiguità è un lusso che ci possiamo permettere. Diciamo “fammi sapere” e l’altro capisce, dal contesto, dal tono, dalla relazione, da duemila piccole cose non dette, cosa intendiamo davvero. La comunicazione umana è un sistema ad altissima ridondanza che funziona nonostante la nostra pigrizia, non grazie alla nostra precisione.

Con un LLM, questa ridondanza non c’è. Non c’è l’occhiataccia che ti corregge la rotta, non c’è la cena di ieri, non c’è il fatto che vi conoscete da dieci anni. C’è solo il testo. E il testo, se è ambiguo, resta ambiguo. Solo che il modello, gentile com’è, non lo dice. Tira a indovinare. E quando un sistema probabilistico tira a indovinare, le sue probabilità si distribuiscono — democraticamente, equamente — su tutte le interpretazioni plausibili. Inclusa quella sbagliata.

L’entropia, che non perdona

C’è un fenomeno tecnico dietro a tutto questo, e ha un nome quasi poetico: i modelli linguistici, davanti a un input ambiguo, allargano lo spazio delle risposte possibili. È fisica statistica, non malafede. Se la richiesta è cristallina, il modello converge. Se è nebulosa, il modello esplora. E “esplora” è una parola elegante per dire che produce risposte mediocri con maggior frequenza, perché le risposte mediocri sono, statisticamente, le più numerose.

In altre parole: l’ambiguità non degrada le performance perché il modello sia stupido. Le degrada perché il modello sta facendo esattamente il suo lavoro — e il suo lavoro, di fronte all’incertezza, è restituirti una media. La media di un insieme grande è quasi sempre poco interessante. È così che si finisce con quei testi piatti, quei riassunti vagamente corretti, quelle email che sembrano scritte da un consulente in burnout.

La cortesia che inganna

Il vero problema, però, non è tecnico. È psicologico. Gli LLM sono stati addestrati a essere cortesi, accomodanti, eternamente disponibili. Non ti chiedono “scusa, ma cosa intendi esattamente?” come farebbe un collega irritato. Producono. Sempre. Anche quando non dovrebbero.

Questa cortesia, paradossalmente, ci fa credere di essere stati chiari. Riceviamo una risposta plausibile, la leggiamo, annuiamo, e proseguiamo. Non sapremo mai quanto migliore sarebbe stata la risposta a un prompt fatto bene, perché il modello non ce lo dice. È come avere un dipendente che annuisce a tutto: non capisci di averlo messo in difficoltà finché non vedi il risultato finale, e a quel punto è troppo tardi per chiedere chiarimenti.

Il piccolo zoo degli errori quotidiani

Se l’ambiguità è la tassa, ci sono modi specifici di pagarla — vizi conversazionali talmente diffusi da essere diventati invisibili. Vale la pena guardarli da vicino, perché chiunque usi un LLM ne commette almeno uno al giorno, spesso senza accorgersene.

La domanda che contiene già la risposta. È il classico “non è vero che…?”, oppure “non pensi anche tu che…?”. Sembrano domande, ma sono dichiarazioni travestite. Stiamo chiedendo conferma, non informazione. E il modello, addestrato a essere utile e a non contraddire gratuitamente l’interlocutore, tende a darcela. Non perché menta, ma perché il segnale che gli abbiamo inviato — il modo in cui abbiamo formulato la frase, le parole che abbiamo scelto, la posizione delle virgole — gli dice che la nostra ipotesi è probabilmente corretta. Il modello completa il pattern che gli abbiamo suggerito. Se vuoi un’analisi vera, devi togliere il pollice dalla bilancia. “Cosa pensi di X?” produce risposte molto diverse da “X è una buona idea, vero?”.

La domanda che cerca un fan, non un interlocutore. È la variante più imbarazzante della precedente. Una quantità sorprendente di persone usa gli LLM come specchio adulatorio: incolla un proprio testo, una propria idea, un proprio piano, e chiede “cosa ne pensi?” sperando, neanche troppo segretamente, in un applauso. Il modello, ovviamente, applaude. È il suo mestiere essere d’aiuto, e per molti utenti “essere d’aiuto” significa “farmi sentire bene”. È una forma di solitudine travestita da produttività. Se cerchi conferme, le otterrai. Se cerchi feedback, devi chiederlo esplicitamente — e devi avere il coraggio di chiedere quello brutto, non quello carino.

La domanda che lascia troppo spazio. All’estremo opposto c’è il prompt-blob: “parlami di marketing”, “aiutami con il mio business”, “scrivi qualcosa sulla sostenibilità”. Sono richieste che non orientano nulla, e infatti il modello produce esattamente quello che ti aspetti: il centro statistico dell’argomento. La cosa più ovvia, più già detta, più rassicurante. Una richiesta vaga è un permesso a essere banali, e i modelli accettano il permesso con entusiasmo.

La domanda che mente sul contesto. Capita di chiedere consigli omettendo i dettagli che renderebbero il consiglio utile — perché sono noiosi da scrivere, o perché li diamo per scontati, o perché un po’ ci vergogniamo di averli. Il modello allora risponde sul caso medio, e noi ci lamentiamo che la risposta non si adatta al nostro caso specifico. Ma non gli avevamo detto il caso specifico. Avevamo solo sperato che lo intuisse.

La domanda che cambia mentre la scrivi. È quella in cui inizi chiedendo una cosa, a metà ne aggiungi un’altra, e alla fine concludi con una terza, scollegata dalle prime due. Il modello tenta di accontentarle tutte e tre e produce una risposta che non accontenta nessuna. Spesso è il sintomo che non sapevamo davvero cosa volevamo. E in quel caso, lo strumento giusto non è un LLM: è un foglio bianco e cinque minuti di silenzio.

Il filo che lega tutti questi errori è uno solo: in ogni caso, stiamo chiedendo al modello di compensare qualcosa che dovremmo fare noi. Pensare prima. Decidere cosa vogliamo davvero. Distinguere tra il bisogno di sapere e il bisogno di essere rassicurati. Sono operazioni umane. Nessun modello, per quanto potente, le farà al posto nostro.

La responsabilità è nostra

Qui sta la parte scomoda — quella che molti preferiscono non vedere. Se il prompt è ambiguo e la risposta è scadente, la colpa non è del modello. È nostra. Abbiamo affidato un compito complesso a un sistema che non legge nel pensiero, sperando che leggesse nel pensiero, e ci siamo lamentati del fatto che non leggesse nel pensiero.

C’è qualcosa di profondamente umano in questo: trattiamo gli LLM come trattiamo i colleghi, i partner, i genitori. Diciamo metà delle cose e ci aspettiamo che l’altro capisca l’altra metà. Con le persone, qualche volta, funziona — perché c’è una storia condivisa che riempie i buchi. Con un modello che ti incontra per la prima volta a ogni messaggio, no.

Una piccola igiene mentale

Non sto suggerendo di scrivere prompt da ingegneri NASA. Sto suggerendo qualcosa di più sobrio: rileggere quello che si è scritto e chiedersi, onestamente, tre cose. Primo: un estraneo capirebbe? Secondo: sto chiedendo davvero, o sto cercando un applauso? Terzo: ho dato al modello le informazioni di cui ha bisogno, o ho sperato che le indovinasse?

È un esercizio noioso. È anche, probabilmente, il singolo investimento con il rendimento più alto che si possa fare in qualunque attività che coinvolga un LLM. Più del modello che usi, più della temperature, più dei trucchetti del prompt engineering del momento.

Morale

I modelli linguistici sono specchi statistici molto sofisticati. Se gli mostri una domanda chiara, ti restituiscono una risposta chiara. Se gli mostri un pasticcio, ti restituiscono un pasticcio elegante. Se gli mostri il tuo bisogno di approvazione, ti restituiscono approvazione.

L’ambiguità, in fondo, è una forma di pigrizia camuffata da spontaneità. Funzionava con gli umani perché gli umani la perdonavano. Le macchine no: la pagano, e te la fanno pagare. Con gli interessi.

E forse — forse — il vero motivo per cui certe conversazioni con un LLM ci sembrano deludenti non è che il modello non capisca noi. È che noi, parlando con lui, abbiamo capito qualcosa di scomodo su come parliamo davvero.

There is a picturesque and fairly widespread misunderstanding that language models “understand” what we tell them. From this mistake comes a second, even more tender one: that it is enough to write something — anything, as long as it is sufficiently long — for the model, with its air of digital oracle, to infer the rest.

It is a moving faith. It is also why millions of prompts, every day, produce answers that sound as if they were generated by an unmotivated cousin on a rainy afternoon.

The Foundational Misunderstanding

When we speak to another human being, ambiguity is a luxury we can afford. We say “let me know” and the other person understands, from context, tone, relationship, and two thousand small unsaid things, what we actually mean. Human communication is a very high-redundancy system that works despite our laziness, not thanks to our precision.

With an LLM, that redundancy is not there. There is no dirty look to correct your course, no dinner from yesterday, no fact that you have known each other for ten years. There is only the text. And text, when it is ambiguous, stays ambiguous. Except the model, polite as it is, does not say so. It guesses. And when a probabilistic system guesses, its probabilities spread — democratically, fairly — across every plausible interpretation. Including the wrong one.

Entropy Does Not Forgive

There is a technical phenomenon behind all this, and it has an almost poetic name: when language models face ambiguous input, they widen the space of possible answers. It is statistical physics, not bad faith. If the request is crystal clear, the model converges. If it is foggy, the model explores. And “explores” is an elegant way of saying that it produces mediocre answers more often, because mediocre answers are, statistically, the most numerous.

In other words: ambiguity does not degrade performance because the model is stupid. It degrades performance because the model is doing exactly its job — and its job, when faced with uncertainty, is to give you an average. The average of a large set is almost always uninteresting. That is how you end up with those flat texts, those vaguely correct summaries, those emails that sound like they were written by a consultant in burnout.

The Courtesy That Misleads

The real problem, though, is not technical. It is psychological. LLMs have been trained to be polite, accommodating, eternally available. They do not ask you “sorry, but what exactly do you mean?” the way an irritated colleague would. They produce. Always. Even when they should not.

This courtesy, paradoxically, makes us believe we were clear. We receive a plausible answer, read it, nod, and move on. We will never know how much better the answer could have been with a well-made prompt, because the model does not tell us. It is like having an employee who nods at everything: you do not realize you have put them in trouble until you see the final result, and by then it is too late to ask clarifying questions.

The Little Zoo of Everyday Mistakes

If ambiguity is the tax, there are specific ways to pay it — conversational vices so common they have become invisible. They are worth looking at closely, because anyone who uses an LLM commits at least one of them a day, often without noticing.

The question that already contains the answer. The classic “isn’t it true that…?”, or “don’t you also think that…?”. They look like questions, but they are statements in disguise. We are asking for confirmation, not information. And the model, trained to be helpful and not contradict the interlocutor for free, tends to give it to us. Not because it is lying, but because the signal we sent — the way we phrased the sentence, the words we chose, the position of the commas — tells it our hypothesis is probably correct. The model completes the pattern we suggested. If you want a real analysis, you need to take your thumb off the scale. “What do you think of X?” produces very different answers from “X is a good idea, right?”.

The question looking for a fan, not an interlocutor. This is the more embarrassing variant of the previous one. A surprising number of people use LLMs as flattering mirrors: they paste in their own text, idea, or plan and ask “what do you think?”, hoping, not even that secretly, for applause. The model, of course, applauds. Its job is to be helpful, and for many users “being helpful” means “making me feel good”. It is a form of loneliness disguised as productivity. If you want reassurance, you will get it. If you want feedback, you have to ask for it explicitly — and you have to have the courage to ask for the ugly kind, not the nice kind.

The question that leaves too much room. At the opposite extreme is the prompt blob: “tell me about marketing”, “help me with my business”, “write something about sustainability”. These are requests that orient nothing, and so the model produces exactly what you should expect: the statistical center of the topic. The most obvious thing, the most already-said thing, the most reassuring thing. A vague request is permission to be banal, and models accept that permission enthusiastically.

The question that lies about context. It happens when we ask for advice while leaving out the details that would make the advice useful — because they are boring to write, or because we take them for granted, or because we are slightly ashamed of them. The model then answers for the average case, and we complain that the answer does not fit our specific case. But we had not told it the specific case. We had only hoped it would intuit it.

The question that changes while you write it. This is the one where you start by asking one thing, add another halfway through, and end with a third, disconnected from the first two. The model tries to satisfy all three and produces an answer that satisfies none. Often it is a symptom that we did not really know what we wanted. And in that case, the right tool is not an LLM: it is a blank sheet of paper and five minutes of silence.

The thread tying all these mistakes together is only one: in every case, we are asking the model to compensate for something we should do ourselves. Think first. Decide what we really want. Distinguish the need to know from the need to be reassured. These are human operations. No model, however powerful, will do them for us.

The Responsibility Is Ours

Here lies the uncomfortable part — the one many people prefer not to see. If the prompt is ambiguous and the answer is poor, it is not the model’s fault. It is ours. We assigned a complex task to a system that does not read minds, hoping it would read minds, and then complained that it did not read minds.

There is something deeply human about this: we treat LLMs the way we treat colleagues, partners, parents. We say half of things and expect the other person to understand the other half. With people, sometimes, it works — because there is a shared history that fills the gaps. With a model that meets you for the first time at every message, it does not.

A Little Mental Hygiene

I am not suggesting we write prompts like NASA engineers. I am suggesting something more sober: reread what you have written and ask yourself, honestly, three things. First: would a stranger understand it? Second: am I really asking, or am I looking for applause? Third: did I give the model the information it needs, or did I hope it would guess?

It is a boring exercise. It is also, probably, the single highest-return investment you can make in any activity involving an LLM. More than the model you use, more than temperature, more than whatever prompt-engineering trick is fashionable this week.

Moral

Language models are very sophisticated statistical mirrors. If you show them a clear question, they give you a clear answer back. If you show them a mess, they give you an elegant mess back. If you show them your need for approval, they give you approval back.

Ambiguity, in the end, is a form of laziness disguised as spontaneity. It worked with humans because humans forgave it. Machines do not: they pay for it, and they make you pay too. With interest.

And perhaps — perhaps — the real reason some conversations with an LLM feel disappointing is not that the model does not understand us. It is that, while talking to it, we have understood something uncomfortable about how we really talk.