~/portfolio
← back to all posts
★ featured•Aug 16, 2026•4 min read

When AI Has to Choose What to Believe

How multimodal LLMs navigate conflicting sensory signals, modality bias vs. reasoning uncertainty, and the subtle challenge of source dependence in AI truth discovery.

#Multimodal AI#MLLMs#AI Research#Truth Discovery#Epistemology

Multimodal AI is supposed to be good at combining information.

Give it an image and some text. Give it a video, its audio, a caption, some OCR. In theory, the model should look at everything together and figure out what is true.

But what happens when the modalities disagree?

Say an image suggests A, while the text says B.

Which one does the model believe?

It turns out this isn't as neutral as it sounds.


Modality Preferences in MLLMs

A recent paper, Evaluating and Steering Modality Preferences in Multimodal Large Language Models, tries to study exactly this. The authors build a benchmark called MC², where different modalities are deliberately given conflicting information, and test 18 multimodal large language models.

The setup is pretty intuitive. If the image points toward one answer and the text points toward another, see which one the model follows.

And the answer isn't always "it depends."

The models show consistent modality preferences. Some tend to follow one modality more strongly than another, even when both contain conflicting information. The paper also shows that these preferences can be manipulated, suggesting this isn't simply random behavior.

That result is fascinating, but it immediately raises a harder question:

What exactly do we mean when we say a model "prefers" a modality?

Maybe it just trusts text more.

But maybe not.


Preference vs. Uncertainty

A follow-up paper, When Modalities Conflict: How Unimodal Reasoning Uncertainty Governs Preference Dynamics in MLLMs, digs deeper into this. Its argument is that what looks like modality preference can actually contain two different things:

  • An inherent preference for a modality
  • The model's relative uncertainty when reasoning from that modality

That distinction matters.

Imagine the model keeps choosing the textual answer over the visual one.

You might conclude:

"The model trusts text."

But there is another possibility:

"The model is simply better or more confident at solving this particular problem from text."

Those sound similar, but they're very different phenomena.

And the newer generation of multimodal models makes this even stranger. Recent work on native omni-modal models has found cases where the preference shifts toward vision, rather than the text preference observed in earlier systems.

So perhaps there isn't a single universal hierarchy like:

$$\text{text} > \text{image} > \text{audio}$$

Maybe modality preference depends on the model, the task, the architecture, and the uncertainty of the evidence.


The Four-Signal Paradox

Here is the question I find much more interesting.

Suppose the information isn't coming from independent sources.

Imagine one person making a short video:

  • They say A.
  • They put B as text on the screen.
  • Their caption says C.
  • The visual itself suggests D.

Now we have four modalities, but not necessarily four pieces of evidence. They all came from the same person.

This creates a weird epistemic problem.

If audio, caption, and OCR all say the same thing, should an AI think:

"Three signals agree. Confidence is high."

Or should it think:

"One person said the same thing three times."

Those aren't the same thing.


Truth Discovery & Source Dependence

This starts connecting to a much older problem in computer science: truth discovery.

Classical truth-discovery systems often deal with multiple sources making conflicting claims. A central idea is that sources have different reliability, and agreement between multiple sources can help estimate which claim is true.

But there is an assumption hiding underneath all of this:

How independent are the sources?

  • If Reuters, a government document, and an academic paper all independently report $X$, their agreement is useful evidence.
  • If one person says $X$ in their video, repeats $X$ in the caption, and writes $X$ on screen, the apparent agreement is far less informative.

The number of signals has gone up. The number of independent sources has not.


A Hidden Challenge in Plain Sight

And that makes me wonder whether multimodal systems have a slightly strange problem hiding in plain sight.

Maybe the important question isn't simply:

Which modality does an AI prefer?

Maybe it's:

How should an AI count evidence when several modalities are really just different expressions of the same source?

We already know models exhibit modality preferences. We know those preferences can be confounded with uncertainty.

What I'm less sure about is what happens when modality preference and source dependence interact.

And that seems like a much more interesting question to investigate.

Because eventually an AI doesn't just have to answer a question.

It has to decide what deserves to be remembered.

And when four signals disagree, that decision suddenly becomes much harder.