Download
← Morning Tech & AI
Ricerca2 min read

DEEPO Tackles Hallucination in Multimodal LLMs with Dual-Entropy RL

Researchers propose DEEPO, a dual‑entropy reinforcement learning approach designed to curb hallucinations in multimodal LLMs.

September 25, 2026·Source: arXiv cs.AI·AI-assisted
DEEPO Tackles Hallucination in Multimodal LLMs with Dual-Entropy RL

## New Approach to Hallucination Control A recent arXiv pre‑print (arXiv:2609.28570v1) presents Dual‑Entropy Enhanced Policy Optimization (DEEPO), a reinforcement‑learning framework aimed at reducing hallucinations in multimodal large language models (MLLMs).

## Identified Weak Points The authors pinpoint two vulnerabilities in the standard RL correction chain. First, *hard queries*—inputs with high semantic entropy—often generate uniformly incorrect sample groups during rollouts, nullifying the advantage signal precisely where hallucination risk peaks. Second, at the optimization stage, tokens that are confident yet wrong become *gradient‑invisible*: as a categorical policy sharpens, the expected score‑gradient norm approaches zero, delivering minimal updates to the most erroneous predictions.

## DEEPO’s Dual‑Stage Solution DEEPO addresses these issues through a two‑phase enhancement. It combines signal variance regularization to preserve informative gradients during rollout with a complementary gradient‑preserving mechanism (the abstract truncates before naming the second component). This dual strategy seeks to maintain a robust advantage signal and ensure that confident‑but‑incorrect tokens receive meaningful parameter updates.

## Implications for Future Research By targeting both rollout entropy and gradient attenuation, DEEPO offers a systematic way to improve reasoning fidelity in MLLMs. The paper suggests that integrating dual‑entropy considerations could become a standard practice for RL‑based fine‑tuning of multimodal models.

*Source: arXiv cs.AI (https://arxiv.org/abs/2609.28570), published 2026‑09‑25.*

#multimodale#RL#ottimizzazione#allucinazioni

Get Morning Tech & AI

Every morning at 7:30, the AI story of the day in your inbox.

Subscribe
Shellonback

Preferenze cookie

Scegli quali categorie di cookie accettare. I cookie tecnici e funzionali sono sempre attivi.

Per maggiori informazioni, consulta la nostra Cookie Policy e la Privacy Policy.

Cookie di profilazione

Utilizzati per creare profili relativi all'utente e inviare messaggi promozionali in linea con le preferenze espresse.

Cookie analitici

Ci permettono di capire come gli utenti navigano il sito per migliorare l'esperienza e i contenuti.

Cookie tecnici

Sempre attivo

Necessari per il funzionamento del sito. Non possono essere disattivati.

Cookie funzionali

Sempre attivo

Consentono funzionalità avanzate come la memorizzazione delle preferenze di navigazione.