Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions
TLDR
Deep Research agents are vulnerable to misleading knowledge, causing false conclusions despite verification, as shown by the MisKnow-Agent framework.
Reasoning
The paper introduces a systematic framework to study a critical failure mode, with controlled experiments and defense evaluations. However, it relies on a synthetic benchmark rather than real-world deployment, and the scope is limited to a single type of misinformation.
Read-first score
Read-first score 59.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 61.
Field roles
Rank sensitivity
Stability: volatile; rank range: 35.
Keyword Scores
Deep Analysis
Innovations
- MisKnow-Agent: a framework for constructing and validating misleading knowledge with controllable authority levels and styles, yielding 5,933 quality-controlled instances on DeepResearch Benchmark tasks.
- Systematic demonstration that even limited exposure to misleading knowledge can induce false-conclusion adoption in final reports of Deep Research agents.
- Evaluation of pre- and post-research defenses showing they mitigate but do not fully prevent false conclusions, revealing the need for evidence verification and correction at both model and framework levels.
Methodology
The paper constructs 5,933 misleading instances using the MisKnow-Agent framework, varying authority and style, built on DeepResearch Benchmark tasks. Extensive experiments on open-source and closed-source Deep Research agents measure false-conclusion adoption in final reports. Search-enabled verifier models and combinations of pre- and post-research defenses are also evaluated.
Key Results
Limited exposure to misleading knowledge induces false-conclusion adoption in final reports; search-enabled verifiers identify misleading instances during focused corpus validation but the same instances are still adopted during long-horizon research, revealing a disconnect; all defense configurations mitigate but do not fully prevent false-conclusion adoption.
Limitations
- Pre- and post-research defenses, individually or combined, mitigate but do not fully prevent false-conclusion adoption.
- There is a disconnect between focused verification capabilities and workflow-level evidence use, as verifier models can identify misleading instances yet fail to stop their adoption during long-horizon research.