Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs
arXiv SecurityArchived Aug 14, 2026✓ Full text saved
arXiv:2608.12911v1 Announce Type: cross Abstract: While the privacy risks of multimodal large language models (MLLMs) have drawn significant attention, the unique vulnerabilities of domain-specific MLLMs remain largely underexplored. Focusing on document understanding MLLMs for identity document processing, this paper investigates the privacy issues inherent in Key Information Extraction (KIE) tasks. We reveal that when input images lack sufficient visual evidence, these models often rely on mem
Full text archived locally
✦ AI Summary· Claude Sonnet
Computer Science > Computer Vision and Pattern Recognition
[Submitted on 13 Aug 2026]
Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs
Beining Xu, Hairui Wang, Jiaxin Wang, Changsheng Chen, Anirban Chakraborty
While the privacy risks of multimodal large language models (MLLMs) have drawn significant attention, the unique vulnerabilities of domain-specific MLLMs remain largely underexplored. Focusing on document understanding MLLMs for identity document processing, this paper investigates the privacy issues inherent in Key Information Extraction (KIE) tasks. We reveal that when input images lack sufficient visual evidence, these models often rely on memorized field relations from training data to infer missing content, thereby leaking multiple correlated fields containing sensitive personal information. To mitigate this risk, we make three key this http URL, we propose the Dynamic Relational Unlearning Framework (DRUF) which comprises a Relational Decoupling Unlearning (RDU) module and a dynamic set update mechanism. It suppresses the leakage of high-risk field pairs while preserving KIE this http URL, we introduce DocPrivacyBench, a novel benchmark to systematically evaluate a model's susceptibility to privacy leakage under conditions of absent or minimal visual this http URL, we evaluate three MLLMs and six unlearning methods using this benchmark, assessing both post-unlearning leakage suppression and utility this http URL results demonstrate that existing MLLMs consistently exhibit privacy leakage when visual evidence is scarce, particularly on noisier datasets. In contrast, DRUF outperforms the strongest baseline by improving leakage suppression by 4.8 percentage points, effectively mitigating privacy risks while maintaining robust document information extraction performance.
Comments: ACM mm 2026
Subjects: Computer Vision and Pattern Recognition (cs.CV); Cryptography and Security (cs.CR); Multimedia (cs.MM)
Cite as: arXiv:2608.12911 [cs.CV]
(or arXiv:2608.12911v1 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2608.12911
Focus to learn more
Related DOI:
https://doi.org/10.1145/3767308.3835901
Focus to learn more
Submission history
From: Beining Xu [view email]
[v1] Thu, 13 Aug 2026 07:53:53 UTC (2,903 KB)
Access Paper:
HTML (experimental)
view license
Current browse context:
cs.CV
< prev | next >
new | recent | 2026-08
Change to browse by:
cs
cs.CR
cs.MM
References & Citations
NASA ADS
Google Scholar
Semantic Scholar
Export BibTeX Citation
Bookmark
Bibliographic Tools
Bibliographic and Citation Tools
Bibliographic Explorer Toggle
Bibliographic Explorer (What is the Explorer?)
Connected Papers Toggle
Connected Papers (What is Connected Papers?)
Litmaps Toggle
Litmaps (What is Litmaps?)
scite.ai Toggle
scite Smart Citations (What are Smart Citations?)
Code, Data, Media
Demos
Related Papers
About arXivLabs
Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)