Please use this identifier to cite or link to this item: https://elib.vku.udn.vn/handle/123456789/7595
Full metadata record
DC FieldValueLanguage
dc.contributor.authorTrinh, The Minh-
dc.contributor.authorVan, Duc Cuong-
dc.contributor.authorLe, Hong Ha Anh-
dc.contributor.authorNguyen, Van Son-
dc.contributor.authorHuynh, Cong Phap-
dc.contributor.authorNguyen, Thi Hanh-
dc.date.accessioned2026-08-05T08:00:33Z-
dc.date.available2026-08-05T08:00:33Z-
dc.date.issued2026-06-
dc.identifier.isbn979-8-3315-9372-8-
dc.identifier.isbn979-8-3315-9373-5-
dc.identifier.urihttps://doi.org/10.1109/DEFI67526.2025.11551630-
dc.identifier.urihttps://elib.vku.udn.vn/handle/123456789/7595-
dc.description2025 Conference on Digital Economy and Fintech Innovation (DEFI): pp: 81-88.vi_VN
dc.description.abstractProduct manuals are critical in helping users understand and operate consumer products. Yet, their increasing length and reliance on complex graphical elements often frustrate information retrieval. Traditional Product Manual Question Answering (PMQA) systems primarily focus on text and overlook visual content such as icons and diagrams, resulting in incomplete or ambiguous answers. This gap highlights the need for a multimodal approach that integrates textual and visual cues to ensure accurate and contextually grounded support. In this paper, we introduce MERIT, a novel multimodal QA framework that advances beyond page-level and text-only baselines. MERIT segments manuals into fine-grained, semantically coherent instruction chunks and encodes them using a Mixture-of-Experts Q-Former encoder, which effectively fuses textual summaries with graphical features. Trained under a contrastive learning objective, MERIT constructs a unified embedding space optimized for dense retrieval. A vision-language model then processes chunks and their associated graphics to generate accurate, fluent, and visually grounded answers. Extensive experiments on the PM209 dataset demonstrate that MERIT consistently improves retrieval precision and answer quality over state-of-the-art baselines, underscoring the importance of fine-grained multimodal reasoning for practical e-commerce applications.vi_VN
dc.language.isoenvi_VN
dc.publisherIEEEvi_VN
dc.subjectE-commerce platformvi_VN
dc.subjectProduct Manual Question Answeringvi_VN
dc.subjectRetrievalvi_VN
dc.subjectMixture of Expertsvi_VN
dc.titleMERIT: Multimodal Enhanced Retrieval through Instruction-level Transformer-based Chunking in Product Manualsvi_VN
dc.typeWorking Papervi_VN
Appears in Collections:DEFI 2025

Files in This Item:

 Sign in to read



Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.