Please use this identifier to cite or link to this item:
https://elib.vku.udn.vn/handle/123456789/7595Full metadata record
| DC Field | Value | Language |
|---|---|---|
| dc.contributor.author | Trinh, The Minh | - |
| dc.contributor.author | Van, Duc Cuong | - |
| dc.contributor.author | Le, Hong Ha Anh | - |
| dc.contributor.author | Nguyen, Van Son | - |
| dc.contributor.author | Huynh, Cong Phap | - |
| dc.contributor.author | Nguyen, Thi Hanh | - |
| dc.date.accessioned | 2026-08-05T08:00:33Z | - |
| dc.date.available | 2026-08-05T08:00:33Z | - |
| dc.date.issued | 2026-06 | - |
| dc.identifier.isbn | 979-8-3315-9372-8 | - |
| dc.identifier.isbn | 979-8-3315-9373-5 | - |
| dc.identifier.uri | https://doi.org/10.1109/DEFI67526.2025.11551630 | - |
| dc.identifier.uri | https://elib.vku.udn.vn/handle/123456789/7595 | - |
| dc.description | 2025 Conference on Digital Economy and Fintech Innovation (DEFI): pp: 81-88. | vi_VN |
| dc.description.abstract | Product manuals are critical in helping users understand and operate consumer products. Yet, their increasing length and reliance on complex graphical elements often frustrate information retrieval. Traditional Product Manual Question Answering (PMQA) systems primarily focus on text and overlook visual content such as icons and diagrams, resulting in incomplete or ambiguous answers. This gap highlights the need for a multimodal approach that integrates textual and visual cues to ensure accurate and contextually grounded support. In this paper, we introduce MERIT, a novel multimodal QA framework that advances beyond page-level and text-only baselines. MERIT segments manuals into fine-grained, semantically coherent instruction chunks and encodes them using a Mixture-of-Experts Q-Former encoder, which effectively fuses textual summaries with graphical features. Trained under a contrastive learning objective, MERIT constructs a unified embedding space optimized for dense retrieval. A vision-language model then processes chunks and their associated graphics to generate accurate, fluent, and visually grounded answers. Extensive experiments on the PM209 dataset demonstrate that MERIT consistently improves retrieval precision and answer quality over state-of-the-art baselines, underscoring the importance of fine-grained multimodal reasoning for practical e-commerce applications. | vi_VN |
| dc.language.iso | en | vi_VN |
| dc.publisher | IEEE | vi_VN |
| dc.subject | E-commerce platform | vi_VN |
| dc.subject | Product Manual Question Answering | vi_VN |
| dc.subject | Retrieval | vi_VN |
| dc.subject | Mixture of Experts | vi_VN |
| dc.title | MERIT: Multimodal Enhanced Retrieval through Instruction-level Transformer-based Chunking in Product Manuals | vi_VN |
| dc.type | Working Paper | vi_VN |
| Appears in Collections: | DEFI 2025 | |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.