Mobile QR Code QR CODE : Journal of the Korean Society of Civil Engineers
Title Retrieval-Augmented Question Answering with Post-Generation Evidence Verification for Korean Design Standards: System Design and Empirical Validation
Authors 홍철승(Hong, Chulseung) ; 이수연(Lee, Suyeon) ; 문창(Moon, Chang) ; 은영기(Eun, Youngkee)
DOI https://doi.org/10.12652/Ksce.2026.46.5.0485
Page pp.485-499
ISSN 10156348
Keywords 건설공사 설계기준(KDS); 검색증강생성(RAG); 대규모 언어모델(LLM); 질의응답; 근거 검증 Korean Design Standards (KDS); Retrieval-augmented generation (RAG); Large language models (LLMs); Question answering; Evidence verification
Abstract The Korean Design Standards (KDS) specify technical requirements for the design of various types of infrastructure and construction works. In engineering practice, engineers must not only locate relevant clauses but also identify the values or provisions applicable to the conditions at hand and verify their supporting basis. However, this review still relies largely on manual examination, and document search alone makes it difficult to select provisions applicable to conditions expressed in natural language and to determine whether a generated answer is supported by the source text. This study proposes EC-RAG (Engineering Code RAG), a system that interprets the conditions stated in a natural-language query, retrieves relevant evidence from the full KDS corpus, presents a condition-specific answer with a supporting quotation, and checks whether the answer value is supported by that quotation. Clause-path information and textual representations of tables and formulae were extracted from the KDS source documents to construct retrieval units, and keyword-based retrieval was combined with semantic similarity retrieval. The system was also designed not to present an answer value as definitive when it could not be confirmed from the retrieved evidence. Approximately 50,000 retrieval units were constructed from 568 KDS documents collected at the time of data acquisition, and the system was evaluated using 1,019 questions spanning 18 major KDS categories. The evaluation queries did not include a specific KDS code or clause path, and the system searched the full index for each query. The final answer accuracy was 97.7 % (996/1,019). Of the 23 questions not judged correct, 19 resulted from difficulties in identifying the applicable value in multi-level tables, while four resulted from the relevant evidence not being included among the retrieved candidates. These results indicate that the proposed approach can support the identification of relevant design provisions and applicable values from natural-language queries without requiring users to specify a document or clause in advance. They also identify
table-structure processing and retrieval performance as the main areas for further improvement.