Pre-Training Clustering Models to Summarize Vietnamese Texts.

Saved in:
Bibliographic Details
Title: Pre-Training Clustering Models to Summarize Vietnamese Texts.
Authors: Nguyen, Ti-Hon, Do, Thanh-Nghi
Source: Vietnam Journal of Computer Science (World Scientific); Feb2025, Vol. 12 Issue 1, p83-100, 18p
Subject Terms: LANGUAGE models, AUTOMATIC summarization, TEXT summarization, DOCUMENT clustering, ELECTRONIC newspapers, K-means clustering, ALGORITHMS
Abstract: Our investigation aims at pre-training clustering models to summarize Vietnamese texts. For this purpose, we create a large-scale dataset by collecting Vietnamese articles from newspaper websites and extracting the plain text to build the dataset, including 1,101,101 documents. We propose a new single-document extractive text summarization model based on clustering models. Our proposal clusters the documents with the hard clustering k-means algorithm and the soft clustering LDA (Latent Dirichlet Allocation) algorithm. Then, based on the pre-training clustering models, a summary model is used to select the salient sentence in the input text to construct the summary. The empirical results showed that our summary model achieved 51.22% ROUGE-1, 17.62% ROUGE-2 and 29.16% ROUGE-L on the testing set. Besides the traditional word representation such as BoW (Bag-of-Words), we also use the word meaning-based tools like FastText and BERT (Bidirectional Encoder Representations from Transformers) in our model. The additional benefit of our proposed extractive summary model is that the output summary is a long-text, readable document. Furthermore, the model's architecture is straightforward, easy to understand and runs on cost-efficient resources like arm CPU and GPU too. [ABSTRACT FROM AUTHOR]
Copyright of Vietnam Journal of Computer Science (World Scientific) is the property of World Scientific Publishing Company and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Complementary Index
Be the first to leave a comment!
You must be logged in first