Remote sensing image captioning via Variational Autoencoder and Reinforcement Learning

Image captioning, i.e., generating the natural semantic descriptions of given image, is an essential task for machines to understand the content of the image. Remote sensing image captioning is a part of the field. Most of the current remote sensing image captioning models suffered the overfitting p...

Celý popis

Uložené v:

Podrobná bibliografia
Vydané v:	Knowledge-based systems Ročník 203; s. 105920
Hlavní autori:	Shen, Xiangqing, Liu, Bing, Zhou, Yong, Zhao, Jiaqi, Liu, Mingming
Médium:	Journal Article
Jazyk:	English
Vydavateľské údaje:	Amsterdam Elsevier B.V 05.09.2020 Elsevier Science Ltd
Predmet:	Convolutional neural network Learning Machinery Reinforcement Reinforcement learning Remote sensing Remote sensing image captioning Self-attention mechanisms Semantic features Semantics Transfer learning Transformer Variational Autoencoder Transformer Remote sensing image captioning Transfer learning Variational Autoencoder Convolutional neural network Reinforcement learning Self-attention mechanisms
ISSN:	0950-7051, 1872-7409
On-line prístup:	Získať plný text
Tagy:	Pridať tag Žiadne tagy, Buďte prvý, kto otaguje tento záznam!

Popis
Shrnutí:	Image captioning, i.e., generating the natural semantic descriptions of given image, is an essential task for machines to understand the content of the image. Remote sensing image captioning is a part of the field. Most of the current remote sensing image captioning models suffered the overfitting problem and failed to utilize the semantic information in images. To this end, we propose a Variational Autoencoder and Reinforcement Learning based Two-stage Multi-task Learning Model (VRTMM) for the remote sensing image captioning task. In the first stage, we finetune the CNN jointly with the Variational Autoencoder. In the second stage, the Transformer generates the text description using both spatial and semantic features. Reinforcement Learning is then applied to enhance the quality of the generated sentences. Our model surpasses the previous state of the art records by a large margin on all seven scores on Remote Sensing Image Caption Dataset. The experiment result indicates our model is effective on remote sensing image captioning and achieves the new state-of-the-art result. •Introducing VAE to regularize the shared encoder and extract image features more effectively by reconstructing input images.•Improving the performance of image caption significantly by virtue of low-level and high-level image features simultaneously.•Enhancing the final text description quality by adding self-attention to spatial features.•Our proposed model outperforms the state-of-the-art models in the remote sensing image captioning.
Bibliografia:	ObjectType-Article-1 SourceType-Scholarly Journals-1 ObjectType-Feature-2 content type line 14
ISSN:	0950-7051 1872-7409
DOI:	10.1016/j.knosys.2020.105920