← 返回资源分享
Hierarchical Question-Image Co-Attention for Visual Question Answering
论文
论文
发布时间2016-05-31
发表NeurIPS 2016 12 · arXiv:1606.00061
作者:Dhruv Batra,Devi Parikh,Jiasen Lu,Jianwei Yang
详细介绍
A number of recent works have proposed attention models for Visual Question
Answering (VQA) that generate spatial maps highlighting image regions relevant
to answering the question. In this paper, we argue that in addition to modeling
"where to look" or visual attention, it is equally important to model "what
words to listen to" or question attention. We present a novel co-attention
model for VQA that jointly reasons about image and question attention. In
addition, our model reasons about the question (and consequently the image via
the co-attention mechanism) in a hierarchical fashion via a novel 1-dimensional
convolution neural networks (CNN). Our model improves the state-of-the-art on
the VQA dataset from 60.3% to 60.5%, and from 61.6% to 63.3% on the COCO-QA
dataset. By using ResNet, the performance is further improved to 62.1% for VQA
and 65.4% for COCO-QA.
代码仓库 (9)
jiasenlu/HieCoAttenVQA官方PyTorch
karunraju/VQAPyTorch
ritvikshrivastava/ADL_VQA_Tensorflow2TensorFlow
miohana/vqaTensorFlow
WillSuen/VQATensorFlow
phisad/keras-hicoattTensorFlow
SkyOL5/VQA-CoAttentionPyTorch
arya46/VQA_HieCoAttTensorFlow
Rabahjamal/Visual-Question-AnsweringTensorFlow
