naiveHobo/CaptionNet---TensorFlow

An encoder-decoder based deep neural network for image captioning

★ 6Forks 0PythonGitHub ↗Compare

README

HoboCaptionNet

An encoder-decoder based deep neural network for image captioning

Encoder

The encoder uses the Inception-v4 architecture to extract features from the image. The final softmax layer is removed and a fully-connected layer is added in its place which converts the extracted features to the same size as the word-embeddings and feeds it to the decoder network. The pre-trained inception-v4 weights are available here. The weights must be places in the 'inception' directory.

Decoder

The decoder network is a basic lstm network that takes in the output of the encoder and predits the most-probable word at each time-step with a maximum length of 20 words for any predicted caption.

Pre-trained model

HoboCaptionNet was trained on a combination of COCO-2017 and flickr30k datasets using tensorflow-1.4. The pre-trained model is available here. The file must be placed in the 'model/Trained_Graphs' directory.

Testing

To test the pre-trained model, run the test.py script.

python test.py --h

Training and evaluation instructions will be updated soon.

Contributors

naiveHobo

Issues