ZhihuaGao/TinyLLaVA_Factory
A Framework of Small-scale Large Multimodal Models
A Framework of Small-scale Large Multimodal Models
A minimal vision-language model prototype that aligns image and text embeddings via a projection layer (inspired by LLaVA)
link
Detectron2 is FAIR's next-generation platform for object detection and segmentation.
Google's MobileNets definition with Alpha = 0.25, 0.5, 0.75, 1.0, in Caffe prototxt.
Antialiasing cnns to improve stability and accuracy. In ICML 2019.
Open MMLab Detection Toolbox and Benchmark
ImageNet pre-trained models with batch normalization for the Caffe framework
some notes for reading papers