The number of RGB image inputs

#3 · closed · 1 comments

View on GitHub ↗

haofuly

Hi @DorayakiLin , thanks for your open-minded work! I'm not sure which camera you used to observe the images. In my opinion, one of the powerful ability of VGGT is processing more than two images together. So integrating VGGT helps VLM understand the whole scenerio. But the figure 7 in paper just show two images from wrist and scene camera. Thanks for your response!

Comments

DorayakiLin

Sorry for the late reply. We only use the wrist and scene cameras in our real-world experiments, which is a common setup. We also found that two images are sufficient to achieve strong geometric effects.