haofuly
Hi @DorayakiLin , thanks for your open-minded work! I'm not sure which camera you used to observe the images. In my opinion, one of the powerful ability of VGGT is processing more than two images together. So integrating VGGT helps VLM understand the whole scenerio. But the figure 7 in paper just show two images from wrist and scene camera. Thanks for your response!