Hello, I was wondering whether any research was done on how well the smaller V-JEPA variants perform (e.g. ViT-Large) in terms of e.g. alignment with human evaluation. Does the ViT-Huge variant of V-JEPA have sufficiently strong performance relative to the smaller checkpoints to justify the increased computational cost?
Hello, I was wondering whether any research was done on how well the smaller V-JEPA variants perform (e.g. ViT-Large) in terms of e.g. alignment with human evaluation. Does the ViT-Huge variant of V-JEPA have sufficiently strong performance relative to the smaller checkpoints to justify the increased computational cost?