High-throughput, resume-safe vision and video embedding pipelines for VLM datasets—26 pinned encoders, streaming Hugging Face ingestion, and Safetensors shard commits.
-
Updated
Aug 9, 2026 - Python
High-throughput, resume-safe vision and video embedding pipelines for VLM datasets—26 pinned encoders, streaming Hugging Face ingestion, and Safetensors shard commits.
Open-source vision stack with stereo camera hardware, GPU processing, and AI agent for training video classifiers.
vjepa / vjepa2 / vjepa2.1 PCA visualization utility for dense features and world model inspection.
Assess Data Quality Before Annotation or Labelled Data Quality after Annotation (Txt files/Yolo Format). Visualise the patterns covered by each class/activity.
Can the V-JEPA2 model be used as a world model?
Masked Multi-Component Gated Decomposition Architecture
V-JEPA for Gray-Scott dynamics. Initial work produced during the 24-hour Hack the World(s) hackathon. 1st place 🏆
A physics-based video search engine using Meta's V-JEPA 2 world model to find videos with similar motion dynamics.
Patch-level predictive surprise from video foundation model embeddings. The embedding delta is the attention signal.
End-to-end engineering of a V-JEPA-style video world-model trainer: data curation at scale, distributed (FSDP) training, and inference optimization. From-scratch JEPA + SIGReg.
GranularVAR: Multi-scale video understanding with augmentation-graded contrastive learning and calibrated uncertainty. V-JEPA 2 encoder + granularity-conditioned decoder. GWU MS Data Science Capstone, Spring 2026.
SCOUT: frozen-encoder embedding-space prediction for sim-to-real text-based person retrieval. ECCV 2026 Workshop (AI City Challenge Track 4), accepted as a poster.
Locally-Hosted Media Gallery App with AI Similarity Search
To associate your repository with the vjepa topic, visit your repo's landing page and select "manage topics."