TokenMatch: 3D Mesh Correspondence Transformer with Curvature-Guided Tokenisation· TokenMatch:曲率引导标记化的3D网格匹配Transformer
While data-driven 3D shape correspondence estimation has recently seen substantial progres…
While data-driven 3D shape correspondence estimation has recently seen substantial progres…
We introduce S$^3$T (Self-Supervised Self-Distillation over Time), which, to the best of o…
Online 3D reconstruction models perform poorly on long videos. This happens because regres…
Evaluating physical reasoning in video models is difficult because absolute motion measure…
We propose Puffin-World, a unified multimodal architecture that integrates physical unders…
Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guid…
Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in …
3D Foundation Models (3DFMs) such as VGGT have recently pushed the boundaries of 3D vision…
Generative image models can now produce high-quality images, follow complex instructions, …
Video models are evolving into vision foundation models, yet they still lack human-like mu…
MeanFlow generators achieve fast few-step sampling by predicting average velocities over t…
Online novel view synthesis from multi-view streaming videos faces a fundamental trade-off…
Myocardial infarction (MI) remains a leading cause of mortality worldwide. Echocardiograph…
We present SceneBind, an omni-modal representation of realistic scenes with joint semantic…
Recent advances in Vision-Language Models (VLMs) have significantly improved image geo-loc…
The reliability of deepfake detectors frequently degrades under black-box adversarial tran…
Distribution shift in medical imaging remains a central bottleneck for the clinical transl…
Video generative models commonly rely on latent spaces learned by 3D Variational Autoencod…
Building interactive worlds that respond coherently to player actions has long been a shar…
Historical Manchu OCR must accommodate various visually distinct writing styles, including…
Driving-world generation has emerged as a core capability for scalable autonomous-driving …
The 11th Affective Behavior Analysis in-the-wild (ABAW11) Multi-Task Learning Challenge re…
Dermatological practice routinely involves measuring and tracking lesion size, morphology …
We present X-lens, a compact feed-forward model for metric depth estimation from a variabl…
Accurate dermatological diagnosis naturally necessitates equitable performance across dive…
LiDAR-based collaborative 3D perception in Vehicle-to-Everything (V2X) systems typically r…
Point tracking in surgery is crucial to enable applications in downstream tasks such as se…
In this paper, we propose SpectraReward, a training-free reward function that turns pretra…
Generating and editing a person's face demands high precision, as even minor modifications…
Current Video Large Language Models (Video LLMs) excel in question answering (QA) but larg…
Recent Multimodal Large Language Models (MLLMs) achieve strong performance on single-view …
This paper presents a cascaded Low-Rank Adaptation (LoRA)-based multimodal fusion framewor…
When a large disaster strikes, responders need a map of which buildings are damaged within…
Autoregressive diffusion models have enabled high-quality video generation, yet their sequ…
License plate character detection is a crucial component of intelligent transportation sys…
We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded to…
Long-form audio description (AD) requires more than describing visible actions: it must pr…
In this work, we aim to address the challenge of long-range memory in panoramic world mode…
The rapid progress of large foundation models has been driven predominantly by pretraining…
Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge case…
Vision language models (VLMs) have made remarkable progress in visual reasoning during the…
In many real-world systems, including articulated robots and biomechanical models, rotatio…
Collecting annotated plant images for automated phenotyping is often slow and expensive. P…
Reliable autonomous driving requires full-scene perception that couples foreground objects…
The deployment of large-scale foundation models, such as the Segment Anything Model 3 (SAM…
Generating long-duration, high-definition, and rhythmically synchronized dance videos dire…