CVPR2026 论文笔记 TODO¶
总计: 3337 篇 | 已完成: 3276 | 待更新: 61
- 2-shots in the dark low-light denoising with minimal data acquisition
- 2ndmatch finetuning pruned diffusion models via second-order jacobian matching | arXiv: 2506.05398
- 3d gaussian splatting at arbitrary resolutions with compact proxy anchors
- 3d gaussian splatting from unposed spike stream
- 3d gaussian splatting with self-constrained priors for high fidelity surface rec | arXiv: 2603.19682
- 3d sans 3d scans scalable pre-training from video-generated point clouds | arXiv: 2512.23042
- 3d space as a scratchpad for editable text-to-image generation
- 3d-aware implicit motion control for view-adaptive human video generation
- 3d-aware multi-task learning with cross-view correlations for dense scene unders
- 3d-fixer coarse-to-fine in-place completion for 3d scenes from a single image | arXiv: 2604.04406
- 3d-ide 3d implicit depth emergent | arXiv: 2604.03296
- 3d-vcd hallucination mitigation in 3d-llm embodied agents through visual contras
- 3drawagent teaching llm to draw in 3d with early contrastive experience | arXiv: 2604.08042
- 3dreflecnet a large-scale dataset for 3d reconstruction of reflective transparen | arXiv: 2605.10204
- 3m-ti high-quality mobile thermal imaging via calibration-free multi-camera cros | arXiv: 2511.19117
- 4c4d 4 camera 4d gaussian splatting | arXiv: 2604.04063
- 4d local modeling toward dynamic global perception for ambiguity-free rotation-i
- 4d primitive-mache glueing primitives for persistent 4d scene reconstruction
- 4dequine disentangling motion and appearance for 4d equine reconstruction from m | arXiv: 2603.10125
- 4dp-qa scalable qa for 4d perception in vision language models
- 4dworldbench a comprehensive evaluation framework for 3d4d world generation mode
- a causal marriage between vlm and irm from understanding to reasoning
- a closedform solution for debiasing visionlanguage | arXiv: 2603.12998
- a closer look at cross-domain few-shot object detection fine-tuning matters and | arXiv: 2603.28182
- a combination of noise and bilateral filters achieve supralinear and scalable ad
- a debiased reconstruction-based framework for training-free detection of ai-gene
- a difference-in-difference approach to detecting ai-generated images
- a frame is worth one token efficient generative world modeling with delta tokens | arXiv: 2604.04913
- a geometric algebra-informed 3dgs framework for wireless channel prediction | arXiv: 2605.19065
- a mixed diet makes dino an omnivorous vision encoder | arXiv: 2602.24181
- a more word-like image tokenization for mllms
- a polynomial chaos framework for causal discovery in nonlinear uncertain systems
- a provable energy-guided test-time defense boosting adversarial robustness of la
- a sanity check for multi-in-domain face forgery detection in the real world
- a self-conditioned representation guided diffusion model for realistic text-to-l
- a semantically disentangled unified model for multi-category 3d anomaly detectio | arXiv: 2603.25159
- a stitch in time learning procedural workflow via self supervised plackett luce r | arXiv: 2511.17805
- a style is worth one code unlocking code-to-style image generation with discrete
- a supervised multi-task framework for joint cryo-et restoration enabled by gener
- a temporal and content co-awareness latent diffusion for controllable hand image
- a unified perspective on adversarial membership manipulation in vision models | arXiv: 2604.02780
- a2gc asymmetric aggregation with geometric constraints for locally aggregated de
- abstract 3d perception for spatial intelligence in vision-language models
- accelerating autoregressive video diffusion via history-guided cache and residua
- accelerating diffusion via hybrid data-pipeline parallelism based on conditional
- accelerating diffusion-based video editing via heterogeneous caching beyond full
- accelerating streaming video large language models via hierarchical token compre
- acetone bridging words and colors for conditional image grading | arXiv: 2604.00530
- acot-vla action chain-of-thought for vision-language-action models
- acpv-net all-class polygonal vectorization for seamless vector map generation fr | arXiv: 2603.16616
- act like a pathologist tissue-aware whole slide image reasoning | arXiv: 2603.00667
- act2see emergent active visual perception for video reasoning | arXiv: 2605.01657
- actavatar temporally-aware precise action control for talking avatars
- action motifs self-supervised hierarchical representation of human body movement
- action-geometry prediction with 3d geometric prior for bimanual manipulation | arXiv: 2602.23814
- action-sketcher from reasoning to action via visual sketches for robotic manipul
- actiongeometry prediction with 3d geometric prior | arXiv: 2602.23814
- actionmesh animated 3d mesh generation with temporal 3d diffusion | arXiv: 2601.16148
- activation matters test-time activated negative labels for ood detection with vi | arXiv: 2603.25250
- active inference for micro-gesture recognition efe-guided temporal sampling and | arXiv: 2603.07559
- active perceptual inference a corticothalamic-inspired dynamic nested recurrent
- activead planning-oriented active learning for end-to-end autonomous driving
- activegrasp information-guided active grasping with calibrated energy-based mode | arXiv: 2511.12795
- activepolicy active gaussian reconstruction and optimization strategy based on g
- activevla injecting active perception into vision-language-action models for pre
- activityforensics a comprehensive benchmark for localizing manipulated activity | arXiv: 2604.03819
- actta rethinking test-time adaptation via dynamic activation | arXiv: 2603.26096
- ad-gbc anisotropic granular-ball skip-connection refiner for unet-based medical image seg
- adabet gradient-free layer selection for efficient training of deep neural netwo | arXiv: 2510.03101
- adacluster adaptive query-key clustering for sparse attention in video generatio | arXiv: 2604.18348
- adadextrack dynamic modulation for adaptive and generalizable dexterous manipula
- adaiat adaptively increasing attention to generated text to alleviate hallucinat
- adaprior bayesian-inspired adaptive prior correction for long-tailed continual l
- adapter shield a unified framework with built-in authentication for preventing u
- adapting a pre-trained single-cell foundation model to spatial gene expression g | arXiv: 2603.19766
- adapting point cloud analysis via multimodal bayesian distribution learning | arXiv: 2603.22070
- adaptive 3d perception for small aerial targets under sparse sampling via reinfo
- adaptive action chunking at inference-time for vision-language-action models | arXiv: 2604.04161
- adaptive anisotropic gaussian splatting for multi-contrast mri arbitrary-scale s
- adaptive auxiliary prompt blending for target-faithful diffusion generation | arXiv: 2603.19158
- adaptive bayesian early-exit networks for efficient non-transferable learning
- adaptive capacity autoregressive visual tracking
- adaptive confidence regularization for multimodal failure detection | arXiv: 2603.02200
- adaptive data augmentation with multi-armed bandit sample-efficient embedding ca
- adaptive depth lightweight rgb-t tracking with holistic token routing
- adaptive learned image compression with graph neural networks | arXiv: 2603.25316
- adaptive spatial-temporal window unlocking the potential of event cameras in het
- adaptive video distillation mitigating oversaturation and temporal collapse in f
- adaptok learning adaptive and temporally causal video tokenization in a 1d laten
- adaptvision efficient vision-language models via adaptive visual acquisition | arXiv: 2512.03794
- adaradar rate adaptive spectral compression for radar-based perception | arXiv: 2603.17979
- adasformer adaptive serialized transformers for monocular semantic scene complet | arXiv: 2603.25494
- adaspark adaptive sparsity for efficient long video understanding | arXiv: 2604.08077
- adasvd singular value decomposition with adaptive mechanisms for large multimoda
- addressing exacerbated attention sink for source-free cross-domain few-shot lear | arXiv: 2605.25799
- adseeker a knowledge-grounded reasoning framework for industry anomaly detection
- advancing cancer prognosis with hierarchical fusion of genomic proteomic and pat
- adversarial style optimization enhancing vlm jailbreaks by grpo-based stylistic
- advfm lookahead flow-matching velocity-field attacks for imperceptible and trans
- aergs-slam auto-exposure-robust stereo 3d gaussian splatting slam
- aerodgs physically consistent dynamic gaussian splatting for single-sequence aer | arXiv: 2602.22376
- aerogs scale-aware gaussian splatting for pose-free dynamic uav scene reconstruc
- affordance field intervention enabling vlas to escape memory traps in robotic ma
- affordance-first decomposition for continual learning in video-language understa
- affordgen generating diverse demonstrations for generalizable object manipulatio
- affordgrasp cross-modal diffusion for affordance-aware grasp synthesis | arXiv: 2603.08021
- affordmatcher affordance learning in 3d scenes from visual signifiers | arXiv: 2603.27970
- affostruction 3d affordance grounding with generative reconstruction | arXiv: 2601.09211
- ag-vas anchor-guided zero-shot visual anomaly segmentation with large multimodal
- agentdet a shared-blackboard multi-agent framework for zero-few-shot object dete
- agentic retoucher for texttoimage generation | arXiv: 2601.02046
- agentic video summarization via self-reflecting multimodal understanding
- agentsafe benchmarking the safety of embodied agents on hazardous instructions
- agft alignment-guided fine-tuning for zero-shot adversarial robustness of vision | arXiv: 2603.29410
- agile learning robust long-horizon manipulation via affordance-grounded bidirect
- ahs adaptive head synthesis | arXiv: 2604.15857
- aif adaptive information flow vlm | arXiv: 2604.15809
- aimdepth asymmetric image-event mamba for monocular depth estimation
- air-know arbiter-calibrated knowledge-internalizing robust network for composed | arXiv: 2604.19386
- akcmamba-yolo selective state space models for real-time object detection
- alchemint fine-grained temporal control for multi-reference consistent video gen
- alert-clip abnormality-aware latent-enhanced representation tuning of clip for v
- align once to explain feature alignment for scalable b-cosification of foundatio
- align while search belief-guided exploratory inference for world-grounded embodi
- aligning multi-character narrative image generation with multi-aspect human pref
- alignpose generalizable 6d pose estimation via multi-view feature-metric alignme | arXiv: 2512.20538
- all in one slider attribute manipulation | arXiv: 2508.19195
- all in one unifying deepfake detection tampering localization and source tracing | arXiv: 2602.23523
- all roads lead to rome incentivizing divergent thinking in vision-language model
- all vehicles can lie efficient adversarial defense in fully untrusted-vehicle co | arXiv: 2603.08498
- allnet multi-task dense prediction for degraded images
- alphamatte4k mumatting dataset and model for ultra-micro precision alpha video m
- amb3r accurate feed-forward metric-scale 3d reconstruction with backend
- amuse audio-visual benchmark and alignment framework for agentic multi-speaker u
- an empirical study on how video-llms answer video questions
- an instance-centric panoptic occupancy prediction benchmark for autonomous drivi | arXiv: 2603.27238
- an optimal transport driven approach for cultivating latent space in online incr | arXiv: 2211.16780
- anatomica localized control over geometric and topological properties for anatom
- anatomical domain shifts test-time heterogeneous adaptation for 3d human pose pr
- anchor-guided gradient alignment for incomplete multimodal learning
- anchoring and rescaling attention for semantically coherent inbetweening | arXiv: 2603.17651
- anchoring the mind of multimodal reasoners cognitive bias as a vector for jailbr
- anchorsplat feed-forward 3d gaussian splatting with 3d geometric priors | arXiv: 2604.07053
- ani3dhuman photorealistic 3d human animation with self-guided stochastic samplin | arXiv: 2602.19089
- animimic imitating 3d animation from video priors
- annotation-efficient coreset selection for context-dependent segmentation
- anomalyvfm -- transforming vision foundation models into zero-shot anomaly detec | arXiv: 2601.20524
- anthrotap learning point tracking with real-world motion | arXiv: 2507.06233
- anti-i2v safeguarding your photos from malicious image-to-video generation | arXiv: 2603.24570
- antistyler defending object detection models against adversarial patch attacks u
- ants adaptive negative textual space shaping for ood detection via test-time mll
- any2any 3d diffusion models with knowledge transfer a radiotherapy planning stud | arXiv: 2605.09622
- any4d unified feed-forward metric 4d reconstruction
- anydoc enhancing document generation via large-scale htmlcss data synthesis and | arXiv: 2603.25118
- AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual References
- anypcc compressing any point cloud with a single universal model | arXiv: 2510.20331
- apet approximation-error guided token compression for efficient vlms | arXiv: 2602.19870
- apex a decoupled memory-based explorer for asynchronous aerial object goal navig
- apple attribute-preserving pseudo-labeling for diffusion-based face swapping | arXiv: 2601.15288
- ar2-4fv anchored referring and re-identification for long-term grounding in fixe | arXiv: 2603.07758
- ar2can an architect and an artist leveraging a canvas for multi-human generation | arXiv: 2511.22690
- arcadia toward a full-lifecycle framework for embodied lifelong learning
- archon a unified multimodal model for holistic digital human generation | arXiv: 2605.30311
- archsym detecting 3d-grounded architectural symmetries in the wild
- are image-to-video models good zero-shot image editors
- are we ready for rl in text-to-3d generation a progressive investigation
- AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
- ares unifying asymmetric rgb-event stereo for probabilistic scene flow estimatio
- argus defending against multimodal indirect prompt injection via steering instru
- arm-thinker reinforcing multimodal generative reward models with agentic tool us
- arthoi taming foundation models for monocular 4d reconstruction of hand-articula | arXiv: 2603.25791
- artimuse fine-grained image aesthetics assessment with joint scoring and expert-
- artllm generating articulated assets via 3d llm | arXiv: 2603.01142
- artpro self-supervised articulated object reconstruction with adaptive integrati
- assignment-driven hash learning in a hyper-semantic space for on-the-fly categor
- asynchronous temporal modeling with two-agent framework for streaming dense vide
- at-vla adaptive tactile injection for enhanced feedback reaction in vision-langu
- atomicvla unlocking the potential of atomic skill learning in robots | arXiv: 2603.07648
- attention may i have your decision localizing generative choices in diffusion mo | arXiv: 2604.06052
- attention surgery an efficient recipe to linearize your video diffusion transfor
- attention-aware inference optimizations for large vision-language models with me
- attribution as retrieval modelagnostic aigenerated | arXiv: 2603.10583
- attribution-guided model rectification of unreliable neural network behaviors | arXiv: 2603.15656
- audioavatar personalized audio-driven whole-body talking avatars
- audiostory generating long-form narrative audio with large language models
- authorize-on-demand dynamic authorization with legality-aware intellectual prope
- autocut end-to-end advertisement video editing based on multimodal discretizatio | arXiv: 2603.28366
- autodebias automated framework for debiasing text-to-image models | arXiv: 2508.00445
- autogaze attend before attention efficient video | arXiv: 2603.12254
- autotraces autoregressive trajectory forecasting via multimodal large language m
- av-reasoner improving and benchmarking clue-grounded audio-visual counting for m
- ava vla improving vision language action models with active visual attention | arXiv: 2511.18960
- avatar forcing real-time interactive head avatar generation for natural conversa | arXiv: 2601.00664
- avatarpointillist autoregressive 4d gaussian avatarization | arXiv: 2604.04787
- avfakebench a comprehensive audio-video forgery detection benchmark for av-lmms
- aviasafe a physics-informed data-driven model for aviation safety-critical cloud
- avion aerial visionlanguage instruction from offli | arXiv: 2603.12659
- axg-reasoner error detection and explanation in long task videos with vision-lan
- b-clip text-conditioned contrastive learning for multi-granular vision-language
- b3-seg camera-free training-free 3dgs segmentation via analytic eig and beta-ber
- BA-GS: Bayesian Adaptive Gaussian Splatting for SFM-Free 3D Reconstruction
- babyvlm-v2 toward developmentally grounded pretraining and benchmarking of visio | arXiv: 2512.10932
- back to basics let denoising generative models denoise
- back to point exploring point-language models for zero-shot 3d anomaly detection | arXiv: 2603.21511
- back to source open-set continual test-time adaptation via domain compensation | arXiv: 2604.21772
- back to the feature explaining video classifiers with video counterfactual expla
- backsplit the importance of sub-dividing the background in biomedical lesion seg | arXiv: 2511.19394
- bagger backwards aggregation for mitigating drift in autoregressive video diffus
- balanced dataset distillation via modeling multiple visual pattern distribution
- balanced hierarchical contrastive learning with decoupled queries for fine-grain
- balm a model-agnostic framework for balanced multimodal learning under imbalance | arXiv: 2603.19718
- bami training-free bias mitigation in gui grounding | arXiv: 2605.06664
- barbiegait an identity-consistent synthetic human dataset with versatile cloth-c | arXiv: 2604.12221
- Basis-Oriented Low-rank Transfer for Few-Shot and Test-Time Adaptation
- batch loss score for dynamic data pruning | arXiv: 2604.04681
- batman benign knowledge alignment through malicious null space in federated back
- bayesian decomposition and semantic completion for few-shot semantic segmentatio
- bd-merging bias-aware dynamic model merging with evidence-guided contrastive lea | arXiv: 2603.03920
- bdnetbio-inspired dual-backbone small object detection network
- bea-gs beyond radiance supervision in 3dgs for precise object extraction | arXiv: 2605.09662
- beautygrpo aesthetic alignment for face retouching via dynamic path guidance and | arXiv: 2603.01163
- benchmarking endoscopic surgical image restoration and beyond | arXiv: 2505.19161
- benchmarking phd-level coding in 3d geometric computer vision | arXiv: 2603.30038
- better stronger faster tackling the trilemma in mllm-based segmentation with sim
- better than average spatially-aware aggregation of segmentation uncertainty impr | arXiv: 2603.29941
- bev-car enhancing monocular birds eye view segmentation with context-aware raste
- bev-sld self-supervised scene landmark detection for global localization with li | arXiv: 2603.17159
- beyond 3d vqas injecting 3d spatial priors into vision-language models for enhan | arXiv: 2605.30231
- beyond appearance camouflaged object detection via geometric structure
- beyond binary contrast modeling continuous skeleton action spaces with transitio
- beyond caption-based queries for video moment retrieval | arXiv: 2603.02363
- beyond caption-based queries in video moment retrieval
- beyond cls token query-driven token-level forgery purification for generalizable
- beyond duality a hybrid framework of leveraging shared and private features for
- beyond endpoints path-centric reasoning for vectorized off-road network extracti
- beyond euclidean gossip kl-barycentric consensus on heterogeneous and imbalanced
- beyond explicit language plug-and-play visual-to-linguistic modeling toward gene
- beyond fixed formulas data-driven linear predictor for efficient diffusion model | arXiv: 2604.26365
- beyond geometry artistic disparity synthesis for immersive 2d-to-3d | arXiv: 2603.05906
- beyond global scores fine grained token grounding as robust detector of lvlm hallucinations | arXiv: 2604.04863
- beyond global similarity multi-conditional retrieval for fine-grained cross-moda
- beyond graph model reliable vlm fine-tuning via random graph adapter
- beyond ground-truth leveraging image quality priors for real-world image restora | arXiv: 2603.29773
- beyond heuristic prompting a concept-guided bayesian framework for zero-shot ima | arXiv: 2603.07911
- beyond layer-wise merging chain-of-merging for vision-language models | 📄 paper_cache/CVPR2026/cvf-beyond_layer-wise_merging_chain-of-mergi.txt
- beyond matching to tiles bridging unaligned aerial and satellite views for visio
- beyond mimicry learning whole-body human-humanoid interaction from human-human d
- beyond missing modalities hypergraph conditioned diffusion for uncertainty-aware
- beyond multiple choice verifiable openqa for robust vision-language rft
- beyond myopic alignment lookahead optimization for online class-incremental lear
- beyond objects contextual synthetic data generation for fine-grained classificat | arXiv: 2510.24078
- beyond patches global-aware autoregressive model for multimodal few-shot font ge | arXiv: 2601.01593
- beyond perceptual shortcuts causal-inspired debiasing optimization for generaliz | arXiv: 2605.01324
- beyond pixel simulation pathology image generation via diagnostic semantic token | arXiv: 2512.21058
- beyond prompt degradation prototype-guided dual-pool prompting for incremental o | arXiv: 2603.02286
- beyond rule-based agents active markov games for realistic multi-agent interacti
- beyond scanpaths graph-based gaze simulation in dynamic scenes
- beyond semantic search towards referential anchoring in composed image retrieval | arXiv: 2604.05393
- beyond sequential tools a unified vlm agent system for photographic post-process
- beyond single images a comprehensive benchmark for album-level vision-language u
- beyond single solution multi-hypothesis collaborative deep unfolding network for | arXiv: 2606.03666
- beyond single solution multi-hypothesis deep unfolding network for image compres
- beyond soft label dataset distillation via orthogonal gradient matching
- beyond static frames temporal aggregate-and-restore vision transformer for human
- beyond strict pairing arbitrarily paired training for high-performance infrared
- beyond success refining elegant robot manipulation from mixed-quality data via j
- beyond text visual description assembly by probabilistic model for clip-based we
- beyond the golden data resolving the motion-vision quality dilemma via timestep | arXiv: 2603.25527
- beyond the ground truth enhanced supervision for image restoration | arXiv: 2512.03932
- beyond the static world continual category discovery under visual drift
- beyond the static-world lifelong learning for all-in-one medical image restorati
- beyond tie points satellite image block adjustment based on dense feature consis
- beyond top activations efficient and reliable crowdsourced evaluation of automat
- beyond weak supervision mllms-guided graded knowledge distillation for unsupervi
- beyond whats shared recovering lost unique information from intermediate layers
- bhcast unlocking black hole plasma dynamics from a single blurry image with long | arXiv: 2603.26777
- bi cmpstereo bidirectional cross modal prompting for event frame asymmetric stereo | arXiv: 2604.15312
- bi-bridge bidirectional diffusion bridges for low-light image enhancement
- bias in bias out finding unbiased subnetworks in vanilla models
- bias is a subspace not a coordinate a geometric rethinking of post-hoc debiasing
- bias reward models t2i | arXiv: 2604.13305
- bidirectional cross-modal prompting for event-frame asymmetric stereo
- bidirectional multimodal prompt learning with scale-aware training for few-shot | arXiv: 2408.13516
- bidirectional query-driven generation of parametric cad sketch
- bievlight bi-level learning of task-aware event refinement for low-light image e
- bifm bidirectional flow matching for few-step image editing and generation
- bigmint biologically-guided hierarchical multimodal integration for modeling mul
- bilevel layer-positioning lora for real image dehazing | arXiv: 2603.10872
- bimotion b-spline motion for text-guided dynamic 3d character generation | arXiv: 2602.18873
- binaryattention one-bit qk-attention for vision and diffusion transformers | arXiv: 2603.09582
- biomedccpl causal conditional prompt learning for biomedical vision-language mod
- biotprompt bidirectional optimal transport guided prompting for disease evolutio
- biovita biological dataset model and benchmark for visual-textual-acoustic align | arXiv: 2603.23883
- bipa bilevel prompt adaptation for underwater instance segmentation
- bipremanip learning affordance-based bimanual preparatory manipulation through a | arXiv: 2603.21679
- bit matching-based bi-directional interaction transformation network for visible
- black-box domain adaptation for object detection with retention-driven knowledge
- black-box membership inference attacks on the pre-training data of image-generat | arXiv: 2605.27020
- blackmirror black-box backdoor detection for text-to-image models via instructio | arXiv: 2603.05921
- blind spot of adaptation quantifying and mitigating forgetting in fine tuned driving models | arXiv: 2604.04857
- block-based learned image compression without blocking artifacts
- bluref unsupervised image deblurring with dense-matching references | arXiv: 2603.14176
- boosting document parsing efficiency and performance with coarse-to-fine visual
- boosting quantitive and spatial awareness for zero-shot object counting | arXiv: 2603.16129
- boosting reasoning in large multimodal models via activation replay | arXiv: 2511.19972
- boosting vision-language models towards cross-domain incremental object detectio
- boosting vision-language-action finetuning with feasible action neighborhood pri | arXiv: 2604.01570
- boostslt boosting sign language translation via a plug-and-play diffusion-based
- bootstrap dynamic-aware 3d visual representation for scalable robot learning | arXiv: 2512.00074
- bootstrap your own av-proxies adaptive contrastive and prototype learning for au
- bootstrapping multi-view learning for test-time noisy correspondence
- bootstrapping video semantic segmentation model via distillation-assisted test-t | arXiv: 2604.10950
- bop-ask object-interaction reasoning for vision-language models | arXiv: 2511.16857
- boundary-responsive differentiable gating for superpixel-based segmentation
- breaking multimodal llm safety via video-driven prompting
- breaking semantic boundaries distribution-guided semantic exploration for creati
- breaking smooth-motion assumptions a uav benchmark for multi-object tracking in
- breaking spurious correlations uncertainty-driven causal transformers for au det
- breaking the 3d dataset bottleneck fast scalable generation of aligned 3d assets
- breaking the continuum discrete distribution learning for structural mri reconst
- breaking the regional perception bottleneck of multimodal large language models
- breaking the scalability limit of multi-projector calibration with embedded came
- brepgaussian cad reconstruction from multi-view images with gaussian splatting | arXiv: 2602.21105
- brepvgae variational graph autoencoder with unified latent representation for b-
- brewing stronger features dual-teacher distillation for multispectral earth obse | arXiv: 2602.19863
- bricknet graph-backed generative brick assembly | arXiv: 2604.22984
- bridge basis-driven causal inference marries vfms for domain generalization | arXiv: 2604.26820
- bridgeeqa virtual embodied agents for real bridge inspections
- bridging brain and semantics a hierarchical framework for semantically enhanced | arXiv: 2605.14569
- bridging domain expertise and generalization for performance estimation
- bridging domains through subspace-aware model merging
- bridging facial understanding and animation via language models
- bridging fidelity-reality with controllable one-step diffusion for image super-r
- bridging pixels and words mask-aware local semantic fusion for multimodal media | arXiv: 2603.26052
- bridging privacy and provenance traceable virtual identity generation
- bridging rgb and hematoxylin components an interleaved guidance and fusion frame
- bridging the 2d-3d gap a hierarchical semantic-geometric map for vision language
- bridging the modality gap in compositional zero-shot learning via sparse alignme
- bridging the perception gap in image super-resolution evaluation | arXiv: 2503.13074
- brima bridged modality adaptation for multi-modal continual action quality asses | arXiv: 2602.19170
- bringing your portrait to 3d presence
- buildanypoint 3d building structured abstraction from diverse point clouds | arXiv: 2602.23645
- building a precise video language with human-ai oversight | arXiv: 2604.21718
- building robust vision encoders for cross-dataset evaluation in immunofluorescen
- buildinggpt auto-regressive building wireframe reconstruction model with reinfor
- bulk rna-seq guided multi-modal detection of anomalous regions in human cancer v
- bullettime decoupled control of time and camera pose for video generation
- bussard normalizing flows for bijective universal scene-specific anomalous relat | arXiv: 2603.16645
- bypassing the transport plan dynamic reweighting for out-of-distribution detecti
- bézier degradation modeling for lidar-based human motion capture | arXiv: 2605.19620
- c-genreg training-free 3d point cloud registration by multi-view-consistent geom | arXiv: 2604.16680
- c-lav conditional latent velocity field denoising for weather-robust lidar place
- c2fg control classifier-free guidance via score discrepancy analysis | arXiv: 2603.08155
- ca-lora concept-aware lora for domain-aligned segmentation dataset generation | arXiv: 2503.22172
- cad-refiner a unified framework for cad generation and iterative editing
- cadc content adaptive diffusion-based generative image compression
- cadfs a big cad program dataset and framework for computer-aided design with lar | arXiv: 2605.01925
- camera control for text-to-image generation via learning viewpoint tokens
- can a second-view image be a language geometric and semantic cross-modal reasoni
- can natural image autoencoders compactly tokenize fmri volumes for long-range dy | arXiv: 2604.03619
- can we build scene graphs not classify them flowsg progressive image-conditioned
- canoncgt reference-based color grading via canonical pivot representation | arXiv: 2606.01638
- capnav benchmarking vision language models on capability-conditioned indoor navi
- capt confusion-aware prompt tuning for reducing vision-language misalignment | arXiv: 2603.02557
- captionformer unified segmentation tracking and captioning for spatio-temporal o
- car-sam cross-attention reconstruction for post-training quantization of the seg
- card a multi-modal automotive dataset for dense 3d reconstruction in challenging | arXiv: 2605.05014
- care a molecular-guided foundation model with adaptive region modeling for whole | arXiv: 2602.21637
- care what fails contrastive anchored-reflection for verifiable multimodal reason
- care-edit condition-aware routing of experts for contextual image editing | arXiv: 2603.08589
- careflow cyclic adaptive rectified flow for multimodal fusion | arXiv: 2602.19140
- cari4d category agnostic 4d reconstruction of human object interaction | arXiv: 2512.11988
- carlos retrieval via concise assessment representation of loras at scale
- caspa graph-structured concept anchors for modality-agnostic adaptation in visio
- casr a robust cyclic framework for arbitrary large-scale super-resolution with d
- cast context-aware dynamic latent space transformation for interactive text-to-i
- cast-bench benchmarking causal chain-grounded spatio-temporal reasoning for vide | arXiv: 2605.23216
- cat-gs efficient 3dgs rendering for large-scale scenes with inter-frame caching
- catalyst4d highfidelity 3dto4d scene editing via d | arXiv: 2603.12766
- catnet collaborative alignment and transformation network for cooperative percep
- causallens sensitivity-guided multi-head causal intervention for hallucination m
- causalvad de-confounding end-to-end autonomous driving via causal intervention | arXiv: 2603.18561
- cc-vqa conflict- and correlation-aware method for mitigating knowledge conflict | arXiv: 2602.23952
- cccaption dual-reward reinforcement learning for complete and correct image capt | arXiv: 2602.21655
- ccf complementary collaborative fusion for domain generalized multi-modal 3d obj | arXiv: 2603.23276
- cd-buffer complementary dual-buffer framework for test-time adaptation in advers | arXiv: 2603.26092
- cdics delving into fine-grained attribute for in-context segmentation via compos
- cell-type prototype-informed neural network for gene expression estimation from | arXiv: 2603.18461
- cf-ipt cross-modal fusion interactive prompt tuning of vision-language pre-train
- cfg-ctrl control-based classifier-free diffusion guidance | arXiv: 2603.03281
- cg-reasoner centroid-guided positional reasoning segmentation for medical imagin
- cghair compact gaussian hair reconstruction with card clustering | arXiv: 2604.03716
- cgl advancing continual gui learning via reinforcement fine-tuning
- cgu-bayes causal graph uncertainty-guided bayesian inference for domain generali
- chain of event-centric causal thought for physically plausible video generation | arXiv: 2603.09094
- chain of world world model thinking in latent motion | arXiv: 2603.03195
- chain-of-frames advancing video understanding in multimodal llms via frame-aware
- chain-of-thought guided multi-modal object re-identification
- chal causal-guided hierarchical anomaly-aware learning for moving infrared small
- changebridge spatiotemporal image generation with multimodal controls for remote | arXiv: 2507.04678
- changes in real time online scene change detection with multi-view fusion | arXiv: 2511.12370
- charge a comprehensive novel view synthesis benchmark and dataset to bind them a | 📄 paper_cache/CVPR2026/cvf-charge_a_comprehensive_novel_view_synthe.txt
- chart-fr1 visual focus-driven fine-grained reasoning on dense charts
- chartnet a million-scale high-quality multimodal dataset for robust chart unders | arXiv: 2603.27064
- chartr evaluating reasoning accuracy and robustness in chart question answering
- cheem continual learning by reuse new adapt and skip -- a hierarchical explorati | arXiv: 2303.08250
- chips efficient clip adaptation via curvature-aware hybrid influence-based data | arXiv: 2511.18519
- chirp dataset towards long-term individual-level behavioral monitoring of bird p
- chordedit one-step low-energy transport for image editing | arXiv: 2602.19083
- chorus multi-teacher pretraining for holistic 3d gaussian scene encoding
- cica coupling confidence-aware pretraining with confidence-informed attention fo
- cigma causal information-gain mechanistic attribution of attention heads in visi
- cigpose causal intervention graph neural network for whole-body pose estimation | arXiv: 2603.09418
- ciice intrinsic concept extraction compositional | arXiv: 2603.11795
- cinebrain a large-scale multi-modal audiovisual brain dataset for brain-conditio
- cinematic audio source separation using visual cues | arXiv: 2603.26113
- cinescene implicit 3d as effective scene representation for cinematic video gene
- cinesrd leveraging visual acoustic and linguistic cues for open-world visual med | arXiv: 2603.16966
- circuit mechanisms for spatial relation generation in diffusion models | arXiv: 2601.06338
- circular-dpo aligning multi-stage 3d generative models via preference feedback l
- clair obscur an illumination-aware method for real-world image vectorization
- clay conditional visual similarity | arXiv: 2604.11539
- clay-to-stone phase-wise 3d gaussian splatting for monocular articulated hand-ob
- clcr cross-level semantic collaborative representation for multimodal learning | arXiv: 2602.19605
- cleaning the pool progressive filtering of unlabeled pools in deep active learni | arXiv: 2511.22344
- clex complementary label exchange learning for noisy facial expression recogniti
- climaood improving anomaly segmentation via physically realistic synthetic data | arXiv: 2512.02686
- clip shortsighted beyond first sentence | arXiv: 2602.22419
- clip-like model as a foundational density ratio estimator | arXiv: 2506.22881
- clipgstream clip-stream gaussian splatting for any length and any motion multi-v | arXiv: 2604.13746
- clipoint3d language-grounded few-shot unsupervised 3d point cloud domain adaptat | arXiv: 2602.20409
- clivis unleashing cognitive map through linguistic-visual synergy for embodied v
- CLP: A Real-World Dataset of Contaminated Lens Protectors for Robust Semantic Segmentation
- cluster-aware neural collapse prompt tuning for long-tailed generalization of vi
- cluster-wise spatio-temporal masking for efficient video-language pretraining | arXiv: 2603.22953
- clustermark towards robust watermarking for autoregressive image generators with | arXiv: 2508.06656
- cme-cad heterogeneous collaborative multi-expert reinforcement learning for cad code gen
- cocovideo the high-quality commercial-model-based contrastive benchmark for ai-g | arXiv: 2606.00101
- cod a diffusion foundation model for image compression | arXiv: 2511.18706
- coded-e2lf coded aperture light field imaging from events | arXiv: 2602.22620
- codedance a dynamic tool-integrated mllm for executable visual reasoning | arXiv: 2512.17312
- codepercept code-grounded visual stem perception for mllms | arXiv: 2603.10757
- codev code with images for faithful visual reasoning via tool-aware policy optim
- cofida-m concept-aware feature modulation for cross-domain adaptation with image | arXiv: 2605.31591
- cog confidence-aware optimal geometric correspondence for unsupervised single-re | arXiv: 2603.00493
- cogdriver integrating cognitive inertia for temporally coherent planning in auto
- cogniedit dense gradient flow optimization for fine-grained image editing
- cogniverse revolutionizing multi-modal retrieval-augmented generation with cogni | arXiv: 2605.29602
- coin coverage and informativeness-guided token reduction for efficient large mul
- coin3d revisiting configuration-invariant multi-camera 3d object detection | arXiv: 2603.05042
- colavla leveraging cognitive latent reasoning for hierarchical parallel trajecto | arXiv: 2512.22939
- colc communication-efficient collaborative perception with lidar completion | arXiv: 2603.00682
- cologen progressive learning of concept-localization duality for unified image g | arXiv: 2602.22150
- color the devil is in scene coordinate regression for large-scale visual localiz
- color when it counts grayscale-guided online triggering for always-on streaming | arXiv: 2603.22466
- color-encoded illumination for high-speed volumetric scene reconstruction | arXiv: 2604.26920
- com pt chain of models pretraining | arXiv: 2604.12391
- como learning continuous latent motion from internet videos for scalable robot l | arXiv: 2505.17006
- comp collaborative multi-mode pruning for vision-language models | arXiv: 2604.02956
- compbench benchmarking complex instruction-guided image editing
- competitorformer mitigating query conflicts for 3d instance segmentation via com
- complementary prototype mapping for efficient multimodal anomaly detection
- complet4r geometric complete 4d reconstruction
- compose a unified completion-pose framework for robust category-level object pos | arXiv: 2605.25553
- composing concepts from images and videos via concept-prompt binding | arXiv: 2512.09824
- composite-attribute person re-identification via pose-guided disentanglement
- compositional text-to-image generation via region-aware bimodal direct preferenc
- compressed-domain-aware online video super-resolution | arXiv: 2603.07694
- computation and communication efficient federated unlearning via on-server gradi | arXiv: 2603.13795
- computer vision with a superpixelation camera
- conan progressive learning to reason like a detective over multi-scale visual ev
- concept-aware lora for domain-aligned segmentation dataset generation
- concept-guided fine-tuning steering vits away from spurious correlations to impr | arXiv: 2603.08309
- conceptprism concept disentanglement in personalized diffusion models via residu | arXiv: 2602.19575
- condensed test-time adaptation of vlms for action recognition
- conditional factuality controlled llms with generalization certificates via conf | arXiv: 2603.27403
- conesep cone-based robust noise-unlearning compositional network for composed im | arXiv: 2604.20358
- confidence-guided multi-scale aggregation for sparse-view high-resolution 3d gau
- conflict-aware adaptive cross-reconstruction for multimodal sentiment analysis
- confusion-aware spectral regularizer for long-tailed recognition
- consensus entropy harnessing multi-vlm agreement for self-verifying and self-imp
- consensus vs controversy mapping the decision space where architectures diverge
- consid-gen view-consistent and identity-preserving image-to-video generation
- consistcompose multimodal layout control | arXiv: 2511.18333
- consistency beyond contrast enhancing open-vocabulary object detection robustnes
- consisvla-4d advancing spatiotemporal consistency in efficient 3d-perception and | arXiv: 2605.05126
- contact-aware neural dynamics
- content-aware dynamic patchification for efficient video diffusion
- content-aware frequency encoding for implicit neural representations with fourie
- context-nav context-driven exploration and viewpoint-aware 3d spatial reasoning | arXiv: 2603.09506
- continual distillation of teachers from different domains | arXiv: 2605.04059
- contrastive cross-bag augmentation for multiple instance learning-based whole sl
- conversational image segmentation grounding abstract concepts with scalable supe
- convexity-aware noise calibration a self-supervised framework for noise-level-un
- convolutional neural networks driven by content similarity
- coordspeaker exploiting gesture captioning for coordinated caption-empowered co-
- cope consistent occlusion and prompt enhancement network for occluded person re-
- copo causal-oriented policy optimization for hallucinations of mllms
- copy-transform-paste zero-shot object-object alignment guided by vision-language
- copylens towards copyrighted characters infringement detection via copyright-awa
- core compact object-centric representations as a new paradigm for token merging
- corim conflict-driven risk minimization for dynamic multimodal fusion
- corogs contextual gaussian splatting for robust large-deviation view synthesis
- correspondence-attention alignment for multi-view diffusion models
- cot-edit let cot guide instruction video editing
- counterfactual vla self-reflective vision-language-action model with adaptive re
- countgd generalized prompting for open-world counting | arXiv: 2512.23351
- coupling liquid time-constant encoders with modern hopfield memory
- cov-align efficient fine-grained cross-modal alignment with cohesive visual sema
- cov2pose leveraging spatial covariance for direct manifold-aware 6-dof object po
- covft context-aware visual fine-tuning for multimodal large language models | arXiv: 2603.21077
- crackssm reviving ssms for crack segmentation via dynamic scanning
- craft aligning diffusion models with finetuning is easier than you think | arXiv: 2603.18991
- craft-lora content-style personalization via rank-constrained adaptation and tra
- craftmesh high-fidelity generative mesh manipulation via poisson seamless fusion
- creval an automated interpretable evaluation for creative image manipulation und
- creward a type-specific creativity reward model | arXiv: 2511.19995
- crft consistent-recurrent feature flow transformer for cross-modal image registr | arXiv: 2604.05689
- crit graph-based automatic data synthesis to enhance cross-modal multi-hop reaso | arXiv: 2604.01634
- critical patch-aware sparse prompting with decoupled training for continual lear | arXiv: 2604.07399
- cross from left to right brain adaptive text dreamer for vision-and-language nav
- cross-architecture adaptation cloud-edge continual test-time adaptation with dyn
- cross-axis feature fusion with joint-wise motion difference prediction for text- | arXiv: 2606.01014
- cross-domain dual-stream feature disentanglement for brain disorder prediction w
- cross-domain few-shot segmentation via multi-view progressive adaptation | arXiv: 2602.05217
- cross-hand latent representation for vision-language-action models
- cross-instance gaussian splatting registration via geometry-aware feature-guided | arXiv: 2603.21936
- cross-modal attention calibration for lvlm hallucination mitigation
- cross-modal emotion transfer for emotion editing in talking face video | arXiv: 2604.07786
- cross-modal fuzzy alignment network for text-aerial person retrieval and a large | arXiv: 2603.20721
- cross-modal guided visual synthesis for data-efficient multimodal depression rec
- cross-modal identity mapping minimizing information loss in modality conversion | arXiv: 2603.01696
- cross-scale pansharpening via scaleformer and the panscale benchmark | arXiv: 2603.00543
- cross-slice knowledge transfer via masked multi-modal heterogeneous graph contra | arXiv: 2603.22821
- cross-view distillation and adaptive masking for incomplete multi-view multi-lab
- cross-view splatter feed-forward view synthesis with georeferenced images | arXiv: 2605.19656
- crossearth-gate fisher-guided adaptive tuning engine for efficient adaptation of
- crosshoi-bench a unified benchmark for hoi evaluation across vision-language mod | arXiv: 2508.18753
- crossvl complexity-aware feature routing and paired curriculum for cross-view vi | arXiv: 2605.09802
- crowdgaussian reconstructing high-fidelity 3d gaussians for human crowd from a s | arXiv: 2603.17779
- crown a unified framework for anti-aliased downsampling and phase-calibrated fus
- cryohype reconstructing a thousand cryo-em structures with transformer-based hyp | arXiv: 2512.06332
- cryokraqen kernel-regularized annealing for quantized embedding networks in cryo
- cryosense compressive sensing enables high-throughput microscopy with sparse and | arXiv: 2511.12931
- csf black-box fingerprinting via compositional semantics for text-to-image model | arXiv: 2604.16363
- ctcal rethinking text-to-image diffusion models via cross-timestep self-calibrat | arXiv: 2603.20741
- cube bspline 3d faces | arXiv: 2604.12894
- cubecomposer spatio-temporal autoregressive 4k 360 video generation from perspec | arXiv: 2603.04291
- cubecomposer spatio-temporal autoregressive 4k 360deg video generation from pers | 📄 paper_cache/CVPR2026/cvf-cubecomposer_spatio-temporal_autoregress.txt
- cubic coordinated unified bimanual perception and control framework
- cubic discrete diffusion discrete visual generation on high-dimensional represen | arXiv: 2603.19232
- cue concept-aware multi-label expansion to mitigate concept confusion in long-ta | arXiv: 2605.01309
- cupid generative 3d reconstruction via joint object and pose modeling
- cure curriculum-guided multi-task training for reliable anatomy grounded report | arXiv: 2601.15408
- curriculum group policy optimization adaptive sampling for unleashing the potent
- curvature-aware captioning leveraging geodesic attention for 3d scene understand
- curvature-aware zeroth-order optimization for memory-efficient test-time adaptat
- curve a benchmark for cultural and multilingual long video reasoning
- customized fusion a closed-loop dynamic network for adaptive multi-task-aware in | arXiv: 2604.08924
- customtex high-fidelity indoor scene texturing via multi-reference customization | arXiv: 2603.19121
- cut to the chase training-free multimodal summarization via chain-of-events | arXiv: 2603.06213
- cva context-aware video-text alignment for video temporal grounding | arXiv: 2603.24934
- cycle-consistent tuning for layered image decomposition | arXiv: 2602.20989
- cyclebev regularizing view transformation networks via view cycle consistency fo | arXiv: 2602.23575
- cyclemanip enabling cycle-based manipulation via effective history perception an
- cyclemanip enabling cyclic task manipulation via effective historical percepti | arXiv: 2512.01022
- d-convexity a unified differentiable convex shape prior via quasi-concavity for | arXiv: 2605.19210
- d-prism differentiable primitives for structured dynamic modeling | arXiv: 2604.17082
- d2-fosa dual-diffusion guided eeg-to-image reconstruction with frequency-oriente
- d2c diffusion dataset condensation | arXiv: 2507.05914
- d2cache second-order delta caching for higher video diffusion acceleration
- d2dewarp dual dimensions geometric representation learning based document image | arXiv: 2507.08492
- d2fanet enhancing video object detection with dual-domain feature aggregation ne
- d3d-vlp dynamic 3d vision-language-planning model for embodied grounding and nav
- d3fer dual channel and dual branch network for robust facial expression recognit
- da-mamba learning domain-aware state space model for global-local alignment in d | arXiv: 2603.18757
- da-vae plug-in latent compression for diffusion via detail alignment | arXiv: 2603.22125
- dabo difficulty-aware bayesian optimization with diffusion-learned priors
- dage dual-stream architecture for efficient and fine-grained geometry estimation | arXiv: 2603.03744
- dance across shifts forward-facilitation continual test-time adaptation through | arXiv: 2605.18608
- darc dual adjustment reasoning with counterfactuals for trustworthy chest x-ray
- dark3r learning structure from motion in the dark | arXiv: 2603.05330
- darkact a rgb-thermal dataset and fusion framework for multimodal low-light acti
- darkshake-dvs event-based human action recognition under low-light and shaking c
- dash a meta-attack framework for synthesizing effective and stealthy adversarial | arXiv: 2508.13309
- data leakage detection and de-duplication in large scale geospatial image datase | arXiv: 2304.02296
- data-centric meta-learning for robust few-shot generalization
- dataset distillation by influence matching
- dawn pixel motion diffusion robot control | arXiv: 2509.22652
- dbmsolver a training-free diffusion bridge sampler for high-quality image-to-ima | arXiv: 2605.05889
- DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation
- dcw snr t bias diffusion | arXiv: 2604.16044
- ddsf robust few-shot learning via disentangled subspaces with determinantal poin
- dear fine-grained vlm adaptation by decomposing attention head roles | arXiv: 2603.01111
- debiased sample selection for learning with noisy labels
- deciphering genotype-phenotype mechanisms from high-content profiling via knowle
- deco frequency-decoupled pixel diffusion for end-to-end image generation | arXiv: 2511.19365
- Decoding 3D Perception via BrainSSD: Synergistic Fusion of EEG Representations from Static and Dynamic Visual Streams
- decompose and transfer cot-prompting enhanced alignment for open-vocabulary temp | arXiv: 2603.24030
- decompose mix adapt a unified framework for parameter-efficient neural network r
- deconstructing the failure of ideal noise correction a three-pillar diagnosis | arXiv: 2603.12997
- decouple to generalize context-first self-evolving learning for data-scarce visi
- decouple your discovery and memory in continual generalized category discovery
- decoupled residual denoising diffusion models for unified and data efficient ima | arXiv: 2606.01048
- decoupling bias aligning distributions synergistic fairness optimization for dee
- decoupling defense strategies for robust image watermarking | arXiv: 2602.20053
- decoupling stability and plasticity for multi-modal test-time adaptation | arXiv: 2603.00574
- decoupling vision and language codebook anchored visual adaptation | arXiv: 2602.19449
- decovln decoupling observation reasoning and correction for vision-and-language | arXiv: 2603.13133
- dedelayed deleting remote inference delay via on-device correction | arXiv: 2510.13714
- deepalign mitigating modality conflict through modality-specific alignment
- deeper thought weaker aim understanding and mitigating perceptual impairment dur
- deepfakeimpact a two-stage benchmark with real-world impact in deepfake detectio
- deepprotect proactive face-swapping defense using identity blending and attribut
- DeepScan: A Training-Free Framework for Visually Grounded Reasoning in Large Vision-Language Models
- defect cue-preserved structural feature refinement for few-shot anomaly detectio
- deformable gaussian occupancy decoupling rigid and nonrigid motion with factoriz | arXiv: 2605.28587
- deformation-based in-context learning for point cloud understanding | arXiv: 2604.02845
- degradation-consistent test-time adaptation for all-in-one image restoration
- degradation-robust fusion an efficient degradation-aware diffusion framework for
- dehallu3d hallucination-mitigated 3d generation from a single image via cyclic v
- dejavu towards experience feedback learning for embodied intelligence
- Delta Rectified Flow Sampling for Text-to-Image Editing
- delving aleatoric uncertainty in medical image segmentation via vision foundatio
- demo2tutorial from human experience to multimodal software tutorials | arXiv: 2606.03951
- demofungrasp universal dexterous functional grasping via demonstration-editing r
- den tp a density balanced data curation and evaluation framework for trajectory | arXiv: 2409.17385
- denoise and align towards source-free uda for robust panoramic semantic segmenta
- denoising fast and slow difficulty-aware adaptive sampling for image generation | arXiv: 2604.19141
- dense metric depth completion from sparse direct time-of-flight sensors
- depth any endoscopy towards self-supervised generalizable depth estimation in mo
- depth any panoramas a foundation model for panoramic depth estimation
- depth hypothesis guided iterative refinement for event-image monocular depth est
- depthfocus controllable depth estimation for see-through scenes
- dervos decoupling consistent trajectory generation and multimodal understanding
- design your ad personalized advertising image and text generation with unified a | arXiv: 2605.12138
- designing instance-level sampling schedules via reinforce with james-stein shrin | arXiv: 2511.22177
- designing to forget deep semi-parametric models for unlearning | arXiv: 2603.22870
- detach decomposed spatio-temporal alignment for exocentric video and ambient sen
- detecting ai-generated forgeries via iterative manifold deviation amplification | arXiv: 2602.18842
- detecting unknown objects via energy-based separation | arXiv: 2603.29954
- detectsci toward object-guided roi reconstruction for high-resolution video snap
- deva fine-tuning multimodal large language models for visual perception tasks
- dex-portrait disentangled and expressive portrait animation via explicit and lat
- dexterous world models
- df2-vb dual-level fuzzy fusion with view-specific boosting for multi-view multi-
- dfd-hr generalizable deepfake detection via hierarchical routing learning
- dggt feedforward 4d reconstruction of dynamic driving scenes using unposed image
- dgs dual gradient and semantic-shift guided low-rank adaptation for class increm
- diagnose correct and learn from manipulation failures via visual symbols | arXiv: 2512.02787
- diagnosing and repairing unsafe channels in vision-language models via causal di | arXiv: 2603.27240
- dicart advancing category-level articulated object pose estimation in discrete s
- dictionary aligned concept control for safeguarding multimodal llms | arXiv: 2604.08846
- diff-semier transparency-aware adaptive fusion diffusion model with generative p
- diff4splat controllable 4d scene generation with latent dynamic reconstruction m | arXiv: 2511.00503
- diffbmp differentiable rendering with bitmap primitives | arXiv: 2602.22625
- differences that matter auditing models for capability gap discovery and rectifi
- differentiable adaptive 4d structured illumination for joint capture of shape an | arXiv: 2605.06214
- differentiable laplacian matrix guided superpixel segmentation
- differentiable stroke planning with dual parameterization for efficient and high
- differentiable vector quantization for rate-distortion optimization of generativ
- differentially private 2d human pose estimation
- diffgraph an automated agent-driven model merging framework for in-the-wild text
- diffsoup direct differentiable rasterization of triangle soup for extreme radian
- diffusion forcing planner history-annealed planning with time-dependent guidance
- diffusion mental averages | arXiv: 2603.29239
- diffusion mri transformer with a diffusion space rotary positional embedding d-r
- diffusion probe generated image result prediction using cnn probes | arXiv: 2602.23783
- diffusion sampling path tells more an efficient plug-and-play strategy for sampl
- diffusion with a linguistic compass steering the generation of clinically plausi
- diffusion-based native adversarial synthesis for enhanced medical segmentation g
- diffusion-based srgb real noise generation via prompt-driven noise representatio | arXiv: 2603.04870
- diffusionff a diffusion-based framework for joint face forgery detection and fin
- diffusionharmonizer bridging neural reconstruction and photorealistic simulation
- dig differential grounding for enhancing fine-grained perception in multimodal l
- digraphhal-bench evaluating multimodal large language models on complex directed
- dimos disentangling instance-level moving object segmentation
- dino eats clip adapting beyond knowns for open-set 3d object retrieval | arXiv: 2604.19432
- dip taming diffusion models in pixel space | arXiv: 2511.18822
- direct segmentation without logits optimization for training-free open-vocabular | arXiv: 2604.07723
- directfisheye-gs enabling native fisheye input in gaussian splatting with cross- | arXiv: 2604.00648
- direction-aware 3d large multimodal models
- disca accelerating video diffusion transformers wi | arXiv: 2602.05449
- disco-gs gaussian splatting in dynamic color lighting
- discover segment and select a progressive mechanism for zero-shot camouflaged ob | arXiv: 2602.19944
- discovering adaptive task dependencies for efficient multi-task representation c
- discriminative perception via anchored description for reasoning segmentation | arXiv: 2603.04002
- disentangle-then-align non-iterative hybrid multimodal image registration via cr | arXiv: 2603.19623
- disentangled textual priors for diffusion-based image super-resolution | arXiv: 2603.07430
- disentangling to re-couple resolving the similarity-controllability paradox in s | arXiv: 2604.00849
- distilling balanced knowledge from a biased teacher | arXiv: 2506.18496
- distilling quasi-conformal mapping a generalizable and efficient solution for wi
- distilling unsigned distance function for surface reconstruction from 3d gaussia
- distributed image compression with multimodal side information at extremely low
- distribution-aligned multimodal fusion for robust object detection
- dit-distill open-set fine-grained retrieval via generative curriculum knowledge
- dit360 high-fidelity panoramic image generation via hybrid training
- ditic aligned diffusion transformer for efficient | arXiv: 2603.13162
- diverse video generation with determinantal point process-guided policy optimiza
- diversedit towards diverse representation learning in diffusion transformers | arXiv: 2603.04239
- diversegrpo mitigating mode collapse in image generation via diversity-aware grp
- diversity over uniformity rethinking representation in generated image detection | arXiv: 2603.00717
- divide conquer and aggregate asymmetric experts for class-imbalanced semi-superv
- divide then ground adapting frame selection to query types for long-form video u | arXiv: 2512.04000
- dk-ddil adaptive knowledge retention for dynamic domain-incremental learning in
- dlvp-clip enhancing fine-grained zero-shot anomaly detection via dynamic local v
- dlwm dual latent world models enable holistic gaussian-centric pre-training in a | arXiv: 2604.00969
- dmaligner enhancing image alignment via diffusion model based view synthesis | arXiv: 2602.23022
- dmgd train-free dataset distillation with semantic-distribution matching in diff
- dmllm-tts self-verified and efficient test-time scaling for diffusion multi-moda
- dnf-sr dual-input and negative-aware feature fine-tuning for real-world image su
- do less achieve more do we need every-step optimization for rl fine-tuning of di
- do vision-language models measure up benchmarking visual measurement reading wit
- do vlms perceive or recall probing visual perception vs memory with classic visu
- do you have freestyle expressive humanoid locomotion via audio control
- do you see what i am pointing at gesture-based egocentric video question answeri | arXiv: 2603.12533
- docpruneefficient document question answering via background question and compre | arXiv: 2604.22281
- docseeker long document understanding | arXiv: 2604.12812
- does yolo really need to see every training image in every epoch | arXiv: 2603.17684
- domain-skewed federated learning with feature decoupling and calibration | arXiv: 2603.14238
- dont show pixels show cues unlocking visual tool reasoning in language models vi
- downscaling intelligence exploring perception and reasoning bottlenecks in small | arXiv: 2511.17487
- dp-fedadamw an efficient optimizer for differentially private federated large mo
- dpcache denoising path planning diffusion accel | arXiv: 2602.22654
- dpgf-net dual-prior guided fusion network for joint assessment of perceptual qua
- dpl decoupled prototype learning for enhancing robustness of vision-language tra
- dr seg revisiting grpo training for visual large language models through percept
- draft and refine with visual experts | arXiv: 2511.11005
- drainage a unifying framework for addressing class uncertainty
- drama next-gen dynamic orchestration for resilient multi-agent ecosystems in flu
- dream document recognition with explicit adaptive memory
- dreamshot storyboard synthesis | arXiv: 2604.17195
- dreamsr towards ultra-high-resolution image super-resolution via a receptive-fie
- dreamstereo towards real-time stereo inpainting for hd videos
- dreamstyle a unified framework for video stylization
- driffusion draft-and-refine process parallelizes diffusion models with ease
- drift-resilient temporal priors for visual tracking | arXiv: 2604.02654
- drive my way preference alignment of vision-language-action model for personaliz | arXiv: 2603.25740
- drivecombo benchmarking compositional traffic rule reasoning in autonomous drivi
- drivelaw unifying planning and video generation in a latent driving world | arXiv: 2512.23421
- drivemoe mixture-of-experts for vision-language-action model in end-to-end auton | arXiv: 2505.16278
- drivepi spatial-aware 4d mllm for unified autonomous driving understanding perce
- drivepts a progressive learning framework with textual and structural enhancemen
- drivergaze360 omnidirectional driver attention with object-level guidance | arXiv: 2512.14266
- drivevln towards mapless vision-and-language navigation in autonomous driving
- driving on registers
- drm diffusion-based reward model with step-wise guidance
- drocc depth region guided 3d occupancy | arXiv: 2603.01007
- droid-slam in the wild | arXiv: 2603.19076
- dropping anchor and spherical harmonics for sparse-view gaussian splatting | arXiv: 2602.20933
- drs-gui dynamic region search for training-free gui grounding
- dsca dynamic subspace concept alignment for lifelong vlm editing | arXiv: 2604.07965
- dsert roll robust multi modal perception for diverse driving conditions | arXiv: 2604.03685
- dsert-roll robust multi-modal perception for diverse driving conditions with ste
- dsflash panoptic scene graph realtime | arXiv: 2603.10538
- dual band thermal videography separating time-varying reflection and emission ne | arXiv: 2509.11334
- dual graph regularized deep unfolding network for guided depth map super-resolut
- dual-agent reinforcement learning for adaptive and cost-aware visual-inertial od | arXiv: 2511.21083
- dual-branch distilled transformer for efficient asymmetric uav tracking
- dual-estimator decoupling global and local semantic shift for drift compensation
- dual-granularity memory for efficient video generation
- dual-level adapter boosting prompt-free curvilinear structure segmentation
- dual-level confidence based implicit self-refinement for medical visual question
- dual-level hypergraph generation for addressing feature scarcity in whole-slide
- dual-prototype-guided multi-task learning for unsupervised anomaly detection and
- Duala: Dual-Level Alignment of Subjects and Stimuli for Cross-Subject fMRI Decoding
- dualmirage hunting stealthy multimodal llm agents via captchas with contour and
- dualprim compact 3d reconstruction with positive and negative primitives
- dualreg dual-space filtering and reinforcement for rigid registration | arXiv: 2508.17034
- dualsplat robust 3d gaussian splatting via pseudo-mask bootstrapping from recons
- duet-vlm dual stage unified efficient token reduction for vlm training and infer | arXiv: 2602.18846
- duetmerging synergizing dynamic and static strategies for mitigating task interf
- duetsvg unified multimodal svg generation with internal visual guidance
- duo-vsr dual-stream distillation for one-step video super-resolution | arXiv: 2603.22271
- duogen towards autonomous interleaved multimodal generation
- duomo dual motion diffusion for world-space human reconstruction | arXiv: 2603.03265
- dyadit a multi-modal diffusion transformer for socially favorable dyadic gesture
- dyfclt dynamic frequency-decoupled cross-modal learning transformer for multimod
- dynamic black-hole emission tomography with physics-informed neural fields | arXiv: 2602.08029
- dynamic exposure burst image restoration
- dynamic label noise suppression with optimal teacher pool for facial expression
- dynamic logits adjustment and exploration for test-time adaptation in vision lan
- dynamic magic unleashing restricted knowledge for lifelong person re-identificat
- dynamic momentum recalibration in online gradient learning | arXiv: 2603.06120
- dynamic stream network for combinatorial explosion problem in deformable medical
- dynamic token reweighting for robust vision-language models | arXiv: 2505.17132
- dynamic visual slam using a general 3d prior
- dynamic-editor training-free text-driven 4d scene editing with multimodal diffus
- dynamic-static decomposition for novel view synthesis of dynamic scenes with spi
- dynamicgtr leveraging graph topology representation preferences to boost vlm cap | arXiv: 2602.21864
- dynamics language-based representation for inferring rigid-body dynamics from vi
- dynamics-aware preference optimization for vision-language models
- dynamicsboost dynamic plausible video generation via annotation-free continuatio
- dynamictree interactive real tree animation via sparse voxel spectrum
- dynamicvggt learning dynamic point maps for 4d scene reconstruction in autonomou
- dynavid learning to generate highly dynamic videos using synthetic motion data | arXiv: 2604.01666
- dynbridge bridging imagination and control through interaction dynamics for robo
- dynfusion rethinking condition fusion for adaptive multi-conditional text-to-ima
- e-3dpsm a state machine for event-based egocentric 3d human pose estimation | arXiv: 2604.08543
- e-comiq-zh a human-aligned dataset and benchmark for fine-grained evaluation of | arXiv: 2602.21698
- e-rayzer self-supervised 3d reconstruction as spatial visual pre-training | arXiv: 2512.10950
- e2-sci elastic edge-cloud speculative decoding via credit inertia
- e2egs event-to-edge gaussian splatting for pose-free 3d reconstruction | arXiv: 2603.14684
- e3ad an emotion-aware vision-language-action model for human-centric end-to-end
- eaglenet energy-aware fine-grained relationship learning network for text-video | arXiv: 2603.25267
- eaglevision a dual-stage framework with bev-grounding-based chain-of-thought for | arXiv: 2512.15160
- easy2hard from partially to fully unmatched modalities as negative samples in co
- easy3e feed-forward 3d asset editing via rectified voxel flow | arXiv: 2602.21499
- easyomnimatte taming pretrained inpainting diffusion models for end-to-end video
- EasyV2V: A High-quality Instruction-based Video Editing Framework
- ebmc multimodal sentiment analysis | arXiv: 2604.12518
- echoes of ownership adversarial-guided dual injection for copyright protection i | arXiv: 2602.18845
- echoes over time unlocking length generalization in video-to-audio generation mo | arXiv: 2602.20981
- echofoley event-centric hierarchical control for video grounded creative sound g
- echopose 6d pose estimation of sparse echocardiograms for left-ventricular 3d sh
- echovdiff cardiac-cycle echocardiography video generation from arbitrary single
- ecosplat efficiency-controllable feed-forward 3d gaussian splatting from multi-v
- eda arbitrary noise diffusion design space | arXiv: 2507.18534
- Edge-Focused Super-Resolution for Omnidirectional Images with Spherical Geometric Augmentation
- edges compete for trust group relative edge optimization for building reconstruc
- edit-as-act goal-regressive planning for open-vocabulary 3d indoor scene editing | arXiv: 2603.17583
- EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editing
- editmgt unleashing potentials of masked generative transformers in image editing
- editprint general digital image forensics via editing fingerprint with self-augm
- edudiag a benchmark for educational diagnostic reasoning with error tracing and
- ee-rl vision language guided reinforcement learning with explorer and expert mod
- eegit teaching vision transformers to understand the eeg signal
- effecterase joint video object removal and insertion for high-quality effect era | arXiv: 2603.19224
- effectmaker unifying reasoning and generation for customized visual effect creat
- efficient all-pairs correlation volume sampling for optical flow estimation | arXiv: 2505.16942
- efficient and high-fidelity omni modality retrieval
- efficient and training-free single-image diffusion models | arXiv: 2606.04299
- efficient encoder-free fourier-based 3d large multimodal model
- efficient equivariant transformer for self-driving agent modeling | arXiv: 2604.01466
- efficient frame selection for long video understanding via reinforcement learnin
- efficient hybrid se3-equivariant visuomotor flow policy via spherical harmonics | arXiv: 2603.23227
- efficient real-time raw-to-raw denoising for extreme low-light ultra hd video on
- Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning
- efficient unrolled networks for large-scale 3d inverse problems
- efficient video object segmentation and tracking with recurrent dynamic submodel
- efficient weighted sampling via score-based generative models
- efficientmonohair fast strand-level reconstruction from monocular video via mult | 📄 paper_cache/CVPR2026/cvf-efficientmonohair_fast_strand-level_reco.txt
- efficientvpr toward efficient visual place recognition via scene-aware prompt tu
- ego embedding-guided personalization of vision-language models
- ego-1k -- a large-scale multiview video dataset for egocentric vision | arXiv: 2603.13741
- ego2web a web agent benchmark grounded in egocentric videos | arXiv: 2603.22529
- egocontrol controllable egocentric video generation via 3d full-body poses
- egoedit dataset real-time streaming model and benchmark for egocentric video edi
- egoflow gradient-guided flow matching for egocentric 6dof object motion generati | arXiv: 2604.01421
- egomind activating spatial cognition through linguistic reasoning in mllms | arXiv: 2604.03318
- egoposeformer v2 accurate egocentric human motion estimation for arvr | arXiv: 2603.04090
- egoprox evaluating mllms on egocentric 3d proximity reasoning across a cognitive | arXiv: 2605.24456
- egoroc towards egocentric robotic control via task-agnostic visual alignment
- egox egocentric video generation from a single exocentric video
- egoxtreme a dataset for robust object pose estimation in egocentric views under | arXiv: 2603.25135
- ei-partexplode for completion and implode for refinement
- elastic3d controllable stereo video conversion with guided latent decoding
- elasticformer detecting objects in hrw shots via elastic computing vision transf
- elic efficient lidar geometry compression via cross-bit-depth feature propagatio | arXiv: 2511.14070
- eliciting complex spatial reasoning in mllms through wide-baseline matching | arXiv: 2606.03577
- eliminate distance differences induced by backdoor attacks layer-selective train
- elv-halluc benchmarking semantic aggregation hallucinations in video understandi
- elvis enhance low-light for video instance segmentation in the dark | arXiv: 2512.01495
- emad evidence-centric grounded multimodal diagnosis for alzheimers disease | arXiv: 2602.19178
- embodiedsplat online feed-forward semantic 3dgs for open-vocabulary 3d scene und | arXiv: 2603.04254
- embodmocap in-the-wild 4d human-scene reconstruction for embodied agents
- emergent outlier view rejection in visual geometry grounded transformers
- emf meanflow text to image | arXiv: 2604.18168
- emgauss continuous slice-to-3d reconstruction via dynamic gaussian modeling in v | arXiv: 2512.06684
- emma concept erasure benchmark with comprehensive semantic metrics and diverse c | arXiv: 2512.17320
- emo-r3 reflective reinforcement learning for emotional reasoning in multimodal l | arXiv: 2602.23802
- emotag emotion-aware talking head synthesis on gaussian splatting with few-shot | arXiv: 2603.21332
- emothinker advancing visual-acoustic emotion analysis via structural token selec
- emr-diff edge-aware multimodal residual diffusion model for hyperspectral image
- enabling supervised learning of generative signatures for generalized ai-generat
- enc-bench a benchmark for evaluating multimodal large language models in electro | arXiv: 2603.22763
- end-to-end hyper-relational information extraction for engineering diagrams via
- endless world real-time 3d-aware long video generation
- energy waveify and redistribution for test-time adaptation a control system pers
- energy-gs image energy-guided pose alignment gaussian splatting with redesigned
- enhance-then-balance modality collaboration for robust multimodal sentiment anal
- enhancing accuracy of uncertainty estimation in appearance-based gaze tracking w | arXiv: 2501.14894
- enhancing continual learning of vision-language models via dynamic prefix weight | arXiv: 2604.18075
- enhancing descriptive captions with visual attributes for multimodal perception
- enhancing hands in 3d whole-body pose estimation with conditional hands modulato | arXiv: 2603.14726
- enhancing mixture of experts specialization via cluster aware upcycling | arXiv: 2604.13508
- enhancing out-of-distribution detection with extended logit normalization | arXiv: 2504.11434
- enhancing part-level point grounding for any open-source mllms
- enhancing spatial understanding in image generation via reward modeling | arXiv: 2602.24233
- enhancing the security of visual speaker authentication based on dynamic lip-pri
- enhancing unregistered hyperspectral image super-resolution via unmixing-based a
- enhancing visual representation with textual semantics textual semantics powered p | arXiv: 2503.13543
- envision attend then respond counterfactual hallucination mitigation in large vi
- epiagent agent centric system for ancient inscription restoration | arXiv: 2604.09367
- erasing thousands of concepts towards scalable and practical concept erasure for
- erecu pseudolabel evolution unsupervised camouflage | arXiv: 2603.11521
- eretinexgs retinex modeling for low-light scene enhancement via event streams an
- ermoe eigen-reparameterized mixture-of-experts for stable routing | arXiv: 2511.10971
- esam efficient online 3d perception on the edge
- ethoclip ontology-enhanced video-language pretraining for animal behavior unders
- eulerian gaussian splatting using hashed probability pyramids | arXiv: 2605.29136
- ev-cgnet co-visible focused 3d-guided 2d event keypoint detection network
- evaluating generative models via one-dimensional code distributions
- evatok adaptive length video tokenization for eff | arXiv: 2603.12267
- event structural valley a unified theoretical and practical framework for event
- event-based motion deblurring using task-oriented 3d gaussian event representati
- event-based motion deblurring with unpaired data
- event-illumination collaborative low-light image enhancement with a high-resolut
- event6d event-based novel object 6d pose tracking | arXiv: 2603.28045
- eventgait towards robust gait recognition with event streams
- eventhub data factory for generalizable event-based stereo networks without acti | arXiv: 2604.02331
- every error has its magnitude asymmetric mistake severity training for multiclas | arXiv: 2603.13682
- evidential deep partial label learning to quantify disambiguation uncertainty
- evidential neural radiance fields
- evidential transformation network post hoc uncertainty estimation | arXiv: 2604.08627
- evlf early vision-language fusion for generative dataset distillation | arXiv: 2603.07476
- evo-1 lightweight vision-language-action model with preserved semantic alignment
- evo-retriever llm-guided curriculum evolution with viewpoint-pathway collaborati
- evobj learning evolving object-centric representations for 3d instance segmentat | arXiv: 2605.13152
- evocomp learning visual token compression for multimodal large language models v | arXiv: 2604.17087
- evograph-r1 self-evolving multimodal knowledge hypergraphs for agentic retrieval
- evoid reinforced evolution for identity-preserving video generation
- evolutionary multimodal reasoning via hierarchical semantic representation for i | arXiv: 2603.03827
- evolving contextual safety in multi-modal large language models via inference-ti | arXiv: 2603.15800
- ew-detr evolving world object detection via incremental low-rank detection trans | arXiv: 2602.20985
- exact-gs mathematically rigorous and accurate 3d gaussian splatting for 3d x-ray
- exemplar-free class incremental learning via preserving class-discriminative str
- exmesh explicit mesh reconstruction with topology adaptation | arXiv: 2606.07288
- exotic external vision-driven incomplete multi-view classification
- Expand and Prune: Maximizing Trajectory Diversity for Effective GRPO in Generative Models
- expanding mmwave datasets for human pose estimation with unlabeled data and lida | arXiv: 2603.14507
- Expanding Spatial and Temporal Context for Robotic Imitation Learning With Scene Graphs
- experience transfer for multimodal llm agents in minecraft game
- expert-teacher-student collaborative learning for domain adaptive object detecti
- explaining object detectors via collective contribution of pixels | arXiv: 2412.00666
- explore with long-term memory a benchmark and multimodal llm-based reinforcement | arXiv: 2601.10744
- exploring 6d object pose estimation with deformation | arXiv: 2604.06720
- exploring adaptive masked reconstruction for self-supervised skeleton-based acti
- exploring conditions for diffusion models in robotic control | arXiv: 2510.15510
- exploring spatial intelligence from a generative perspective | arXiv: 2604.20570
- exploring spatiotemporal feature propagation for video-level compressive spectra | arXiv: 2603.00611
- exploring the underwater world segmentation without extra training
- expocm exposure-aware one-step generative single-image hdr reconstruction
- expose reinforcing video generation models for extreme pose estimation
- exposing and evaluating hallucinations for gui grounding
- exposing functional fusion a new class of strategic backdoor in dynamic prompt a
- expportrait expressive portrait generation via personalized representation | arXiv: 2602.19900
- extend3d town-scale 3d generation | arXiv: 2603.29387
- extending embodied question answering from perception to decision
- extrinsplat decoupling geometry and semantics for open-vocabulary understanding | arXiv: 2509.22225
- f2-assist multi-phase fetal growth forecast and report generation from ultrasoun
- f2hdr two-stage hdr video reconstruction via flow adapter and physical motion mo | arXiv: 2603.14920
- f2net a frequency-fused network for ultra-high resolution remote sensing segment
- faar efficient frequency-aware multi-task fine-tuning via automatic rank selecti | arXiv: 2603.20403
- face-guided sentiment boundary enhancement for weakly-supervised temporal sentim
- face2scene using facial degradation as an oracle for diffusion-based scene resto | arXiv: 2603.16570
- FACE: A Face-based Autoregressive Representation for High-Fidelity and Efficient Mesh Generation
- facecam portrait video camera control via scale-aware conditioning | arXiv: 2603.05506
- factorize reconstruct enhance a unified framework for multimodal sentiment analy
- factorized context aggregation for robust cancer risk estimation via soft re-ran
- failure modes for deep learning-based online mapping how to measure and address | arXiv: 2603.19852
- failureatlas mapping the failure landscape of t2i models via active exploration
- fairllava fairness-aware parameter-efficient fine-tuning for large vision-langua | arXiv: 2603.26008
- Faithful Contouring: Near-Lossless 3D Voxel Representation Free from Iso-surface
- faithfusion harmonizing reconstruction and generation via pixel-wise information
- falcon false-negative aware learning of contrastive negatives in vision-language | arXiv: 2505.11192
- fantasyvln unified multimodal chain-of-thought reasoning for vision-and-language
- fape-ir frequency-aware planning and execution framework for all-in-one image re
- fast scenescript fast and accurate language-based 3d scene understanding via mul | arXiv: 2512.05597
- fast spatial tracking with visual geometry transformer
- fast-foundationstereo real-time zero-shot stereo matching
- fast-thinkact efficient vision-language-action reasoning via verbalizable latent | arXiv: 2601.09708
- fasteventdgs deformable gaussian splatting for fast dynamic scenes from a single
- fastgs training 3d gaussian splatting in 100 seconds | arXiv: 2511.04283
- fasthybrid accelerating hybrid autoregressive image generation with lookahead an
- fastlightgen fast and light video generation with fewer steps and parameters | arXiv: 2603.01685
- fastref fast prototype refinement for few-shot industrial anomaly detection
- fave a structured benchmark for fine-grained audio-visual temporal evaluation in
- fb-clip fine-grained zero-shot anomaly detection with foreground-background dise
- fbta enabling single-gpu end-to-end gigapixel wsi classification with feature br
- feast fully connected expressive attention for spatial transcriptomics
- feat federated geometry aware correction for exemplar replay under continual dynamic heterogeneity | arXiv: 2604.08617
- featurising pixels from dynamic 3d scenes with linear in-context learners | arXiv: 2604.26488
- fed-ade adaptive learning rate for federated post-adaptation under distribution | arXiv: 2603.01040
- fedadamom adaptive momentum for improved generalization in federated optimizatio
- fedafd multimodal federated learning via adversarial fusion and distillation | arXiv: 2603.04890
- fedalign differentially private distribution alignment for non-iid federated lea
- fedara resource-adaptive low-rank personalized federated learning via anchor-dri
- fedbprompt federated domain generalization person re-identification via body dis | arXiv: 2603.12912
- fedcart tackling long-tailed distributions in federated adversarial training via
- feddap domain-aware prototype learning for federated learning under domain shift | arXiv: 2604.06795
- federated active learning extreme noniid | arXiv: 2603.10341
- fedharmony harmonizing heterogeneous label correlations in federated multi-label | arXiv: 2604.28024
- fedmop achieving enhanced privacy and performance in federated learning via mome
- fedmpt federated multi-label prompt tuning of vision-language models
- fedrac rolling submodel allocation for collaborative fairness in federated learn
- fedre a representation entanglement framework for model-heterogeneous federated | arXiv: 2511.22265
- fedrg unleashing the representation geometry for federated learning with noisy c
- fedsdr federated graph learning with structural noise detection and reconstructi
- feed-forward one-shot animatable textured mesh avatar reconstruction
- few-for-many personalized federated learning
- few-shot acoustic synthesis with multimodal flow matching | arXiv: 2603.19176
- few-shot hybrid incremental learningcontinually learning under data scarcity and
- few-shot incremental 3d object detection in dynamic indoor environments | arXiv: 2604.07997
- few-step diffusion sampling through instance-aware discretizations
- ffp-300k scaling first-frame propagation for generalizable video editing
- fg-portrait 3d flow guided editable portrait animation | arXiv: 2603.23381
- fidesr high-fidelity and detail-preserving one-step diffusion super-resolution | arXiv: 2603.02692
- fighting hallucinations with counterfactuals diffusion-guided perturbations for | arXiv: 2603.10470
- filtergs traversal-free parallel filtering and adaptive shrinking for large-scal
- fine factorizing knowledge for initialization of variable-sized diffusion models
- fine-grained image aesthetic assessment learning discriminative scores from rela | arXiv: 2603.03907
- fine-grained multi image object hallucination benchmark
- fine-grained post-training quantization for large vision language models with qu | arXiv: 2603.17809
- fine-tuning impairs the balancedness of foundation models in long-tailed persona | arXiv: 2605.02247
- fine-vad towards fine-grained video anomaly detection via progressive cross-gran
- finer mllms hallucinate under fine-grained negative queries | arXiv: 2603.17662
- finpercep rm a fine grained reward model and co evolutionary curriculum for rl ba | arXiv: 2512.22647
- first logit boosting visual grounding method to mitigate object hallucination in
- fisherposer human motion estimation from sparse observations with hierarchical r
- fishuman fine-grained single-image 3d human reconstruction via multi-view 4d rem
- fixed anchors are not enough dynamic retrieval and persistent homology for datas | arXiv: 2602.24144
- flare a failure-aware framework for autonomous correction and recovery in visual
- flash-dmd towards high-fidelity few-step image generation with efficient distill
- flashcap millisecond-accurate human motion capture via flashing leds and event-b | arXiv: 2603.19770
- FlashIn: Fast and Accurate Image Inversion for Real-time Image Editing
- flashlips 100-fps mask-free latent lip-sync using reconstruction instead of diff
- flashmesh faster and better autoregressive mesh synthesis via structured specula
- flashmotion fewstep controllable video generation | arXiv: 2603.12146
- flashportrait 6x faster infinite portrait animation with adaptive latent predict
- flashvggt efficient and scalable visual geometry transformers with compressed descr | arXiv: 2512.01540
- flat-pack bench evaluating spatio-temporal understanding in large vision-languag | arXiv: 2605.21625
- flexavatar flexible large reconstruction model for animatable gaussian head avat
- flexavatar learning complete 3d head avatars with partial supervision | arXiv: 2512.15599
- flexivideo variation-aware temporal dynamics modeling for efficient video unders
- flextraj image-to-video generation with flexible point trajectory control
- flow matching for multimodal distributions
- flow optimal transport-driven feature warping for generalized remote physiologic
- flow3r factored flow prediction for scalable visual geometry learning | arXiv: 2602.20157
- flow4dgs-slam optical flow-guided 4d gaussian splatting slam
- Flowception: Temporally Expansive Flow Matching for Video Generation
- flowcomposer composable flows for compositional zeroshot learning | arXiv: 2603.16641
- flowdc flow-based decoupling-decay for complex image editing
- flowdirector training-free flow steering for precise text-to-video editing
- flowhijack a dynamics-aware backdoor attack on flow-matching vision-language-act
- flowhijack dynamics aware backdoor attack on flow matching vla models | arXiv: 2604.09651
- flowmotion training-free flow guidance for video motion transfer | arXiv: 2603.06289
- flowpalm optical flow driven non-rigid deformation for geometrically diverse pal
- flowportal residual-corrected flow for training-free video relighting and backgr
- flowsteer guiding few-step image synthesis with authentic trajectories
- fluidgaussian propagating simulation-based uncertainty toward functionally-intel | arXiv: 2603.21356
- fluoclip stain-aware focus quality assessment in fluorescence microscopy | arXiv: 2602.23791
- fluxmem adaptive hierarchical memory for streaming video understanding | arXiv: 2603.02096
- fmri-lm towards a universal foundation model for language-aligned fmri understan
- focal-general diffusion model with semantic consistent guidance for sign languag
- focus dont prune identifying instruction-relevant regions for information-rich i | arXiv: 2603.22815
- focus on background exploring sams potential in few-shot medical image segmentat
- focus-to-perceive representation learning a cognition-inspired hierarchical fram | arXiv: 2603.25778
- focusui efficient ui grounding via position-preserving visual token selection
- foleydesigner immersive stereo foley generation with precise spatio-temporal ali
- foleydirector fine-grained temporal steering for video-to-audio generation via s
- follow the saliency supervised saliency for retrieval-augmented dense video capt | arXiv: 2603.11460
- fontcrafter high-fidelity element-driven artistic font creation with visual in-c | arXiv: 2603.22054
- force transferable visual jailbreaking attacks via feature over reliance correct | arXiv: 2509.21029
- forcevla2 unleashing hybrid force-position control with force awareness for cont | arXiv: 2603.15169
- forecast the principal stabilize the residual subspace-aware feature caching for
- forecasting 3d scanpaths in egocentric video
- forehoi feed-forward 3d object reconstruction from daily hand-object interaction
- forge continual learning for fmri based brain disorder diagnosis | arXiv: 2604.14259
- forging a dynamic memory retrieval-guided continual learning for generalist medi
- foss modeling long range dependencies and multimodal uncertainty in trajectory p | arXiv: 2603.01284
- foundation model priors enhance object focus in feature space for source-free ob | arXiv: 2512.17514
- foundir-v2 optimizing pre-training data mixtures for image restoration foundatio
- foundry distilling 3d foundation models for the edge | arXiv: 2511.20721
- fourier angle alignment for oriented object detection in remote sensing | arXiv: 2602.23790
- fov-net rotation-invariant cad b-rep learning via field-of-view ray casting | arXiv: 2602.24084
- fozo forward-only zeroth-order prompt optimization for test-time adaptation | arXiv: 2603.04733
- fps-bench a benchmark for high frame-rate video understanding
- fractal camouflage a bio-inspired approach for multi-scale adversarial attacks i
- frame2freq spectral adapters for fine-grained video understanding | arXiv: 2602.18977
- framer frequency-aligned self-distillation with adaptive modulation leveraging d | arXiv: 2512.01390
- franca nested matryoshka clustering for scalable visual representation learning
- frankenmotion part-level human motion generation and composition
- free-grained hierarchical visual recognition | arXiv: 2510.14737
- free-lunch long video generation via layer-adaptive ood correction | arXiv: 2603.25209
- freeartgs articulated gaussian splatting under free-moving scenario | arXiv: 2603.22102
- freeform reduced-order deformable simulation from particle-based skinning eigenm | arXiv: 2605.29318
- freescale scaling 3d scenes | arXiv: 2604.10512
- freqedit preserving high-frequency features for robust multi-turn image editing
- freqflow frequency aware flow matching | arXiv: 2604.15521
- freqsic frequency-aware stereo image compression with bi-directional checkerboar
- frequency switching mechanism for parameter-ecient multi-task learning | arXiv: 2603.21111
- frequency-aware affinity for weakly supervised semantic segmentation
- fresco frequency-spatial consistent optimization for fine-grained head avatar mo
- from 2d alignment to 3d plausibility unifying heterogeneous 2d priors and penetr | arXiv: 2503.17788
- from 3d pose to prose biomechanics-grounded vision-language coaching
- from attraction to equilibrium physics-inspired semantic gravitons for zero-shot
- from contrast to consistency rethinking event-based continuous-time optical flow | arXiv: 2605.25570
- from corners to fiducial tags revisiting checkerboard calibration for event came
- from detection to association learning discriminative object embeddings for mult
- from editor to dense geometry estimator | arXiv: 2509.04338
- from exploration to exploitation a two-stage entropy rlvr approach for noise-tol
- from failure to feedback group revision unlocks hard cases in object-level groun
- from feature learning to spectral basis learning a unifying and flexible framewo
- from indoor to open world revealing the spatial reasoning gap in mllms
- from infusion to assimilation distillation for medical image segmentation
- from inpainting to layer decomposition repurposing generative inpainting models | arXiv: 2511.20996
- from intuition to investigation a tool-augmented reasoning mllm framework for ge | arXiv: 2603.01038
- from manuals to actions a unified vla model for chain-of-thought manual generati
- from measurement to mitigation quantifying and reducing identity leakage in imag
- from none to all self-supervised 3d reconstruction via novel view synthesis
- from observation to action latent action-based primitive segmentation for vla pr | arXiv: 2511.21428
- from pairs to sequences track-aware policy gradients for keypoint detection | arXiv: 2602.20630
- from panel to pixel zoom-in vision-language pretraining from biomedical scientif
- from pixel to precision enhancing handwritten mathematical expression recognitio
- from rays to projections better inputs for feed-forward view synthesis
- from selection to scheduling federated geometry-aware correction makes exemplar
- from sketch to fresco efficient diffusion transformer with progressive resolutio
- from softmax to dirichlet evidential learning for semi-supervised semantic segme
- from spots to pixels dense spatial gene expression prediction from histology ima
- from static to dynamic exploring self-supervised image-to-video representation t | arXiv: 2603.26597
- from weights to concepts data-free interpretability of clip via singular vector | arXiv: 2603.24653
- from where things are to what they are for benchmarking spatial-functional intel | arXiv: 2605.02130
- fsfsplatter geometrically accurate reconstruction with free sparse-view images w
- fslora harmonizing detection and re-identification via freq-spatial low-rank ada
- fuel gauge estimating chain-of-thought length ahead of time in large multimodal
- funfact building probabilistic functional 3d scene graphs via factor-graph reaso
- funrec reconstructing functional 3d scenes from egocentric interaction videos | arXiv: 2604.05621
- fusar-gpt a spatiotemporal feature-embedded and two-stage decoupled visual langu
- fuser feed-forward multiview 3d registration transformer and se3n diffusion refi | arXiv: 2512.09373
- fusion in your way aligning image fusion with heterogeneous demands via direct p | arXiv: 2605.06049
- fusionagent a multimodal agent with dynamic model selection for human recognitio | arXiv: 2603.26908
- fvar next-focus prediction for visual autoregressive modeling
- fvbench benchmarking deepfake video detection capability of large multimodal mod
- g mixer geodesic mixup based implicit semantic expansion for zero shot cir | arXiv: 2604.14710
- g2vlm geometry grounded vision language model with unified 3d reconstruction and
- ga-vln geometry-aware bev representation for efficient vision-language navigatio
- gallant voxel grid-based humanoid locomotion and local-navigation across 3-d con
- gamba mamba-based graph convolutional network with dynamic graph topology learni
- gardendesigner encoding aesthetic principles into jiangnan garden construction v | arXiv: 2604.01777
- garments2look a multi-reference dataset for high-fidelity outfit-level virtual t | arXiv: 2603.14153
- gastric-x a multimodal multi-phase benchmark dataset for advancing vision-langua
- gated condition injection without multimodal attention towards controllable line
- gau-occ geometry-completed gaussians for multi-modal 3d occupancy prediction
- gaussfusion improving 3d reconstruction in the wild with a geometry-informed vid | arXiv: 2603.25053
- gaussian splatting-based low-rank tensor representation for multi-dimensional im
- gaussiandwm 3d gaussian driving world model for unified scene understanding and | arXiv: 2512.23180
- gaussianfluent gaussian simulation for dynamic scenes with mixed materials
- gaussiangrow geometry-aware gaussian growing from 3d point clouds with text guid | arXiv: 2604.05721
- gaussianmatch semi-supervised regression with pseudo-label filtering via multi-v
- gaussianpile a unified sparse gaussian splatting framework for slice-based volum | arXiv: 2603.20611
- gaussianvision vision-language alignment from compressed image representations u
- gaussianzoom progressive zoom-in generative 3d gaussian splatting with geometric
- gaze target estimation anywhere with concepts
- gazeonce360 fisheye-based 360 multi-person gaze estimation with global-local fea | arXiv: 2603.17161
- gazeonce360 fisheye-based 360deg multi-person gaze estimation with global-local
- gazeshift unsupervised gaze estimation and dataset for vr
- gdpo-sr group direct preference optimization for one-step generative image super
- gdro group-level reward post-training suitable for diffusion models
- geco geometry-consistent regularization for domain generalized semantic segmenta
- geco-srt geometry-aware continual adaptation for cross-task sim-to-real transfer
- gecosrt geometryaware continual adaptation for rob | arXiv: 2602.20871
- gem generating lidar world model via deformable mamba
- gem-tfl bridging weak and full supervision for forgery localization through em-g | arXiv: 2603.05095
- gen3r 3d scene generation meets feed-forward reconstruction
- gencolorbench a color evaluation benchmark for text-to-image generation
- generalizable co-salient object detection via mixed content-style modulation
- generalizable radio-frequency radiance fields for spatial spectrum synthesis
- generalizable sparse-view 3d reconstruction from unconstrained images
- generalizable structure-aware keypoint correspondence for category-unified 3d si
- generalizable video quality assessment via weak-to-strong learning | arXiv: 2505.03631
- generalized and personalized federated learning with black-box foundation models
- generalized-cvo fast and correspondence-free local point cloud registration with
- generalizing visual geometry priors to sparse gaussian occupancy prediction | arXiv: 2602.21552
- GenErase: Generalizable and Semantically-Aware Concept Erasure in Diffusion Models
- generate analyze and refine training-free sound source localization via mllm met | arXiv: 2604.06824
- generative adversarial perturbations with cross-paradigm transferability on loca | arXiv: 2603.24821
- generative diffusion priors for 3d mapping of the dark universe | arXiv: 2606.00803
- generative modeling of weights generalization or memorization | 📄 paper_cache/CVPR2026/cvf-generative_modeling_of_weights_generaliz.txt
- generative neural video compression via video diffusion prior | arXiv: 2512.05016
- generative point tracking and forecasting
- generative video compression with one-dimensional latent representation | arXiv: 2603.15302
- genhoi towards object-consistent hand-object interaction with temporally balance
- geninav generative model driven image-goal navigation via imagination-guided con
- genmask adapting dit for segmentation via direct mask generation | arXiv: 2603.23906
- genmatter perceiving physical objects with generative matter models | arXiv: 2604.22160
- geoagent learning to geolocate everywhere with reinforced geographic characteris
- geobridge a semantic-anchored multi-view foundation model bridging images and te
- geobridge semantic-anchored multi-view foundation model for geo-localization | arXiv: 2512.02697
- geocot towards reliable remote sensing reasoning with manifold perspective
- geodesicnvs probability density geodesic flow matching for novel view synthesis | arXiv: 2603.01010
- geodexgrasp geometry-aware generation for data-efficient and physics-plausible d
- geoflow real-time fine-grained cross-view geolocalization | arXiv: 2603.21943
- geofree-coseg unsupervised point cloud-image cross-modal co-segmentation without
- geoguide hierarchical geometric guidance for open-vocabulary 3d semantic segment | arXiv: 2603.26260
- geoint-r1 formalizing multimodal geometric reasoning with dynamic auxiliary cons
- geometric neural distance fields for learning human motion priors
- geometric-aware hypergraph reasoning for novel class discovery in point cloud se
- geometric-photometric event-based 3d gaussian ray tracing
- geometry-as-context modulating explicit 3d in scene-consistent video generation | arXiv: 2602.21929
- geometry-aware cross-modal graph alignment for referring segmentation in 3d gaus
- geometry-driven ood detectors are class-incremental learners
- geommbench and geommagent toward expert level multimodal intelligence in geoscience and remote sensing | arXiv: 2604.08896
- geomotion rethinking motion segmentation via latent 4d geometry
- geopredict leveraging predictive kinematics and 3d gaussian geometry for precise
- georelight learning joint geometrical relighting and reconstruction with flexibl | arXiv: 2604.20715
- geork2 geometry-guided runge-kutta integration for diffusion transformer acceler
- geosemba reconstructing state space model for cross paradigm representation in m
- geosurge geo-localization using semantic fusion with hierarchy of geographic emb | arXiv: 2510.01448
- geotikzbridge advancing multimodal code generation for geometric perception and | arXiv: 2603.22687
- geovis geospatially rewarded visual search for remote sensing visual grounding
- geoworld geometric world models | arXiv: 2602.23058
- gfrrn explore the gaps in single image reflection removal
- ggbench a geometric generative reasoning benchmark for unified multimodal models
- gh-naf grid-adaptive hash-level-attended neural attenuation fields for discrepan
- ghost-fwl a large-scale full-waveform lidar dataset for ghost detection and remo | arXiv: 2603.28224
- ghpt real-time relightable gaussian splatting using hybrid path tracing
- gifsplat generative prior-guided iterative feed-forward 3d gaussian splatting fr
- gift global irreplaceability frame targeting for efficient video understanding
- gkd generalizable knowledge distillation vfm | arXiv: 2603.02554
- glint modeling scene-scale transparency via gaussian radiance transport | arXiv: 2603.26181
- global prior meets local consistency dual-memory augmented vision-language-actio | arXiv: 2602.20200
- global structure-from-motion meets feedforward reconstruction | arXiv: 2605.26103
- global underwater geolocation from time-lapse polarization imagery
- global-aware edge prioritization for pose graph initialization | arXiv: 2602.21963
- global-graph guided and local-graph weighted contrastive learning for unified cl
- glove2hand synthesizing natural hand-object interaction from multi-modal sensing | arXiv: 2603.20850
- glyphprinter region-grouped direct preference optimization for glyph-accurate vi | arXiv: 2603.15616
- gm-r2 generative matching learning for unsupervised geometric representation and
- gmt effective global framework for multi-camera multi-target tracking
- goal force teaching video models to accomplish physics-conditioned goals | arXiv: 2601.05848
- good can sometimes be bad a unified attack against 3d point cloud classifier by
- GOR-IS: 3D Gaussian Object Removal In the Intrinsic Space
- gp-4dgs probabilistic 4d gaussian splatting from monocular video via variational | arXiv: 2604.02915
- gpflow gaussian prototype probability flow for unsupervised multi-modal anomaly
- gqir generative quanta image reconstruc tion | arXiv: 2602.20417
- gr-gauge cost-efficient training configuration by gauging the gradient redundanc
- gradient knows best mixed-precision quantization via gradient-guided bit allocat
- granulon awakening pixel-level visual encoders with adaptive multi-granularity s
- graph-to-frame rag visual-space knowledge fusion for training-free and auditable | arXiv: 2604.04372
- graph2eval automatic multimodal task generation for agents via knowledge graphs | arXiv: 2510.00507
- graphformer a multimodal graph persistent homology transformer for the analysis
- graphvlm benchmark vlm graph learning | arXiv: 2603.13370
- graspall adaptive structural compensation from illumination variation for roboti
- graspgen-x cross-embodiment 6-dof diffusion-based grasping
- graspldp towards generalizable grasping policy via latent diffusion | arXiv: 2602.22862
- grid distillation compositional image distillation via structured generative gri
- groce graph-guided online concept erasure for text-to-image diffusion models | arXiv: 2511.12968
- ground reaction inertial poser physics-based human motion capture from sparse im
- grounded 3d-aware spatial vision-language modeling | arXiv: 2605.30307
- grounding everything in tokens for multimodal large language models
- groundingme exposing the visual grounding gap in mllms through multi-dimensional
- groundvts visual token sampling in multimodal large language models for video te | arXiv: 2604.02093
- group diffusion enhancing image generation by unlocking cross-sample collaborati
- group editing edit multiple images in one go | arXiv: 2603.22883
- grpo-guard mitigating implicit over-optimization in flow matching via regulated
- gs-asm 2dgs-supervised active stereo matching
- gs-clip zero-shot 3d anomaly detection by geometry-aware prompt and synergistic | arXiv: 2602.19206
- gs2 graph-based spatial distribution optimization for compact 3d gaussian splatt
- gsnr graph smooth null space representation for inverse problems | arXiv: 2602.20328
- gsv2x geometry-aware uncertainty modeling and orthogonal fusion for robust roads
- gt-svj generative-transformer-based self-supervised video judge
- gthinker towards general multimodal reasoning via cue-guided rethinking
- gtr turbo merged checkpoint free teacher | arXiv: 2512.13043
- guardians of the hair rescuing soft boundaries in depth stereo and novel views
- guardtrace-vl detecting unsafe multimodel reasoning via iterative safety supervi
- gui-ceval a hierarchical and comprehensive chinese benchmark for mobile gui agen | arXiv: 2603.15039
- guide a benchmark for understanding and assisting users in open-ended gui tasks | arXiv: 2603.25864
- guideflow constraint-guided flow matching for planning in end-to-end autonomous
- guiding a diffusion model by swapping its tokens | arXiv: 2604.08048
- guiding a diffusion transformer with the internal dynamics of itself | arXiv: 2512.24176
- guiding diffusion models with fine-grained conditions and semantics-preserving s
- guiding diffusion models with semantically degraded conditions | arXiv: 2603.10780
- guiding diffusion-based reconstruction with contrastive signals for balanced vis
- guiding token-sparse diffusion models
- gyro-based deep video deblurring
- h-sets hessian-guided discovery of set-level feature interactions in image class | arXiv: 2604.22045
- h2-surv hierarchical hyperbolic multimodal representation learning for survival
- h2a2 homogeneity-aware and heterogeneity-aware feature perception for unified in
- hallugen synthesizing realistic and controllable hallucinations for evaluating i
- hamipose hamiltonian optimization for unsupervised domain adaptive pose estimati
- hammer harnessing mllm via cross-modal integration for intention-driven 3d affor | arXiv: 2603.02329
- hammer harnessing mllms via cross-modal integration for intention-driven 3d affo
- handdreamer zero shot text to 3d hand model generation | arXiv: 2604.04425
- handdreamer zero-shot text to 3d hand model generation using corrective hand sha
- handvqa diagnosing and improving fine-grained spatial reasoning about hands in v | arXiv: 2603.26362
- handx scaling bimanual motion and interaction generation | arXiv: 2603.28766
- Haptic Neural Fields: Bringing Tactile Interactions to 3D Rendered Scenes
- harmonic canvas inversion-free editing for visually-guided music style transfer
- harmonious parameter adaptation in continual visual instruction tuning for safet
- harmonious parameter adaptation in continual visual instruction tuning for safet
- harmonized feature conditioning and frequency-prompt personalization for multi-r | arXiv: 2605.08210
- harnessing the power of foundation models for accurate material classification
- hats hardness-aware trajectory synthesis for gui agents | arXiv: 2603.12138
- haven hierarchical long video understanding with audiovisual entity cohesion | arXiv: 2601.13719
- hawk head importance-aware visual token pruning in multimodal models | arXiv: 2604.07812
- hbridge h-shape bridging of heterogeneous experts for unified multimodal underst
- hcl-ff hierarchical and contrastive learning for forward-forward algorithm | arXiv: 2605.24797
- hdr-vlm hdr-domain adaptation of vlms and preference-aligned quality assessment
- Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
- hear what you see video-to-audio generation with diffusion transformer and seman
- hear you are teaching llms spatial reasoning with vision and spatial sound
- hearing the room through the shape of the drum modal-guided sound recovery from
- herbench a benchmark for multi-evidence integration in video question answering | arXiv: 2512.14870
- hermite radial basis function for surface reconstruction via differentiable rend
- hero hierarchical embedding-refinement for open-vocabulary temporal sentence gro
- herod heuristic inspired reasoning data efficient rod | arXiv: 2603.24166
- herogs hierarchical guidance for robust 3d gaussian splatting under sparse views
- hess head sensitivity score for sparsity redistribution in vggt | arXiv: 2603.25336
- heterogeneous decentralized diffusion models | arXiv: 2603.06741
- heuristic self-paced learning for domain adaptive semantic segmentation under ad | arXiv: 2603.24322
- heuristic-inspired reasoning priors facilitate data-efficient referring object d
- hfedatm hierarchical federated domain generalization via optimal transport and r
- hfr and hdr video from multi-attenuated spikes using a rapidly rotating spokend
- hg-i2p bridging modalities for generalizable image-to-point-cloud registration v | arXiv: 2603.27969
- hg-lane high-fidelity generation of lane scenes under adverse weather and lighti | arXiv: 2603.10128
- hi-lo prune look at what youll lose before pruning with hierarchical token selec
- hicogen hierarchical compositional text-to-image generation in diffusion models
- hidden dangers of compositional generation diagnosing semantic safety failures i
- hidra hierarchical degradation representation and adaptation with generative pri
- hier-cos making deep features hierarchy-aware via composition of orthogonal subs | arXiv: 2503.07853
- hieramamba video temporal grounding via hierarchical anchor-mamba pooling | arXiv: 2510.23043
- hieramp coarse-to-fine autoregressive amplification for generative dataset disti | arXiv: 2603.06932
- hierarchical codec diffusion for video-to-speech generation | arXiv: 2604.15923
- hierarchical concept embedding pursuit for interpretable image classification
- hierarchical enhancement of semantic priors for disentangled text-driven motion
- hierarchical point-patch fusion with adaptive patch codebook for 3d shape anomal
- hierarchical process reward models are symbolic vision learners
- hierarchical visual relocalization with nearest view synthesis from feature gaus | arXiv: 2603.29185
- hieruq hierarchical uncertainty quantification with adaptive granularity reconci
- hif-vla hindsight insight and foresight through motion representation for vision | arXiv: 2512.09928
- hifi-inpaint towards high-fidelity reference-based inpainting for generating det | arXiv: 2603.02210
- hificl highfidelity incontext learning for multimo | arXiv: 2603.12760
- high-fidelity diffusion face swapping with id-constrained facial conditioning | arXiv: 2503.22179
- high-fidelity virtual try-on beyond paired data scarcity via diffusion-based cyc
- high-precision dichotomous image segmentation via depth integrity-prior and fine
- high-quality and efficient turbulence mitigation with events | arXiv: 2603.20708
- hilbert curve-based attention enabling topology-preserving image tensor represen
- hispatial taming hierarchical 3d spatial understanding in vision-language models | arXiv: 2603.25411
- history to future evolving agent with experience and thought for zero-shot visio
- hog layout hierarchical 3d scene generation optimization and editing | arXiv: 2604.10772
- holo homography-guided pose estimator network for fine-grained visual localizati
- honeybee data recipes for vision-language reasoners | arXiv: 2510.12225
- hops hierarchical open-vocabulary part segmentation with attention-aware filteri
- horizonforge driving scene editing with any trajectories and any vehicles | arXiv: 2602.21333
- housemind tokenization mllm floor plan | arXiv: 2603.11640
- how far can we go with synthetic data for audio-visual sound source localization
- how much 3d do video foundation models encode
- how to take a memorable picture empowering users with actionable feedback | arXiv: 2602.21877
- hp-edit a human-preference post-training framework for image editing
- hsi-gpt2 a dual-granularity large motion reasoning model with diffusion refineme
- htnav a hybrid navigation framework with tiered structure for urban aerial visio
- hugging visual prompt and segmentation tokens consistency learning for fine-grai
- hulluedit single-pass evidence-consistent subspace editing for mitigating halluc | arXiv: 2602.22727
- hulluedit subspace editing hallucination | arXiv: 2602.22727
- human interaction-aware 3d reconstruction from a single image | arXiv: 2604.05436
- human-centric multi-exposure fusion benchmark and bi-level cognition distillatio
- human-like abstract visual reasoning via understanding and solving reasoning loo
- humanba human-aware bundle adjustment via global human-camera decoupling
- humannova photorealistic universal and rapid 3d human avatar modeling from a sin | arXiv: 2606.02573
- humanvbench probing human centric video understanding in mllms with automatica | arXiv: 2412.17574
- humaps-4d a multimodal dataset for human motion analysis with physiological and
- humorchain theory-guided multi-stage reasoning for interpretable multimodal humo
- hunting normality from query sample via residual learning for generalist anomaly
- hvg-3d bridging real and simulation domains for 3d-conditional hand-object inter
- hybrid robust collaborative perception with lidar-4d radar fusion under adverse
- hybriddrivevla vision-language-action model with visual cot reasoning
- hycal training free prototype calibration for cross discipline fscil | arXiv: 2604.15678
- hyper-pcn hypergraph-based point cloud completion via high-order correlation mod
- hyperbolic busemann neural networks | arXiv: 2602.18858
- hyperbolic defect feature synthesis for few-shot defect classification
- hyperbolic gramian volumes for multimodal alignment
- Hyperbolic Prototype Learning with Uncertainty-Aware Consistency for Continual Test-Time Segmentation
- hypergait unleashing the power of parsing for gait recognition in the wild via h
- hypergaussians high-dimensional gaussian splatting for high-fidelity animatable | arXiv: 2507.02803
- hypergraph-state collaborative reasoning for multi-object tracking
- hypernas enhancing architecture representation for nas predictor via hypernetwor
- hyperst hierarchical hyperbolic learning for spatial transcriptomics prediction
- hypevpr exploring hyperbolic space for perspective to equirectangular visual pla | arXiv: 2506.04764
- i2i-bench a comprehensive benchmark suite for image-to-image editing models | 📄 paper_cache/CVPR2026/cvf-i2i-bench_a_comprehensive_benchmark_suit.txt
- iafmnet information-aware feature modulation for efficient super-resolution
- ibisagent reinforcing pixel-level visual reasoning in mllms for universal biomed
- ictpolarreal a polarized reflection and material dataset of real world objects | arXiv: 2603.24912
- ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
- id-sim an identity-focused similarity metric
- idesplat iterative depth probability estimation for generalizable 3d gaussian sp
- idperturb enhancing variation in synthetic face generation via angular perturbat | arXiv: 2602.18831
- iebglan interpretability-enhanced brain graph learning framework with llm-instru
- if-bench benchmarking and enhancing mllms for infrared images with generative vi
- if-prune information-flow guided token pruning for efficient vision-language mod
- ifcsr inference-free fidelity-realism control for one-step diffusion-based real-
- illuminating visual identity in universal multimodal embeddings
- illumination-consistent human-scene reconstruction from monocular video
- illustrators depth monocular layer index prediction for image decomposition
- ilrm an iterative large 3d reconstruction model
- image diffusion preview with consistency solver | arXiv: 2512.13592
- image generation from contextually-contradictory prompts
- image guides images consistent video amodal completion with rectified in-context
- Image-Guided Geometric Stylization of 3D Meshes
- image-to-point cloud feature back-projection for multimodal training of 3d seman
- ImageRAGTurbo: Towards One-step Text-to-Image Generation with Retrieval-Augmented Diffusion Models
- imagine before concentration diffusion-guided registers enhance partially releva | arXiv: 2604.03653
- imaia interactive maps ai assistant for travel planning and geo-spatial intellig
- imbalanced view contribution evaluation and refinement for deep incomplete multi
- immeriris a large-scale dataset and benchmark for off-axis and unconstrained iri | arXiv: 2510.10113
- immunizing models against harmful long-horizon fine-tuning via contractive optim
- imontage unified versatile highly dynamic many-to-many image generation
- improved mean flows on the challenges of fastforward generative models
- improving calibration in test-time prompt tuning for vision-language models via | arXiv: 2604.27715
- improving controllable generation faster training and better performance via x0-
- improving motion in image-to-video models via adaptive low-pass guidance
- improving sparse autoencoder with dynamic attention
- ims3 breaking distributional aggregation in diffusion-based dataset distillation
- imu-hoi a symbiotic framework for coherent human-object interaction and motion c
- incentivizing generative zero-shot learning via outcome-reward reinforcement lea
- incentivizing versatile video reasoning in mllms via data-efficient reinforcemen
- increfa breaking the static wall of generative model attribution | arXiv: 2604.17736
- incremental object detection via future-aware decoupled cross-head distillation
- inference-time physics alignment of video generative models with latent world mo
- infinibench infinite benchmarking for visual spatial reasoning with customizable
- InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields
- infinity-rope action-controllable infinite video generation emerges from autoreg | arXiv: 2511.20649
- information-theoretic decomposition for multimodal interaction learning | arXiv: 2606.11614
- innoads-composer efficient condition composition for e-commerce poster generatio | arXiv: 2603.05898
- inscal calibrated multi-source fully test-time prompt tuning for object detectio
- insid3 training-free in-context segmentation with dinov3 | arXiv: 2603.28480
- inside-out measuring generalization in vision transformers through inner working | arXiv: 2604.08192
- insight bench towards grounded in-situ guidance for robotic manipulation
- instantretouch efficient and high-fidelity instruction-guided image retouching w
- InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding
- instruction-guided lesion segmentation for chest x-rays with automatically gener | arXiv: 2511.15186
- instructmix2mix consistent sparse-view editing through multi-view model personal
- inter-edit first benchmark for interactive instruction-based image editing
- interact2ar full-body human-human interaction generation via autoregressive diff
- interactive episodic memory with user feedback | arXiv: 2604.24893
- interactive tracking a human-in-the-loop paradigm with memory-augmented adaptati
- interndata-a1 pioneering high-fidelity synthetic data for pre-training generalis
- interphys physics-aware human motion synthesis in a dynamic scene
- interpretable and steerable concept bottleneck sparse autoencoders | arXiv: 2512.10805
- interpretable cross-domain few-shot learning with rectified target-domain local | arXiv: 2603.17655
- interpretable motion-attentive maps spatio-temporally localizing concepts in vid | arXiv: 2603.02919
- interpretable prompts made edit-friendly token-to-token similarity reduction in
- intervention-aware multiscale representation learning from imaging phenomics and | arXiv: 2604.22832
- intra-class distribution-guided generative hashing with neighbor refinement for
- intrinsic concept extraction based on compositional interpretability | arXiv: 2603.11795
- intrinsic geometry-appearance consistency optimization for sparse-view gaussian
- intrinsicweather controllable weather editing in intrinsic space | arXiv: 2508.06982
- introsvg learning from rendering feedback for text-to-svg generation via an intr
- invad inversion-based reconstruction-free anomaly detection with diffusion model | arXiv: 2504.05662
- invcoss inversion-driven continual self-supervised learning in medical multi-mod
- inverfill one-step inversion for enhanced few-step diffusion inpainting
- ip-adapter is all you need towards fine-tuning-free diffusion-based talking face
- ipr-1 interactive physical reasoner | arXiv: 2511.15407
- ir-hgp physically-aware gaussian inverse rendering for high-illumination scenes
- iris bringing realworld priors into diffusion model for monocular depth estimation | arXiv: 2603.16340
- iris integrating language into diffusion-based monocular depth estimation
- irisfp adversarial-example-based model fingerprinting with enhanced uniqueness a | arXiv: 2603.24996
- is bin generation indispensable a bin-generation-free dataset quantization via s
- is your vlm sky-ready a comprehensive spatial intelligence benchmark for uav nav
- ishift lightweight slow-fast gui agent with adaptive perception
- isoclip decomposing clip projectors for efficient intramodal alignment | arXiv: 2603.19862
- isplat iterative learning for fine-grained gaussian splatting
- it takes two a duet of periodicity and directionality for burst flicker removal | arXiv: 2603.22794
- iterative closed-loop motion synthesis for scaling the capabilities of humanoid
- its never too late noise optimization for collapse recovery in trained diffusion | arXiv: 2601.00090
- ivaan instance-level vision-language alignment via attribute-guided text prompts
- jailbreaking vision-language models via dissonance-guided suffix optimization an
- janus a lightweight framework for jailbreaking text-to-image models via distribu
- jarvisevo towards a self-evolving photo editing agent with synergistic editor-ev
- joint learning of general and diverse patterns with mixture of memory experts fo
- joint spectral image reconstruction and semantic segmentation with cooperative u
- joint-aligned latent action towards scalable vla pretraining in the wild | arXiv: 2602.21736
- joppo hierarchical photography assessment via contrastive joint conditional prob
- jump-hand learning joint-wise uncertainty to gate mixture of view experts for mu
- kalos finds consensus a meta-algorithm for evaluating inter-annotator agreement
- kamp knowledge-anchored multimodal pretraining framework for medical image repre
- kasalv2 fully automatic 3d rotational symmetry classification and axis localizat
- keep it frozen domain-routed conditional residual modulation for multi-domain vi
- keep it sympl symbolic projective layout for allocentric spatial reasoning in vi
- klip localized distribution shift detection via kl-divergence with diffusion pri | arXiv: 2605.31596
- knowval a knowledge-augmented and value-guided autonomous driving system | arXiv: 2512.20299
- kontinuous kontext continuous strength control for instruction-based image editi
- kvsmooth mitigating hallucination in multi-modal large language models through k | arXiv: 2602.04268
- kαlos finds consensus a meta-algorithm for evaluating inter-annotator agreement | arXiv: 2603.27197
- l2dgs low-light dynamic gaussian splatting
- l3dr 3d-aware lidar diffusion and rectification
- label what matters modality-balanced and difficulty-aware multimodal active lear
- lactokgen latent consistency tokenizer for 1024-pixel image generation by 256 to
- lada robotic manipulation | arXiv: 2603.12967
- lady lagrangian-dynamic informed network for skeleton-based action segmentation
- LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis
- lamogen language to motion generation through llm-guided symbolic inference | arXiv: 2603.11605
- lamp language-assisted motion planning for controllable video generation | arXiv: 2512.03619
- lamp localization aware multi-camera people tracking in metric 3d world | arXiv: 2605.05390
- langfield4d learning identity-adaptive and spatio-temporal continuous 4d languag
- langref3dgs natural language-guided 3d referential segmentation from partial obs
- language does matter for cross-domain few-shot visual feature enhancement
- language models can explain visual features via steering | arXiv: 2603.22593
- language-free generative editing from one visual example | arXiv: 2603.25441
- language-guided frequency modulation for large vision-language models
- laof robust latent action learning with optical flow constraints | arXiv: 2511.16407
- large-scale robust enhanced ensemble clustering via outlier decoupling
- las-comp zero-shot 3d completion with latent-spatial consistency | arXiv: 2602.18735
- laser layer-wise scale alignment for training-free streaming 4d reconstruction | arXiv: 2512.13680
- lata laplacian-assisted transductive adaptation for conformal uncertainty in med
- latent chain-of-thought world modeling for end-to-end autonomous driving | arXiv: 2512.10226
- latent implicit visual reasoning
- lattice democratize high-fidelity 3d generation at scale
- lavr scene latent conditioned generative video trajectory re-rendering using lar
- layer consistency matters elegant latent transition discrepancy for generalizabl | arXiv: 2603.10598
- Layer-wise Instance Binding for Regional and Occlusion Control in Text-to-Image Diffusion Transformers
- Layered 4D-Rotor Gaussian Splatting: A Compressed Representation for Long Dynamic Scenes
- layoutad exploring semantic-geometric misalignment reasoning for scene layout an
- lazyvar accelerating visual autoregressive models via scale-wise token pruning a
- lca large-scale codec avatars the unreasonable effectiveness of large-scale avata | arXiv: 2604.02320
- LDP-Slicing: Local Differential Privacy for Images via Randomized Bit-Plane Slicing
- leader lidar relocalization | arXiv: 2604.11355
- leapalign post training flow matching models at any generation step | arXiv: 2604.15311
- learnability-driven submodular optimization for active roadside 3d detection | arXiv: 2601.01695
- learnability-guided diffusion for dataset distillation | arXiv: 2604.00519
- learnable motion-focused tokenization for effective and efficient video unsuperv
- learning 3d representations for spatial intelligence from unposed multi-view ima
- learning a unified latent action space from videos with action-centric cycle con
- learning and aligning click-aware shape prior for interactive amodal instance se
- learning compact 3d representations from feed-forward novel view synthesis
- learning complete and explainable visual representations from itemized text supe
- learning convex decomposition via feature fields
- learning coordinate-based convolutional kernels for continuous se3 equivariant a | arXiv: 2603.17538
- learning cross-view object correspondence via cycle-consistent mask prediction | arXiv: 2602.18996
- Learning Diffeomorphism for Medical Image Registration with Time-Embedded Architectures Using Semigroup Regularization
- learning differentiable hierarchies in 3d gaussian splatting
- learning effective sign features without text for gloss-free sign language trans
- learning explicit continuous motion representation for dynamic gaussian splattin | arXiv: 2603.25058
- learning forgery-aware lip representations without forgery priors
- learning from itself mining internal knowledge from vision language models for c
- learning from noisy supervision a denoising-debiasing framework for weakly super
- learning from oblivion predicting knowledge overflowed weights via retrodiction | arXiv: 2508.05059
- learning from semantic dictionaries discriminative codebook contrastive learning
- learning generalizable 3d medical image representations from mask-guided self-su | arXiv: 2603.13660
- learning hierarchical hyperbolic mixture model for part-aware 3d generation
- learning latent concepts for detecting out-of-distribution objects
- learning latent proxies for controllable single-image relighting | arXiv: 2603.15555
- learning latent transmission and glare maps for lens veiling glare removal | arXiv: 2511.17353
- learning like humans analogical concept learning for generalized category discov | arXiv: 2603.19918
- learning mutual view information graph for adaptive adversarial collaborative pe | arXiv: 2602.19596
- learning scene coordinate reconstruction from unposed images via pose graph opti
- learning spatial-temporal consistency for 3d semantic scene completion
- learning straight flows variational flow matching for efficient generation
- learning to adapt self-improving web agent via cognitive-aware exploration
- learning to assist physics-grounded human-human control via multi-agent reinforc | arXiv: 2603.11346
- learning to control physically-simulated 3d characters via generating and mimick
- learning to diversify and focus a reinforcement framework for open-vocabulary ho
- learning to drive is a free gift large-scale label-free autonomy pretraining fro | arXiv: 2602.22091
- learning to focus and precise croppinga reinforcement learning framework with in
- learning to generate highly dynamic videos using synthetic motion data
- learning to generate via understanding understanding-driven intrinsic rewarding | arXiv: 2603.06043
- learning to identify out-of-distribution objects for 3d lidar anomaly segmentati
- learning to reason in 4d dynamic spatial understanding for vision language model
- learning to refuse refusal-aware reinforcement fine-tuning for hard-irrelevant q
- learning to see and act task-aware virtual view exploration for robotic manipula | arXiv: 2508.05186
- learning to see through a babys eyes early visual diets enable robust visual int
- learning to see through illumination extremes with event streaming in multimodal
- learning to solve pdes on neural shape representations | arXiv: 2512.21311
- learning to track instance from single nature language description | arXiv: 2605.07064
- learning transferable temporal primitives for video reasoning via synthetic vide
- learning what helps task-aligned context selection for vision tasks
- learning what matters prioritized concept learning via relative error-driven sam | arXiv: 2506.01085
- learning where to look and how to judge resolution-agnostic image quality assess
- lemon a large endoscopic monocular dataset and foundation model for perception in | arXiv: 2503.19740
- lenses toward polysemous vision-language understanding
- lenswalk agentic video understanding by planning how you see in videos | arXiv: 2603.24558
- less is more data-efficient adaptation for controllable text-to-video generation
- let it snow animating 3d gaussian scenes with dynamic weather effects via physic | arXiv: 2504.05296
- let vlms grade their own thoughts a self-quantification approach to reasoning-aw
- let your image move with your motion -- implicit multi-object multi-motion trans | arXiv: 2603.01000
- leveraging class distributions in clip for weakly supervised semantic segmentati
- leveraging multispectral sensors for color correction in mobile cameras | arXiv: 2512.08441
- leveraging verifier-based reinforcement learning in image editing
- lf-bvn blind-view network for self-supervised light field denoising
- lidar prompted spatio-temporal multi-view stereo for autonomous driving
- lidar-to-4dradar diffusion bridge via cross-modal alignment and translation in l
- lidere a lightweight readout for fast and data-efficient dense prediction
- life-iqa boosting blind image quality assessment through gcn-enhanced layer inte
- lifeeval a multimodal benchmark for assistive ai in egocentric daily life tasks
- lifelong imitation learning multimodal latent rep | arXiv: 2603.10929
- lift and place a simple stable and effective knowledge distillation framework fo | arXiv: 2605.19729
- lifting unlabeled internet-level data for 3d scene understanding | arXiv: 2604.01907
- lightmover generative light movement with color and intensity controls | arXiv: 2603.27209
- lightsplat fast and memory-efficient open-vocabulary 3d scene understanding in f | arXiv: 2603.24146
- linear image generation by synthesizing exposure brackets
- linking modality isolation in heterogeneous collaborative perception | arXiv: 2603.00609
- linking perception confidence and accuracy in mllms | arXiv: 2603.12149
- linvideo a post-training framework towards on attention in efficient video gener | arXiv: 2510.08318
- lipschitz optimization for formal verification of homographies | arXiv: 2605.23203
- lirec-net a target-free and learning-based network for lidar rgb and event calib | arXiv: 2602.21754
- lite any stereo efficient zero-shot stereo matching | arXiv: 2511.16555
- litept lighter yet stronger point transformer | arXiv: 2512.13689
- litesense lifting lightweight tof with rgb for high-resolution metric depth esti
- live interactive training for video segmentation | arXiv: 2603.26929
- livegesture streamable co-speech gesture generation model
- llada-medv exploring large language diffusion models for biomedical image unders
- llamo scaling pretrained language models for unified motion understanding and ge
- llavashield multimodal multiturn safety | arXiv: 2509.25896
- llm-guided probabilistic fusion for label-efficient document layout analysis
- llmind bio-inspired training-free adaptive visual representations for vision-lan | arXiv: 2603.14882
- local motion matters a deconstruct-recompose paradigm for reinforcement learning
- local precise refinement a dual-gated mixture-of-experts for enhancing foundatio
- localizing structuring and rendering bridging 3d and 2d vision-language-action m
- locate-then-examine grounded region reasoning improves detection of ai-generated
- locate-then-sparsify attribution guided sparse strategy for visual hallucination | arXiv: 2603.16284
- lod-loc v3 generalized aerial localization in dense cities using instance silhou | arXiv: 2603.19609
- lofa learning to predict personalized prior for fast adaptation of visual genera
- logcd local-to-global consistency distillation for few-step image generation
- logit-margin repulsion for backdoor defense
- lol longer than longer scaling video generation to hour
- long-rvos a comprehensive benchmark for long-term referring video object segment
- longstream long-sequence streaming autoregressive visual geometry | arXiv: 2602.13172
- longvideo-r1 smart navigation for low-cost long video understanding | arXiv: 2602.20913
- look before you fuse 2d-guided cross-modal alignment for robust 3d detection | arXiv: 2507.16861
- lookasidevln direction-aware aerial vision-and-language navigation | arXiv: 2604.17190
- looking beyond the window global-local aligned clip for training-free open-vocab | arXiv: 2603.23030
- loreal mitigating low-resolution challenges in vision-language models with attri
- lost level of semantics tokenization for 3d shapes | arXiv: 2603.17995
- lottiegpt vector animation generation | arXiv: 2604.11792
- love me love my label rethinking the role of labels in prompt retrieval for visu | arXiv: 2604.03657
- low-rank test-time training for pre-trained point cloud models
- low-resolution editing is all you need for high-resolution editing | arXiv: 2511.19945
- ls-vit least-squares hessian based block reconstruction for low-bit post-trainin
- lumimotion gaussian relighting dynamics | arXiv: 2604.10994
- lumina a multi-vendor mammography benchmark with energy harmonization protocol | arXiv: 2603.14644
- luxremix lighting decomposition and remixing for indoor scenes | arXiv: 2601.15283
- lyapunov probes for hallucination detection in large foundation models
- Lynx: Towards High-Fidelity Personalized Video Generation
- m3dlayout a multi-source dataset of 3d indoor layouts and structured description | arXiv: 2509.23728
- m3docdep multi-modal multi-page multi-document dependency chunking with large vi
- m3grounder mask-based multi-span and multi-granular grounding for document qa
- m3kg rag multi hop multimodal knowledge graph enhanced retrieval augmented genera | arXiv: 2512.20136
- m4-sam multi-modal mixture-of-experts with memory-augmented sam for rgb-d video
- m4human a large-scale multimodal mmwave radar benchmark for human mesh reconstru
- m4v multimodal mamba for efficient text-to-video generation
- machine mental imagery empower multimodal reasoning with latent visual tokens
- machine unlearning via adaptive gradient reweighting and multi-stage objective o
- mactok robust continuous tokenization for image generation
- mad modality-adaptive decoding for mitigating cross-modal hallucinations in mult
- magicfuse single image fusion for visual and semantic reinforcement | arXiv: 2602.01760
- magician efficient long-term planning with imagined gaussians for active mapping | arXiv: 2603.22650
- MagicQuill V2: Precise and Interactive Image Editing with Layered Visual Cues
- majutsucity language-driven aesthetic-adaptive city generation with controllable
- makeanything harnessing diffusion transformers for multi-domain procedural seque
- making training-free diffusion segmentors scale with the generative power | arXiv: 2603.06178
- mamba learns in context structure-aware domain generalization for multi-task poi | arXiv: 2603.20739
- mambaliteunet cross-gated adaptive feature fusion for robust skin lesion segment | arXiv: 2604.20286
- mambasic mamba-based stereo image compression with bi-directional multi-referenc
- mamma markerless accurate multi-person motion acquisition
- mangobench a benchmark for multi-agent goal-conditioned offline reinforcement le
- manifoldgd training-free hierarchical manifold guidance for diffusion-based data
- manifoldneus manifold-aware view optimizability for pose-free neural surface rec
- mansion multi-floor language-to-3d scene generation for long-horizon tasks
- mantis a versatile vision-language-action model with disentangled visual foresig
- mapo motion-aware partitioning of deformable 3d gaussian splatting for high-fide
- mapreduce lora advancing the pareto front in multi-preference optimization for g | arXiv: 2511.20629
- maproute semantic routing concept erasure
- maprouteprecise-concept erasing mappers via semantic routing
- maps preserving vision-language representations via module-wise proximity schedu
- marco semantic correspondence | arXiv: 2604.18267
- maris marine open-vocabulary instance segmentation
- mark4d temporally-consistent watermarking for 4d gaussian splatting
- markovian scale prediction a new era of visual autoregressive generation | arXiv: 2511.23334
- markushgrapher-2 end-to-end multimodal recognition of chemical structures | arXiv: 2603.28550
- marss radar semantic segmentation via modular attention and state space models
- mask to align weight to disambiguate reliable unsupervised cross-modal hashing w
- maskadapt learning flexible motion adaptation via mask-invariant prior for physi | arXiv: 2603.29272
- maskdexgrasp generative masked modeling for part-aware dexterous grasp synthesis
- maskdime adaptive masked diffusion for precise and efficient visual counterfactu | arXiv: 2602.18792
- masked auto-regressive variational acceleration fast inference makes practical r
- masked region transformer for layered image generation and editing at scale
- masked-diffusion autoencoders for 3d medical vision representation learning
- maskfocus focusing policy optimization on critical steps for masked image genera
- masking matters unlocking the spatial reasoning capabilities of llms for 3d scen | arXiv: 2512.02487
- masquant modality-aware smoothing quantization for multimodal large language mod | arXiv: 2603.04800
- matanyone 2 scaling video matting via a learned quality evaluator | arXiv: 2512.11782
- match-and-fuse consistent generation from unstructured image sets | arXiv: 2511.22287
- matched crisp edge detection using end-to-end matching-based supervision | arXiv: 2602.20689
- matching every pair to track every point pairformer for all-pairs tracking and v
- matchmask mask-centric generative data augmentation for label-scarce semantic se
- material magic wand material-aware grouping of 3d parts in untextured meshes
- matmart material reconstruction of 3d objects via diffusion
- maxmark high-capacity diffusion-native watermarking via robust and invertible la
- mchdoc a comprehensive benchmark for reading multi-carrier chinese historical do
- md2e modeling depth-to-edge cues for monocular metric depth estimation
- mdcs-moame multi-directional composite scanning with mixture of attention and ma
- meanfuser fast one-step multi-modal trajectory generation and adaptive reconstru | arXiv: 2602.20060
- measure the feature universe topology-based pseudo labeling and gravity consiste
- measuring the unfaithfulness of concept-based explanations | arXiv: 2504.10833
- mechanisms of object localization in vision-language models | arXiv: 2605.19792
- medclipseg probabilistic vision-language adaptation for data-efficient and gener | arXiv: 2602.20423
- medgrpo multi-task reinforcement learning for heterogeneous medical video unders | arXiv: 2512.06581
- medic-ad towards medical vision-language models clinical intelligence | arXiv: 2603.27176
- medkco medical vision-language pretraining via knowledge-driven cognitive orches | arXiv: 2603.09101
- medlime a distribution-aligned and evidence-supported framework for medical sali
- medloc-r1 performance-aware curriculum reward scheduling for grpo-based medical
- medmo grounding and understanding multimodal large language model for medical im
- memflow a lightweight forward memorizing framework for quick domain adaptive fea
- memo human-like crisp edge detection using masked edge prediction | arXiv: 2603.20782
- memory efficient transfer learning with fading side networks | arXiv: 2604.09088
- memory matters boosting training-free zero-shot temporal action localization wit
- memory-augmented scene understanding and exploration for open-world aerial objec
- memory-efficient fine-tuning diffusion transformers via dynamic patch sampling a | arXiv: 2603.20755
- mer-tracker towards high-speed 3d point tracking via multi-view event-rgb hybrid
- merg3r a divide-and-conquer approach to large-scale neural visual geometry
- mergevla cross-skill model merging toward a generalist vision-language-action ag | arXiv: 2511.18810
- mesh-pro asynchronous advantage-guided ranking preference optimization for artis | arXiv: 2603.00526
- Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video
- meshflow efficient artistic mesh generation via meshvae and flow-based diffusion
- meshlam feed-forward one-shot animatable textured mesh avatar reconstruction | arXiv: 2604.22865
- MeshMosaic: Scaling Artist Mesh Generation via Local-to-Global Assembly
- meshripple structured autoregressive generation of artist-meshes
- meshsplatting differentiable rendering with opaque meshes
- meshweaver sparse-voxel-guided surface weaving for autoregressive mesh generatio | arXiv: 2606.04688
- meta meta evolution of tool trajectory adaptation for long-video understanding
- meta-fc meta-learning with feature consistency for robust and generalizable wate
- metaspectra a compact broadband metasurface camera for snapshot hyperspectral im | arXiv: 2603.09116
- meteorpred a meteorological multimodal large model and dataset for severe weathe
- mfen multi-frequency expert network for visible-infrared person re-id
- mgdhand multi-granularity prior-to-inertial distillation framework for sequentia
- mhopreg efficient hierarchical multi-hop graph search for point cloud registrati
- miburi towards expressive interactive gesture synthesis | arXiv: 2603.03282
- mico-150k a comprehensive dataset advancing multi-image composition
- micon-bench benchmarking and enhancing multi-image context image generation in u | arXiv: 2602.19497
- mimic human cognition master multi-image reasoning a meta-action framework for e
- mimicat mimic with correspondence-aware cascade-transformer for category-free 3d | arXiv: 2511.18370
- mimictalker a multimodal interactive and memory-enhanced framework for real-time
- mind the gap transferring labels to align object detection datasets
- mind the generative details direct localized detail preference optimization for | arXiv: 2601.04068
- mind the hitch dynamic calibration and articulated perception for autonomous tru | arXiv: 2603.23711
- mind the way you select negative texts pursuing the distance consistency in ood | arXiv: 2603.02618
- minddriver introducing progressive multimodal reasoning for autonomous driving | arXiv: 2602.21952
- mindpower enabling theoryofmind reasoning in vlmba | arXiv: 2511.23055
- minerva-ego spatiotemporal hints for egocentric video understanding
- minicpm-v 45 cooking efficient mllms via architecture data and training recipe
- mining attribute subspaces for efficient fine-tuning of 3d foundation models
- mining instance-centric vision-language contexts for human-object interaction de | arXiv: 2604.02071
- missing no more dictionary-guided cross-modal image fusion under missing infrare | arXiv: 2603.08018
- mistake attribution fine-grained mistake understanding in egocentric videos | arXiv: 2511.20525
- mitigating error amplification in fast adversarial training
- mitigating instance entanglement in instance-dependent partial label learning | arXiv: 2603.04825
- mitigating multimodal hallucinations via gradient-based self-reflection | arXiv: 2509.03113
- mitigating objectness bias and region-to-text misalignment for open-vocabulary p
- mitigating simplicity bias in ood detection through object co-occurrence analysi
- mitigating the distribution shift of diffusion-based dataset distillation
- mixercseg an efficient mixer architecture for crack segmentation via decoupled m | arXiv: 2603.01361
- mixture of prototypes for test-time adaptive segmentation
- mixture of states routing token-level dynamics for multimodal generation | arXiv: 2511.12207
- mixture-of-experts based feature decoupling for open vocabulary scene graph gene
- mllm-hwsi a multimodal large language model for hierarchical whole slide image u
- mllmsplat a 2d mllm-powered framework for 3d gaussian splatting understanding ge
- mm-act learn from multimodal parallel generation to act
- mm-ovseg multimodal optical-sar fusion for open-vocabulary segmentation in remot
- mm-recoder advancing chart-to-code generation with reinforcement learning and se | arXiv: 2604.01600
- mm-ser multimodal self-refinement for lightweight image captioning
- mmbench-gui a unified hierarchical evaluation framework for multi-platform gui a
- mmcp-gen a modality-extensible diffusion language model for conditional protein
- mmdir multimodal instruction-driven framework for mixed-degradation document ima
- mmface-dit a dual-stream diffusion transformer for high-fidelity multimodal face
- mmgait multi modal gait recognition | arXiv: 2604.15979
- mmr-ad a large-scale multimodal dataset for benchmarking general anomaly detecti
- mmrad multimodal anomaly detection | arXiv: 2604.10971
- mmsd30 a multi-image benchmark for real-world multimodal sarcasm detection
- mmtit-bench a multilingual and multi-scenario benchmark with cognition-perceptio | arXiv: 2603.23896
- MMVIP: A Visible-infrared Paired Dataset for Multi-weather Marine Vision
- mmwaveflow unified enhancement and generation of mmwave human point clouds
- mobile vton ondevice virtual tryon | arXiv: 2603.00947
- mocap-2-to-3 multi-view lifting for monocular motion recovery with 2d pretrainin
- mocapanything unified 3d motion capture for arbitrary skeletons from monocular v | arXiv: 2512.10881
- mocodiff a controllable autoregressive diffusion model for expressive motion gen
- mod-dpo towards mitigating cross-modal hallucinations in omni llms using modalit | arXiv: 2603.03192
- modeling spatiotemporal neural frames for high resolution brain dynamic | arXiv: 2603.24176
- modeling the brains grammar roi-guided fmri pretraining for transferable and int
- modeling the visual ambiguity of human sketches
- models as lego builders assembling malice from benign blocks via semantic bluepr
- modes accelerating mixture-of-experts multimodal large language models via dynam | arXiv: 2511.15690
- modix a training-free multimodal information-driven positional index scaling for
- modix positional index scaling | arXiv: 2604.12537
- modularagent a task-aware modular framework for joint optimization of multimodal
- moe-grpo optimizing mixture-of-experts via reinforcement learning in vision-lang | arXiv: 2603.24984
- moeactok a moe-based action tokenizer for vision-language-action models
- moeclip patch-specialized experts for zero-shot anomaly detection | arXiv: 2603.03101
- mofa-vton more fashion possibilities with fine-grained adaptations in virtual tr | arXiv: 2606.11148
- molingo motion-language alignment for text-to-motion generation | arXiv: 2512.13840
- molmo2 open weights and data for vision-language models with video understanding
- momo mars orbital model foundation model for mars orbital applications | arXiv: 2604.02719
- monet reasoning in latent visual space beyond image and language
- monocular open vocabulary occupancy prediction for indoor scenes | arXiv: 2602.22667
- monosaod monocular 3d object detection with sparsely annotated label | arXiv: 2604.01646
- moocap a multi-view benchmark for cow-object-human interaction and behavior dyna
- moon20 dynamic modality-balanced multimodal representation learning for e-commer
- more 3d visual geometry reconstruction meets mixture-of-experts
- more motion-aware feed-forward 4d reconstruction transformer | arXiv: 2603.05078
- more natural more real object-aware gaussian splatting for 3d visual decoding fr
- more than meets the eye a unified image fusion framework via semantic-pixel entr
- more than the sum panorama-language models for adverse omni-scenes | arXiv: 2603.09573
- more-stem long-short memory recall and spatio-temporal consistency model for que
- moregen multi-agent motion-reasoning engine for code-based text-to-video synthes
- morel long-range flicker-free 4d motion
- morel long-range flicker-free 4d motion modeling via anchor relay-based bidirect | arXiv: 2512.09270
- morphany3d unleashing the power of structured latent in 3d morphing | arXiv: 2601.00204
- morphseek fine-grained latent representation-level policy optimization for defor
- mos mitigating optical-sar modality gap for cross-modal ship re-identification | arXiv: 2512.03404
- mos mixture of states multimodal generation | arXiv: 2511.12207
- mosaic-gs monocular scene reconstruction via advanced initialization for complex
- mostly text smart visuals asymmetric text-visual pruning for large vision-langua | arXiv: 2603.16001
- motion-aware animatable gaussian avatars deblurring | arXiv: 2411.16758
- motionaware animatable gaussian avatars deblurring | arXiv: 2411.16758
- MotionEdit: Benchmarking and Learning Motion-Centric Image Editing
- motionenhancer leveraging video diffusion for motion-enhanced vision-language mo | arXiv: 2606.06853
- motionhiflow text-to-motion via hierarchical flow matching | arXiv: 2604.23264
- motionscale reconstructing appearance geometry and motion of dynamic scenes with | arXiv: 2603.29296
- movie broaden your views with human motion for action detection
- movierecapsqa a multimodal open-ended video question-answering benchmark | arXiv: 2601.02536
- movies motion-aware 4d dynamic view synthesis in one second | arXiv: 2507.10065
- mpdit multi-patch global-to-local transformer architecture for efficient flow ma | arXiv: 2603.26357
- mpl match-guided prototype learning for few-shot action recognition
- mr illuminate zero-shot low-light image enhancement with diffusion prior
- mr-rag multimodal relevance-aware retrieval-augmented generation for medical vis
- mrd multi-resolution retrieval-detection fusion for high-resolution image unders | arXiv: 2512.02906
- mri contrast enhancement kinetics world model | arXiv: 2602.19285
- mrm masked representation modeling domain adaptive | arXiv: 2509.13801
- mrt masked region transformer for layered image generation and editing at scale | arXiv: 2605.27235
- mscd-gs motion-separated cooperative deblurring dynamic reconstruction via gauss
- msgnav unleashing the power of multi-modal 3d scene graph for zero-shot embodied | arXiv: 2511.10376
- msjoe jointly evolving mllm and sampler for efficient long-form video understand | arXiv: 2602.22932
- mspt efficient large-scale physical modeling via parallelized multi-scale attent
- msrl scaling generative multimodal reward modeling | arXiv: 2603.25108
- mta multimodal task alignment for bev perception and captioning
- mu-generf multi-view uncertainty-guided generalizable neural radiance fields for | arXiv: 2604.17965
- mufasa a multi-layer framework for slot attention
- mukv multi-grained kv cache compression for long streaming video question-answer
- multi-crit benchmarking multimodal judges on pluralistic criteria-following | arXiv: 2511.21662
- multi-level causal llm-based text-to-motion generation with human alignment
- multi-metric representation learning strategy based on clustering for fine-grain
- multi-modal frequency decomposition network for semantic scene completion
- multi-modal image fusion via intervention-stable feature learning | arXiv: 2603.23272
- multi-modal representation learning via semi-supervised rate reduction for gener | arXiv: 2602.19910
- multi-modal test-time adaptation via adaptive probabilistic gaussian calibration
- multi-paradigm collaborative adversarial attack against multi-modal large langua | arXiv: 2603.04846
- multi-patch global-to-local transformer architecture for efficient flow matching
- multi-prototype compactness and boundary-aware synthesis for unsupervised anomal
- multi-scale gaussian-language map for zero-shot embodied navigation and reasonin | arXiv: 2605.01736
- multi-scale gradient-guided unrolling architecture with adaptive mamba for compr
- multi-scale local speculative decoding for image generation | arXiv: 2601.05149
- multi-spatialmllm multi-frame spatial understanding with multi-modal large langu | arXiv: 2505.17015
- multi-view consistent 3d gaussian head avatars without multi-view generation | arXiv: 2605.25220
- multi-view crowd tracking transformer with view-ground interactions under large | arXiv: 2604.19318
- multi-view hierarchical alignment learning for spatial transcriptomics
- multibanana a challenging benchmark for multi reference text to image generation | arXiv: 2511.22989
- multicrafter high-fidelity multi-subject generation via disentangled attention a
- multigrain-aware semantic prototype scanning and tri-token prompt learning embra
- multimodal causal-driven representation learning for generalizable medical image | arXiv: 2508.05008
- multimodal continual instruction tuning with dynamic gradient guidance
- multimodal distribution matching for vision-language dataset distillation | arXiv: 2605.23482
- multimodal learning on low-quality data with conformal predictive self-calibrati | arXiv: 2605.03820
- multimodal protein language models for enzyme kinetic parameters from substrate | arXiv: 2603.12845
- multimodal rewardbench 2 evaluating omni reward models for interleaved text and
- multimodal semantic bias mitigation for diverse text-to-3d generation
- multimodalpfn extending prior-data fitted networks for multimodal tabular learni | arXiv: 2602.20223
- multishotmaster a controllable multi-shot video generation framework
- mupo all roads lead to rome incentivizing divergent thinking in vlms | arXiv: 2604.00479
- muse harnessing precise and diverse semantics for few-shot whole slide image cla | arXiv: 2602.20873
- muses designing composing generating nonexistent fantasy 3d creatures without tr
- musicinfuser making video diffusion listen and dance | arXiv: 2503.14505
- must modality-specific representation-aware transformer for diffusion-enhanced s | arXiv: 2603.26071
- muvit multi-resolution vision transformers for learning across scales in microsc | arXiv: 2602.24222
- mv-fashion towards enabling virtual try-on and size estimation with multi-view p
- mv-roma from pairwise matching into multi-view track reconstruction | arXiv: 2603.27542
- mv3dis multi-view mask matching via 3d guides for zero-shot 3d instance segmenta
- mvggt multimodal visual geometry grounded transformer for multiview 3d referring | arXiv: 2601.06874
- mvlm a vision language model for mnpus
- mvlm template-free tracking via vision-language margin confidence and memory-gat
- naf zero-shot feature upsampling via neighborhood attention filtering
- nami efficient image generation via bridged progressive rectified flow transform
- narrative weaver towards controllable long-range visual consistency with multi-m | arXiv: 2603.06688
- native and compact structured latents for 3d generation | arXiv: 2512.14692
- natural human motion recovery by aligning high-order temporal dynamics from mono | arXiv: 2605.26879
- NEAF: Natural Image Editing with Attention Fusion for Generalizable Test-time Optimization in Text-Guided Image Editing
- nec-diff noise-robust event-raw complementary diffusion for seeing motion in ext | arXiv: 2603.20005
- negative binomial variational autoencoders for overdispersed latent modeling
- neighbor-aware localized concept erasure in text-to-image diffusion models | arXiv: 2603.25994
- neighbormae exploiting spatial dependencies between neighboring earth observatio
- neoverse enhancing 4d world model with in-the-wild monocular videos | arXiv: 2601.00393
- nerfify multiagent nerf paper to code | arXiv: 2603.00805
- nestwork conditional 3d furnished house layout generation through latent heterog
- neu-pig neural preconditioned grids for fast dynamic surface reconstruction on l | arXiv: 2602.22212
- neural collapse in test-time adaptation | arXiv: 2512.10421
- neural differentiation in deep networks a theoretical framework for expressivity
- neural distribution prior for lidar ood detection | arXiv: 2604.09232
- neural dynamic gi random-access neural compression for temporal lightmaps in dyn | arXiv: 2604.12625
- neural field-based 3d surface reconstruction of microstructures from multi-detec | arXiv: 2508.04728
- neural gabor splatting | arXiv: 2604.15941
- neural mixture density processes
- Neuro-Cognitive Reward Modeling for Human-Centered Autonomous Vehicle Control
- neurodynamics-driven coupled neural p systems for multi-focus image fusion | arXiv: 2509.17704
- neuroflow toward unified visual encoding and decoding from neural activity | arXiv: 2604.09817
- neuroseg meets dinov3 transferring 2d self-supervised visual priors to 3d neuron | arXiv: 2603.23104
- next-scale autoregressive models for text-to-motion generation | arXiv: 2604.03799
- next-scale prediction a self-supervised approach for real-world image denoising
- nexusflow unifying disparate tasks under partial supervision via invertible flow
- ng gs nerf guided 3d gaussian splatting segmentation | arXiv: 2604.14706
- nimbusgs unified 3d scene reconstruction under hybrid weather | arXiv: 2603.27228
- no calibration no depth no problem cross-sensor view synthesis with 3d consisten | arXiv: 2602.23559
- no hard negatives required concept centric learning leads to compositionality wi | arXiv: 2603.25722
- no labels no look-ahead unsupervised online video stabilization with classical p | arXiv: 2602.23141
- no need for real anomaly mllm empowered zero-shot video anomaly detection | arXiv: 2602.19248
- no way to steal my face proactive defense against identity-preserving personaliz
- node-rf learning generalized continuous space-time scene dynamics with neural od | arXiv: 2603.12078
- noise-aware few-shot learning through bi-directional multi-view prompt alignment | arXiv: 2603.11617
- nonparametric deep fine-grained clustering with low-rank guided vision-language
- noovd novel category discovery and embedding for open-vocabulary object detectio | arXiv: 2603.21069
- nord a data-efficient vision-language-action model that drives without reasoning | arXiv: 2602.21172
- nowa null-space optical watermark for invisible capture fingerprinting and tampe
- ns-diff fluid navier-stokes guided video diffusion via reinforcement learning
- ntk-guided implicit neural teaching | arXiv: 2511.15487
- nvgs neural visibility for occlusion culling in 3d gaussian splatting
- object-generalized re-identification a step towards universal instance perceptio
- objectmorpher 3d-aware image editing via deformable 3dgs
- obstruction reasoning for robotic grasping
- occany generalized unconstrained urban 3d occupancy | arXiv: 2603.23502
- occlusion-aware sort observing occlusion for robust multi-object tracking | arXiv: 2603.06034
- occufly a 3d vision benchmark for semantic scene completion from the aerial pers | arXiv: 2512.20770
- octonav towards generalist embodied navigation
- octopus history-free gradient orthogonalization for continual learning in multim
- octot2i a self-evolving agentic text-to-image router
- oddgridbench exposing the lack of fine-grained visual discrepancy sensitivity in | arXiv: 2603.09326
- odgs-slam omnidirectional gaussian splatting slam
- off the grid detection of primitives for feed-forward 3d gaussian splatting | arXiv: 2512.15508
- olbedo an albedo and shading aerial dataset for large-scale outdoor environments | arXiv: 2602.22025
- omg-bench a new challenging benchmark for skeleton-based online micro hand gestu | arXiv: 2512.16727
- omgtex one-stage multi-style facial texture reconstruction without geometry guid | arXiv: 2605.25778
- omni iie bench benchmarking the practical capabilities of image editing models
- omni-ad a large-scale and versatile benchmark for industrial anomaly detection
- omni-attack adversarial attacks on open-ended vqa in black-box multimodal llms
- omni-attribute open-vocabulary attribute encoder for visual concept personalizat | arXiv: 2512.10955
- omni-fake benchmarking unified multimodal social media deepfake detection | arXiv: 2605.01638
- omni-mmsi toward identity-attributed social interaction understanding | arXiv: 2604.00267
- omni-supervised motion editing balancing change and invariance through positive-
- omnibrainbench a comprehensive multimodal benchmark for brain imaging analysis a
- omnidoclayout towards diverse document layout generation via coarse-to-fine llm
- omnifm toward modality-robust and task-agnostic federated learning for heterogen | arXiv: 2603.21660
- omnifood8k nutrition estimation | arXiv: 2604.12356
- omniground a comprehensive spatio-temporal grounding benchmark for real-world co
- omnilottie generating vector animations via parameterized lottie tokens | arXiv: 2603.02138
- omniret efficient and high-fidelity omni modality retrieval | arXiv: 2603.02098
- omnisonic towards universal and holistic audio generation from video and text | arXiv: 2604.04348
- omnivggt omni-modality driven visual geometry grounded transformer
- omnivtg a large-scale dataset and training paradigm for open-world video tempora | arXiv: 2604.25276
- omnizip audio-guided dynamic token compression for fast omnimodal large language | arXiv: 2511.14582
- omnizip learning a unified and lightweight lossless compressor for multi-modal d
- OMoBlur: An Object Motion Blur Dataset and Benchmark for Real-World Local Motion Deblurring
- on tokens dilemma dynamic moe with drift-aware token assignment for continual le | arXiv: 2603.27481
- one algorithm to align them all
- one layers trash is another layers treasure adaptive layer-wise visual token sel
- one model many budgets elastic latent interfaces for diffusion transformers | arXiv: 2603.12245
- one token two fates a unified framework via vision token manipulation against ml
- one-shot flow any-time frame a bidirectional warping framework for event-based v
- one-shot flow any-time frame a bidirectional warping framework for event-based v
- one-step diffusion transformer for controllable real-world image super-resolutio
- onecat decoder-only auto-regressive model for unified understanding and generati
- oneocc semantic occupancy prediction for legged robots with a single panoramic c | arXiv: 2511.03571
- onestory coherent multi-shot video generation with adaptive memory
- onethinker all-in-one reasoning model for image and video | arXiv: 2512.03043
- online data curation for object detection via marginal contributions to dataset-
- online3r online learning for consistent sequential reconstruction based on geome
- onlinehmr video-based online world-grounded human mesh recovery | arXiv: 2603.17355
- onlinepg online open-vocabulary panoptic mapping with 3d gaussian splatting | arXiv: 2603.18510
- open the motion door atomic motion decomposition and recomposition for open-voca
- open-ended instruction realization with llm-enabled multi-planner scheduling in
- open-vocabulary domain generalization in urban-scene segmentation | arXiv: 2602.18853
- open-world hand-object interaction video generation based on structure and conta
- opendance multimodal controllable 3d dance generation with large-scale internet
- opendpr open-vocabulary change detection via vision-centric diffusion-guided pro | arXiv: 2603.27645
- openfs multi-hand-capable fingerspelling recognition with implicit signing-hand | arXiv: 2602.22949
- opening the sim-to-real door for humanoid pixel-to-action policy transfer
- openmarcie dataset for multimodal action recognition in industrial environments | arXiv: 2603.02390
- openmmreasoner pushing the frontiers in multimodal reasoning with an open and ge
- opent2m no-frill motion generation with open-source large-scale high-quality dat
- openvision 2 a family of generative pretrained visual encoders for multimodal le
- openvo open-world visual odometry with temporal dynamics awareness | arXiv: 2602.19035
- openvoxel training-free grouping and captioning voxels for open-vocabulary 3d sc
- opro orthogonal panel-relative operators for panel-aware in-context image genera | arXiv: 2603.27637
- opti-neus neural reconstruction for dual-layered transparent and opaque objects
- optical diffraction-based convolution for semiconductor lithography
- optimvmap offline vectorized map construction via optimal multi-vehicle perspect
- oralgpt-plus learning to use visual tools via reinforcement learning for panoram
- orapo oracle-educated reinforcement learning for data-efficient and factual radi | arXiv: 2509.18600
- orbit benchmarking sfm in the wild with 360deg video
- orbital video 3d foundation priors | arXiv: 2604.12309
- orca orchestrated reasoning with collaborative agents for document visual questi
- oric benchmarking object recognition under contextual incongruity in large visio
- orienpose orientation-guided novel view synthesis for single-image unseen object
- orion orthonormal text encoding for universal vlm adaptation
- OrionEdit: Bridging Reference and Source Images for Generalized Cross-Image Editing
- orsatr-x a foundation model based on differential-and-excitation networks for op
- orthofuse training-free riemannian fusion of orthogonal style-concept adapters f
- orthogonal spatial-aware multi-view anchor graph clustering for incomplete remot
- os-oracle a comprehensive framework for cross-platform gui critic models
- osa echocardiography video segmentation via orthogonalized state update and anat
- oslash source models leak what they shouldnt nrightarrow unlearning zero-shot tr | arXiv: 2604.08238
- ospo object-centric self-improving preference optimization for text-to-image gen
- otil accelerating diffusion model inference via communication-efficient multi-gp
- out of sight out of track adversarial attacks on propagation-based multi-object | arXiv: 2604.00452
- ov3r open-vocabulary semantic 3d reconstruction from rgb videos
- ovod-agent a markov-bandit framework for proactive visual reasoning and self-evo
- p2gs physical prior-guided gaussian splatting for photometrically consistent urb
- pa-attack guiding gray-box attacks on lvlm vision encoders with prototypes and a
- PaCo-RL: Advancing Reinforcement Learning for Consistent Image Generation with Pairwise Reward Modeling
- pact phase-like transition constraints in adapter-based continual learning of vi
- pad-hand physics-aware diffusion for hand motion recovery | arXiv: 2603.26068
- paddleocr vl coarse to fine document parsing | arXiv: 2603.24326
- pai-bench a comprehensive benchmark for physical ai
- palm progress-aware policy learning via affordance reasoning for long-horizon ro | arXiv: 2601.07060
- pamotion physics-aware motion generation for full-body interaction with multiple
- panda unsupervised domain adaptation for multimodal 3d panoptic segmentation in
- pano360 perspective to panoramic vision with geometric consistency | arXiv: 2603.12013
- pano3dcomposer feed-forward compositional 3d scene generation from single panora | arXiv: 2603.05908
- panoenv exploring 3d spatial intelligence in panoramic environments with reinfor
- panovggt feed-forward 3d reconstruction from panoramic imagery | arXiv: 2603.17571
- pantheon360 taming digital twin generation via 3d-aware 360 video diffusion | arXiv: 2605.25449
- pantheon360 taming digital twin generation via 3d-aware 360deg video diffusion
- paparazzo active mapping of moving 3d objects
- paper2figure a multi-agent collaborative system for figure generation towards ac
- paq-detr learning pattern and quality-aware dynamic queries for object detection | arXiv: 2603.06917
- parallax to align them all an omniparallax attention mechanism for distributed m | arXiv: 2603.03615
- parallel jacobi decoding for fast autoregressive image generation | arXiv: 2606.05703
- parallel rigidity matters for bundle adjustment
- parallelised differentiable straightest geodesics for 3d meshes | arXiv: 2603.15780
- parallelvlm lossless video-llm acceleration with visual alignment aware parallel
- parameter-efficient adaptation for mllms via implicit modality decomposition
- parameter-efficient continual learning for enhancing plasticity without forgetti
- parameter-efficient semantic augmentation for enhancing open-vocabulary object d | arXiv: 2604.04444
- parameterized prompt for incremental object detection
- parauni enhance generation in unified multimodal model with reinforcement-driven
- parse search and confirmation training-free aerial vision-and-dialog navigation
- part2gs part-aware modeling of articulated objects using 3d gaussian splatting
- partdiffuser part-wise 3d mesh generation via discrete diffusion
- partial weakly-supervised oriented object detection
- particlegs learning neural gaussian particle dynamics from videos for prior-free
- particulate feed-forward 3d object articulation | arXiv: 2512.11798
- party part-guidance for expressive text-to-motion synthesis | arXiv: 2603.09611
- pas prelim attention score for detecting object hallucinations in large vision-l
- paul uncertainty-guided partition and augmentation for robust cross-view geo-loc
- pavas physics-aware video-to-audio synthesis
- pc-talk precise facial animation control for audio-driven talking face generatio
- pca-seg revisiting cost aggregation for openvocabulary semantic and part segmentat | arXiv: 2603.17520
- pdcr perception-decomposed confidence reward for vision-language reasoning | arXiv: 2605.13467
- pe3r perception-efficient 3d reconstruction | arXiv: 2503.07507
- pearl geometry aligns semantics for training-free open-vocabulary semantic segme | arXiv: 2603.21528
- peccvai overcoming the brittleness of ai image watermarking under visual paraphr
- perceiving the near reasoning the distant coherent long-horizon trajectory predi
- percept-wam perception-enhanced world-awareness-action model for robust end-to-e
- perception characteristics distance measuring stability and robustness of percep | arXiv: 2506.09217
- perceptual 3d simulation with physical world modeling
- perceptual-evidence anchored reinforced learning for multimodal reasoning
- performrecast expression and head pose disentanglement for portrait video editin | arXiv: 2603.19731
- personalized longitudinal medical report generation via temporally-aware federat
- personavlm long term personalized multimodal llms | arXiv: 2604.13074
- pet-dino unifying visual cues into grounding dino with prompt-enriched training | arXiv: 2604.00503
- petar localized findings generation with mask-aware vision-language modeling for
- pfgnet a fully convolutional frequency-guided peripheral gating network for effi | arXiv: 2602.20537
- ph-strips for selective forgetting a blunt but fast diagnostic baseline for mach
- phac promptable human amodal completion | arXiv: 2603.14741
- phantom physical object interactions as dynamic triggers for nms-exploited backd
- phantom physics-infused video generation via joint modeling of visual and latent | arXiv: 2604.08503
- phase-net physics-grounded harmonic attention system for efficient remote photop | arXiv: 2509.24850
- phased dmd few-step distribution matching distillation via score matching within
- phasewin search framework enable efficient object-level interpretation
- phasr generalized image shadow removal with physically aligned priors | arXiv: 2601.17470
- phenoyieldnet learning crop-aware phenological responses for multi-crop yield pr
- photo3d advancing photorealistic 3d generation through structure-aligned detail
- phrase-grounded apo for improving chest x-ray report generation
- phrase-grounding-aware supervised fine-tuning for chart recognition via side-mas
- phyco learning controllable physical priors for generative motion | arXiv: 2604.28169
- phycritic multimodal critic models for physical ai
- phygap physically-grounded gaussians with polarization cues | arXiv: 2603.14001
- physgaia a physics-aware benchmark with multi-body interactions for dynamic nove | arXiv: 2506.02794
- physgm large physical gaussian 4d synthesis | arXiv: 2508.13911
- physgs bayesian-inferred gaussian splatting for physical property estimation | arXiv: 2511.18570
- physhead simulation-ready gaussian head avatars | arXiv: 2604.06467
- physho physics-based dynamic 3d gaussian human and object from monocular video
- physical adversarial clothing evades visible-thermal detectors via non-overlappi | arXiv: 2605.04675
- physical simulator in-the-loop video generation | arXiv: 2603.06408
- physically ground commonsense knowledge for articulated object manipulation with
- physically inspired gaussian splatting for hdr novel view synthesis | arXiv: 2603.28020
- physically-grounded turbulence mitigation with frame-shared degradation paramete
- physics-consistent diffusion for efficient fluid super-resolution via multiscale | arXiv: 2603.00149
- physics-guided multistep deformation reversal for ancient bamboo slip restoratio
- physir-splat physically consistent thermal infrared radiative transfer in 3d gau
- physisinone visual physics learning and reasoning in one suite | arXiv: 2604.09415
- physskin real-time and generalizable physics-based animation via self-supervised | arXiv: 2603.23194
- physvid physics aware local conditioning for generative video models | arXiv: 2603.26285
- pico-banana-400k a large-scale dataset for text-guided image editing
- pilot neural pixel-to-3d registration for uav-based ego and target geo-localizat
- pinpoint evaluation of composed image retrieval with explicit negatives multi-im | arXiv: 2603.04598
- pip-stereo progressive iterations pruner for iterative optimization based stereo | arXiv: 2602.20496
- pix-tab efficient pixel-precise table structure recognition approach with specul
- pix-tab efficient pixel-precise table structure recognition approach with specul
- pixarmesh autoregressive mesh-native single-view scene reconstruction | arXiv: 2603.05888
- pixdlm uav reasoning segmentation | arXiv: 2604.15670
- pixel motion diffusion is what we need for robot control | arXiv: 2509.22652
- pixel2phys distilling governing laws from visual dynamics | arXiv: 2602.19516
- pixeldit pixel diffusion transformers for image generation | arXiv: 2511.20645
- pixelrush ultrafast trainingfree highresolution im | arXiv: 2602.12769
- pixels dont lie but your detector might bootstrapping mllm-as-a-judge for trustw | arXiv: 2602.19715
- placid identity-preserving multi-object compositing via video diffusion with syn
- planareloc camera relocalization in 3d planar primitives via region-based struct | arXiv: 2603.20818
- plannerrft reinforcing diffusion planners
- plannerrft reinforcing diffusion planners through closed-loop and sample-efficie
- planning in 8 tokens a compact discrete tokenizer for latent world model | arXiv: 2603.05438
- plant taxonomy meets plant counting a fine-grained taxonomic dataset for countin | arXiv: 2603.21229
- plenoptic video generation
- plug-and-play incomplete multi-view clustering via janus-faced affinity learning
- plug-and-play pde optimization for 3d gaussian splatting toward high-quality ren
- pluggable pruning with contiguous layer distillation for diffusion transformers | arXiv: 2511.16156
- pmrnet physics-informed multi-scale refinement network for medical image segment
- poga paraphrased and oppositional graph alignment for fine-grained cross-modal r
- poinit-of-view poisoning initialization of views transfers across multiple 3d re | arXiv: 2604.16540
- point4cast streaming dynamic scene reconstruction and forecasting
- pointalign feature-level alignment regularization for 3d vision-language models | arXiv: 2603.00412
- pointcsp cross-sample semantic propagation and stability preservation in self-su
- pointer-cad unifying b-rep and command sequences via pointer-based edges faces s | arXiv: 2603.04337
- pointgs semantic-consistent unsupervised 3d point cloud segmentation with 3d gau
- pointing at parts training-free few-shot grounding in multimodal llms
- pointnsp autoregressive 3d point cloud generation with next-scale level-of-detai
- points-long adaptive dual-mode visual reasoning in mllms
- points-to-3d structure-aware 3d generation with point cloud priors | arXiv: 2603.18782
- pointthinker point-incentivized parallel thinking for multimodal large language
- pointtpa dynamic network parameter adaptation for 3d scene understanding | arXiv: 2604.04933
- polar a portrait olat dataset and generative framework for illumination-aware fa
- PolarGuide-GSDR: 3D Gaussian Splatting Driven by Polarization Priors and Deferred Reflection for Real-World Reflective Scenes
- polarization state tracing for reflection removal and color-consistent reconstru
- polyphony diffusion-based dual-hand action segmentation with alternating vision | arXiv: 2605.31115
- polyslgen online multimodal speaking-listening reaction generation in polyadic i
- pop proof of perception conformal reasoning | arXiv: 2603.00324
- portable active learning for object detection | arXiv: 2605.10349
- portraitdirector a hierarchical disentanglement framework for controllable and r
- pose-free omnidirectional gaussian splatting for 360-degree videos with consiste
- pose-guided enriched feature learning for federated-by-camera person re-identifi
- poseanything general pose-guided video generation with part-aware temporal coher
- posegam robust unseen object pose estimation via geometry-aware multi-view reaso
- PoseGaussian: 6D Pose Estimation for Unseen Objects via Sparse-View Object-Level 3D Gaussian Splatting
- posemaster a unified 3d native framework for stylized pose generation | arXiv: 2506.21076
- post-training feature pruning for fundus images classification
- posteromni generalized artistic poster creation via task distillation and unifie
- posterreward unlocking accurate evaluation for high-quality graphic design gener
- pour a provably optimal method for unlearning representation via neural collapse
- pour a provably optimal method for unlearning representations via neural collaps | arXiv: 2511.19339
- pp-ocrv5 a specialized 5m-parameter model rivaling billion-parameter vision-lang
- ppcl pluggable pruning dit distillation | arXiv: 2511.16156
- ppisp physically-plausible compensation and control of photometric variations in
- ppm-clip probabilistic prompt modeling for generalizable ai-generated image dete
- pqdt pseudo-query dual transformer for robust point cloud restoration
- pr-iqa partial-reference image quality assessment for diffusion-based novel view | arXiv: 2604.04576
- pr-magic prompt refinement via mask decoder gradient flow for in-context segment
- precise object and effect removal with adaptive target-aware attention | arXiv: 2505.22636
- predict before you explore predictive planning with specialized memory for embod
- predicting spatial transcriptomics from histology images via high-order multi-ce
- predictive regularization against visual representation degradation in multimoda | arXiv: 2603.20808
- preference-aligned lora merging preserving subspace coverage and addressing dire | arXiv: 2603.26299
- prefill-time intervention for mitigating hallucination in large vision-language | arXiv: 2604.25642
- premier personalized preference modulation with learnable user embedding in text
- preserving source video realism high-fidelity face swapping for cinematic qualit | arXiv: 2512.07951
- pressure2motion hierarchical human motion reconstruction from ground pressure wi
- prime once then reprogram locally an efficient alternative to black-box service | arXiv: 2604.01474
- primu uncertainty estimation for novel views in gaussian splatting from primitiv
- principled steering via null-space projection for jailbreak defense in vision-la | arXiv: 2603.22094
- prism learning a shared primitive space for transferable skeleton action represe
- prism prototype-based reasoning with inter-modal semantic mining for interpretab
- prism video dataset condensation with progressive refinement and insertion for s | arXiv: 2505.22564
- pritti primitive-based generation of controllable and editable 3d semantic urban | arXiv: 2506.19117
- privi towards a general-purpose video model for primate behavior in the wild | arXiv: 2511.09675
- privsynth alternating and control-based optimization for privacy and utility in
- proactivemobile a comprehensive benchmark for boosting proactive intelligence on
- probabilistic discrepancy learning for roadside lidar scene completion
- probabilistic prompt adaptation for unified image aesthetics and quality assessm
- probing and bridging geometry-interaction cues for affordance reasoning in visio | arXiv: 2602.20501
- processmaker a generalized process visualization framework with adaptive sequenc
- profocus proactive perception and focused reasoning in vision-and-language navig | arXiv: 2603.05530
- progress by pieces test-time scaling for autoregressive image generation
- progressive cross-modal causal intervention for long-term action recognition
- progressive guessing to fixed point rethinking human motion prediction with deep
- progressive neural architecture generation
- progressiveavatars progressive animatable 3d gaussian avatars | arXiv: 2603.16447
- progtrack a multi-object tracking algorithm with progressive matching strategy
- projflow projection sampling with flow matching for zero-shot exact spatial moti
- promo promptable virtual tryon efficient | arXiv: 2603.11675
- prompt-anchored vision-text distillation for lifelong person re-identification | arXiv: 2605.05027
- prompt-free universal region proposal network | arXiv: 2603.17554
- prompt-free unknown label generation for open world detection in remote sensing
- promptdepth efficient and promptable geometric 3d vision model for embodied inte
- promptenhancer taming your rewriter for text-to-image generation via fine-graine
- promptloop plug-and-play prompt refinement via latent feedback for diffusion mod
- promptminer black-box prompt stealing against text-to-image generative models vi
- promptmoe a segmentation refinement framework leveraging mixture of experts for
- promptstereo zero-shot stereo matching via structure and motion prompts | arXiv: 2603.01650
- proood prototype-guided out-of-distribution 3d occupancy prediction | arXiv: 2604.01081
- propfly learning to propagate via on-the-fly supervision from pre-trained video
- prosoftarena benchmarking hierarchical capabilities of multi-modal agents in pro
- prospective dynamic 3d mri reconstruction via latent-space motion tracking from
- protect to adapt orthogonal subspace control with ranked negative-prompt curricu
- protect to adapt orthogonal subspace control with ranked negative-prompt curricu
- protego user-centric pose-invariant privacy protection against face recognition-
- prototype-as-prompt multimodal sentiment prototypes endowing large language mode
- prototype-based causal intervention for multi-label image classification
- prototype-guided concept erasure in diffusion models | arXiv: 2603.08271
- prototypical action reasoning facilitated by vision-language alignment for egoce
- proxy-gs unified occlusion priors for training and inference in structured 3d ga
- proxy-tuning tailoring multimodal autoregressive models for subject-driven image
- proxy3d efficient 3d representations for vision-language models via semantic clu | arXiv: 2605.08064
- proxyfl a proxy-guided framework for federated semi-supervised learning | arXiv: 2602.21078
- prue a practical recipe for field boundary segmentation at scale | arXiv: 2603.27101
- prune wisely reconstruct sharply compact 3d gaussian splatting via adaptive prun | arXiv: 2602.24136
- prune2drive a plug-and-play framework for accelerating vision-language models in | arXiv: 2508.13305
- psdesigner automated graphic design with a human-like creative workflow | arXiv: 2603.25738
- psr scaling multi-subject personalized image generation with pairwise subject-co | arXiv: 2512.01236
- ptc-depth pose-refined monocular depth estimation with temporal consistency | arXiv: 2604.01791
- purecc pure learning for text-to-image concept customization | arXiv: 2603.07561
- pureproof diffusion-resistant black-box targeted attack on large vision-language
- push-and-step from rl-based balance recovery to physical simulation of dense cro
- push-and-step from rl-based balance recovery to physical simulation of dense cro
- pushing the frontier of audiovisual perception with large-scale multimodal corre
- pv-ground text-guided point-voxel interaction for 3d visual grounding
- PvP: Data-Efficient Humanoid Robot Learning with Proprioceptive-Privileged Contrastive Representations
- pyramidalwan on making pretrained video model pyramidal for efficient inference
- PyraTok: Language-Aligned Pyramidal Tokenizer for Video Understanding and Generation
- qd-pcqa quality-aware domain adaptation for point cloud quality assessment | arXiv: 2603.03726
- qkd quantum gated incremental learning | arXiv: 2604.11112
- quadsync quadrifocal tensor synchronization via tucker decomposition | arXiv: 2602.22639
- quant experts token aware vlm quantization | arXiv: 2602.24059
- quantiphy a quantitative benchmark evaluating physical reasoning abilities of vi
- quantized residuals to continuous prompts for few-shot class incremental learning
- quantum-gated task-interaction knowledge distillation for pre-trained model-base
- quantvla scale-calibrated post-training quantization for vision-language-action | arXiv: 2602.20309
- qucnet quantum deep learning driven multi-circuit network for remote sensing ima
- query2uncertainty robust uncertainty quantification and calibration for 3d objec | arXiv: 2605.05328
- queryme query-driven open-vocabulary 3d object affordances grounding from multim
- queryocc query-based self-supervision for 3d semantic occupancy
- question-guided visual compression with memory feedback for long-term video unde | arXiv: 2603.15167
- r-4b incentivizing general-purpose auto-thinking in mllms via bi-mode annealing
- r-c2 cycle-consistent reinforcement learning improves multimodal reasoning
- r2-seg training-free ood medical tumor segmentation via anatomical reasoning and
- r2g multi view circuit graph benchmark suite from rtl to gdsii | arXiv: 2604.08810
- r2tua reconstruction-residual based targeted and untargeted attack against text-
- r3-pcqa ray-reprojection-reinforcement for no-reference 3d point cloud quality a
- r4 retrieval-augmented reasoning for vision-language models in 4d spatio-tempora
- r4-cgqa retrieval-based vision language models for computer graphics image quali
- r4det 4d radar-camera fusion for high-performance 3d object detection | arXiv: 2603.11566
- raas llm agentic system architecture search with grpo
- radar vq-vae decoder of var is a good student for restoring against degradation
- radar-guided polynomial fitting for metric depth estimation | arXiv: 2503.17182
- radiance meshes for volumetric reconstruction
- rag-tp a general framework for vehicle trajectory prediction via retrieval-augme
- rags unleashing 3d gaussian splatting from 4d radar and monocular cue for 3d obj
- raise requirement-adaptive evolutionary refinement for training-free text-to-ima | arXiv: 2603.00483
- ram recover any 3d human motion in-the-wild | arXiv: 2603.19929
- rank-guided pseudo-bias learning for robust black-box adaptation
- rankood - class ranking-based out-of-distribution detection
- rap fast feedforward rendering-free attribute-guided primitive importance score | arXiv: 2602.19753
- rapid reusing attention sparsity with inter-step adaptation for efficient video
- rar restore assess repeat a unified framework for iterative image restoration | arXiv: 2603.26385
- rascene high-fidelity 3d scene imaging with mmwave communication signals | arXiv: 2604.02603
- raven radar adaptive vision encoders for efficient chirp-wise object detection a
- raven radar adaptive vision encoders for efficient chirp-wise object detection a
- rawdomain degradation models smartphone sr | arXiv: 2603.12493
- rawmetadiff unlocking extreme darkness from dual-exposure raw with meta-guided d
- raynova scale-temporal autoregressive world modeling in ray space | arXiv: 2602.20685
- rc-nf robot-conditioned normalizing flow for real-time anomaly detection in robo | arXiv: 2603.11106
- rdf-mig a robust diffusion framework for masked image generation to augment sema
- rdface a benchmark dataset for rare disease facial image analysis under extreme | arXiv: 2604.03454
- rdvq differentiable vq image compression | arXiv: 2604.10546
- re-align structured reasoning-guided alignment for in-context image generation a
- re-evaluating continual vqa toward fair and robust evaluation for multimodal con
- re-vlm event-augmented vision-language model for scene understanding
- reading or reasoning format decoupled reinforcement learning for document ocr
- reading your actions learning generalizable action representations via pre-train
- reag reasoning-augmented generation for knowledge-based visual question answerin | arXiv: 2511.22715
- reagen adaptive generation of structured chains-of-thought for efficient multimo
- real iisr infrared image super resolution autoregressive | arXiv: 2603.04745
- real-time dynamic scene rendering with controlled compressibility and contact aw
- real-time generation of streamable talking portrait video with reference-guided | arXiv: 2606.01620
- real-time long horizon air quality forecasting via group-relative policy optimiz
- real-time multimodal fingertip contact detection via depth and motion fusion for
- real2edit2real generating robotic demonstrations via a 3d control interface | arXiv: 2512.19402
- real2sim2real retinaldepth-64k for depth estimation in posterior segment ophthal
- realappiance let high-fidelity appliance assets controllable and workable as ali
- realbirdid benchmarking bird species identification in the era of mllms
- realign generalizable image forgery detection via reasoning-aligned representati | arXiv: 2605.16080
- realiz3d 3d generation made photorealistic via domain-aware learning | arXiv: 2605.13852
- reallocating attention across layers to reduce multimodal hallucination | arXiv: 2510.10285
- realm mllm agent 3d reasoning gaussian | arXiv: 2510.16410
- realunify do unified models truly benefit from unification a comprehensive bench | arXiv: 2509.24897
- realvlg-r1 a large-scale real-world visual-language grounding benchmark for robo | arXiv: 2603.14880
- realworld point tracking with verifierguided pseud | arXiv: 2603.12217
- reartgs generalizable articulation reconstruction with temporal geometry constra
- reasoning diffusion for unpaired test time out-of-distribution text-image to vid
- reasoning palette modulating reasoning via latent contextualization for controll
- reasoning palette modulating reasoning via latent contextualization for controll
- reasoning-driven anomaly detection and localization with image-level supervision | arXiv: 2603.27179
- reasonmap towards fine-grained visual reasoning from transit maps | arXiv: 2505.18675
- reattnclip training-free open-vocabulary remote sensing image segmentation via r
- reattnclip training-free open-vocabulary remote sensing image segmentation via r
- rebrl reinforcing discrete visual diffusion models with rebalanced timestep cred
- recall recalibrating capability degradation for mllm-based composed image retrie | arXiv: 2602.01639
- recedit-drive 3d reconstruction-guided spatiotemporal video editing for autonomo
- reclaiming lost text layers for source-free cross-domain few-shot learning | arXiv: 2603.05235
- reconstructing spiking neural networks using a single neuron with autapses
- reconstruction-guided slot curriculum addressing object over-fragmentation in vi | arXiv: 2603.22758
- recover to predict progressive retrospective learning for variable-length trajec | arXiv: 2603.10597
- recovermark robust watermarking for localization and recovery of manipulated fac | arXiv: 2602.20618
- recs4r bridging semantics and geometry for referring remote sensing interpretati
- Rectifying Latent Space for Generative Single-Image Reflection Removal
- rectok reconstruction distillation along rectified flow
- recurrent reasoning with vision-language models for estimating long-horizon embo | arXiv: 2603.17312
- recurrent video masked autoencoders | 📄 paper_cache/CVPR2026/cvf-recurrent_video_masked_autoencoders.txt
- ref4d-videobench four-dimensional reference-based evaluation of text-to-video ge
- refacade editing object with given reference texture
- refact empowering multimodal web agents with visual and context focusing
- refer-agent a collaborative multi-agent system with reasoning and reflection for
- reflexsplit single image reflection separation via layer fusion-separation | arXiv: 2601.17468
- reflow self-correction motion learning for dynamic scene reconstruction
- refracting reality generating images with realistic transparent objects | arXiv: 2511.17340
- reframing long-tailed learning via loss landscape geometry | arXiv: 2603.21217
- refta breaking the weight reconstruction bottleneck in tensorized parameter-effi
- refton reference person shot assist virtual try-on | arXiv: 2511.00956
- regenhoi unifying reconstruction and generation for 3d human-object interaction
- regformer transferable relational grounding for efficient weakly-supervised huma
- regformer transferable relational grounding for weakly-supervised hoi detection
- region-adaptive sampling for diffusion transformers
- region-wise correspondence prediction between manga line art images
- regionfuse region-adaptive pixel distribution learning for infrared and visible
- regionroute regional style transfer with diffusion model
- registration-free learnable multi-view capture of faces in dense semantic corres | arXiv: 2605.01450
- regulating rather than constraining adaptive guidance for complex spectral recon
- rehearsevla simulated post-training for vlas with physically-consistent world mo | arXiv: 2509.24948
- rehyat recurrent hybrid attention for video diffusion transformers
- reinforcement-guided synthetic data generation for privacy-sensitive identity re
- rejection mixing fast semantic propagation of mask tokens for efficient dllm inf
- rel-sf4pass panoramic semantic segmentation with rel depth representation and sp | arXiv: 2601.16788
- rel-zero harnessing patch-pair invariance for robust zero-watermarking against a | arXiv: 2603.17531
- relags relational language gaussian splatting | arXiv: 2603.17605
- relational visual similarity | arXiv: 2512.07833
- relax reasoning with latent exploration for large reasoning models
- reliable policy transfer for safety-aware end-to-end driving with deep reinforce
- reliev3r relieving feed-forward 3d reconstruction from multi-view geometric annot | arXiv: 2604.00548
- relightable holoported characters capturing and relighting dynamic human perform
- RelightAnyone: A Generalized Relightable 3D Gaussian Head Model
- relightful video portrait harmonization
- remannet a riemannian manifold network for monocular 3d lane detection
- rematch boosting representation through matching for multimodal retrieval
- remedying target-domain astigmatism for cross-domain few-shot object detection | arXiv: 2603.18541
- remogen real-time human interaction-to-reaction generation via modular learning | arXiv: 2604.01082
- remora multimodal large language model based on refined motion representation fo | arXiv: 2602.16412
- remot reinforcement learning with motion contrast triplets | arXiv: 2603.00461
- remote sensing image super-resolution for imbalanced textures a texture-aware di
- render-to-adapt unsupervised personal adaptation for gaze estimation
- renderflow single-step neural rendering via flow matching | arXiv: 2601.06928
- reparameterized tensor ring functional decomposition for multi-dimensional data | arXiv: 2603.01034
- representation-steered incremental adapter-tuning for class-incremental learning
- representing 3d faces with learnable b-spline volumes
- resad normalized residual trajectory modeling for end-to-end autonomous driving
- resam refine requery and reinforce self-prompting point-supervised segmentation
- rescene4d temporally consistent semantic instance segmentation of evolving indoo | arXiv: 2601.11508
- residual decoder adapter id-preserving tokenizer adaption for autoregressive tex | arXiv: 2606.01911
- residual decoding mitigating hallucinations in large vision-language models via | arXiv: 2602.01047
- resihmr residual-limb aware single-image 3d human mesh recovery for individuals | arXiv: 2604.28025
- resolving evidence sparsity agentic context engineering for long-document unders
- resolving the identity crisis in text-to-image generation | arXiv: 2510.01399
- resolving the stability-plasticity dilemma in reinforcement learning via complem
- restore assess repeat a unified framework for iterative image restoration
- restore text first enhance image later two-stage scene text image super-resoluti
- retformer multimodal retrieval for enhancing image recognition
- Rethinking 2D-3D Registration: A Novel Network for High-Value Zone Selection and Representation Consistency Alignment
- rethinking camera choice an empirical study on fisheye camera properties in robo | arXiv: 2603.02139
- rethinking concept bottleneck models from pitfalls to solutions | arXiv: 2603.05629
- rethinking dataset distillation hard truths about soft labels | arXiv: 2604.18811
- rethinking diffusion model-based video super-resolution leveraging dense guidanc
- rethinking glyph spatial information in font generation
- rethinking knowledge transfer in image quality assessment a perceptual preferenc
- rethinking mllm itself as a segmenter with a single segmentation token | arXiv: 2603.19026
- rethinking model selection in vlm through the lens of gromov-wasserstein distanc | arXiv: 2605.01325
- rethinking pose refinement in 3d gaussian splatting under pose prior and geometr | arXiv: 2603.16538
- rethinking position embedding as a context controller for multi-reference and mu | arXiv: 2604.03738
- rethinking snn online training and deployment grad | arXiv: 2410.07547
- rethinking two-stage referring-by-tracking in referring multi-object tracking ma | arXiv: 2503.07516
- rethinking umm visual generation masked modeling for efficient image-only pre-tr
- retimegs continuous-time reconstruction of 4d gaussian splatting | arXiv: 2603.13783
- retouchiq mllm agents for instruction-based image retouching with generalist rew
- retrieve and segment are a few examples enough to bridge the supervision gap in
- retrieve-to-restore efficient all-in-one image restoration with a retrieval-base
- retrieving counterfactuals improves visual in-context learning | arXiv: 2603.16737
- revinn an end-to-end invertible neural network for reversible adversarial exampl
- revisiting 2d foundation models for scalable 3d medical image classification
- revisiting 3d reconstruction kernels as low-pass filters
- revisiting f-measure optimization in multi-label classification a sampling-based
- revisiting geometric obfuscation with dual convergent lines for privacy-preservi | arXiv: 2604.22310
- revisiting model stitching in the foundation model | arXiv: 2603.12433
- revisiting monocular slam with spatio-temporal scene modeling
- revisiting pose sensitivity in splat-based computed tomography under sparse-view
- revisiting sparsity constraint under high-rank property in partial multi-label l
- revisiting the necessity of full accuracy weakly supervised object-level offset
- revisiting the necessity of lengthy chain-of-thought in vision-centric reasoning
- revisiting token compression for accelerating vit-based sparse multi-view 3d obj
- revisiting unknowns towards effective and efficient open-set active learning | arXiv: 2603.07898
- revisiting visual corruptions in lvlms a shape-texture perspective on model fail
- revisor beyond textual reflection towards multimodal introspective reasoning in
- revive 3d refinement via encoded voluminous inflated prior for volume enhancemen | arXiv: 2604.27504
- reviving convnext for efficient convolutional diffusion models | arXiv: 2603.09408
- reward forcing efficient streaming video generation with rewarded distribution m
- reward sharpness-aware fine-tuning for diffusion models
- rewardflow generate images by optimizing what you reward | arXiv: 2604.08536
- reweaver towards simulation-ready and topology-accurate garment reconstruction | arXiv: 2601.16672
- rewis3d reconstruction improves weaklysupervised s | arXiv: 2603.06374
- rf4dneural radar fields for novel view synthesis in outdoor dynamic scenes
- rfdm residual flow diffusion models for video editing
- rgb-event based pedestrian attribute recognition a benchmark dataset and an asym
- RHCNet: Residual-Guided Hierarchical Calibration Network for Robust Underwater Object Detection
- rhino reconstructing human interactions with novel objects from monocular videos | arXiv: 2605.17014
- rho robust holistic osm-based metric cross-view geo-localization | arXiv: 2603.27758
- riskprop collision-anchored self-supervised risk propagation for early accident | arXiv: 2603.27165
- rl-scaniqa reinforcement-learned scanpaths for blind 360deg image quality assess
- RL-ScanIQA: Reinforcement-Learned Scanpaths for Blind 360° Image Quality Assessment | arXiv: 2603.14297
- rlftsim realistic and controllable multi-agent traffic simulation via reinforcem | arXiv: 2605.19033
- rmae-progress advancing semantic segmentation in unstructured environments
- rmir a benchmark dataset for reasoning-intensive multimodal image retrieval
- rned rotary number encoding and decoding for medical vlms
- rng a unified transformer for complete 3d modeling from partial observations | arXiv: 2603.01194
- rnn as linear transformer a closer investigation into representational potential
- roadgie towards a global-scale aerial benchmark for generalizable interactive ro
- robo-sgg exploiting layout-oriented normalization and restitution can improve ro
- roboagent chaining basic capabilities for embodied task planning | arXiv: 2604.07774
- robotseg a model and dataset for segmenting robots in image and video | arXiv: 2511.22950
- robowheel a data engine from real-world human demonstrations for cross-embodimen
- robust remote sensing image-text retrieval with noisy correspondence
- robust3dgsw toward robust watermarking for quantization-aware 3d gaussian splatt
- robustness under data scarcity few-shot continual adversarial training for evolv
- robustvisrag causality-aware vision-based retrieval-augmented generation under v | arXiv: 2602.22013
- role-synthclip a role-play driven diverse synthetic data approach
- romo a large-scale richly organized dataset and semantic taxonomy for human moti
- roots beneath the cut uncovering the risk of concept revival in pruning-based un
- rosamdepth robust self-supervised depth estimation leveraging segment anything m
- rose rotate your large language model to see
- rosetta stone for unified mllms a unified tokenizer to decipher understanding an
- rotation invariant and symmetry aware pixel difference network for remote sensin
- rounded or streamlined head bridging concept bottleneck models and attribute-des
- routing on demand dsnet for efficient progressive point cloud denoising
- rpgfusion 4d radar prior-guided multi-modal fusion for 3d detection
- rppg vqa video quality assessment | arXiv: 2604.11156
- rs-ssm refining forgotten specifics in state space model for video semantic segm | arXiv: 2603.24295
- rt-splatting joint reflection-transmission modeling with gaussian splatting | arXiv: 2605.18263
- rxncaption reformulating reaction diagram parsing as visual prompt guided captio
- s2-mllm boosting spatial reasoning capability of mllms for 3d visual grounding w
- s2am3d scale-controllable part segmentation of 3d point cloud | arXiv: 2512.00995
- s2c2seg semantic-spatial consistency and category optimization for open-vocabula
- s2d selective spectral decay for quantization-friendly conditioning of neural ac
- s2d sparse to dense lifting for 3d reconstruction with minimal inputs
- s2ft parameter-efficient fine-tuning in sparse spectrum domain | arXiv: 2605.08589
- saber spatially consistent 3d universal adversarial objects for bev detectors | arXiv: 2505.22499
- safedrive fine-grained safety reasoning for end-to-end driving in a sparse world | arXiv: 2602.18887
- safegrpo self-rewarded multimodal safety alignment via rule-governed policy opti
- safelogo turning your logos into jailbreak shields via micro-regional adversaria
- saferope risk-specific head-wise embedding rotation for safe generation in recti
- sage scalable agentic 3d scene generation for embodied ai
- sage style-adaptive generalization for privacy-constrained semantic segmentation
- sage training smart any-horizon agents for long video reasoning with reinforceme
- saido generalizable detection of ai-generated images via scene-aware and importa
- sail similarity-aware guidance and inter-caption augmentation-based learning for | arXiv: 2603.05437
- saliency-guided representation with consistency policy learning for visual unsup
- saliency-r1 enforcing interpretable and faithful vision-language reasoning via s | arXiv: 2604.04500
- salmubench a benchmark for sensitive association-level multimodal unlearning | arXiv: 2603.26316
- same attention different truths put logit-lens over visual attention to detect a
- same content different answers cross-modal inconsistency in mllms | arXiv: 2512.08923
- same or not enhancing visual perception in vision-language models
- same sparse and anchored model editing for heterogeneous incremental learning un
- samix reinforcing sam2 with semantic adapter and reference selecting policy for
- saner switchable adapter with non-parametric enhanced routing for person de-reid
- sapave active perception manipulation vla roboti | arXiv: 2603.12193
- saqn semantic-based adaptive query network for 3d referring expression segmentat
- sar2net learning spatially anchored representations for retrieval-guided cross-s
- sarl-stg a spatially aware reinforcement learning framework for refining mllms i | 📄 paper_cache/CVPR2026/cvf-sarl-stg_a_spatially_aware_reinforcement.txt
- sarmae masked autoencoder for sar representation learning | arXiv: 2512.16635
- sat-rrg llm-guided self-adaptive training for radiology report generation with t
- sattc structure-aware label-free test-time calibration for cross-subject eeg-to- | arXiv: 2603.20738
- savax egotoexo imitation error detection via scene | arXiv: 2603.12764
- save speech-aware video representation learning for video-text retrieval | arXiv: 2603.08224
- say cheese detail-preserving portrait collection generation via natural language
- scal3r scalable test-time training for large-scale 3d reconstruction
- scalable multi-view subspace clustering with tensorized anchor guidance
- scalable object relation encoding for better 3d spatial reasoning in large langu | arXiv: 2603.24721
- scalable trajectory generation for whole-body mobile manipulation
- scale space diffusion
- scaling dense event-stream pretraining from visual foundation models
- Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
- scaling multi-identity consistency for image customization via multi-to-multi ma
- scaling parallel sequence models to vision foundation models
- scaling self-supervised and cross-modal pretraining for volumetric ct transforme
- scaling spatial intelligence with multimodal foundation models | arXiv: 2511.13719
- scaling test-time robustness of vision-language models via self-critical inferen | arXiv: 2603.07659
- scaling the long video understanding of multimodal large language models via vis | arXiv: 2603.29252
- scaling view synthesis transformers | arXiv: 2602.21341
- scaling zero-shot reference-to-video generation
- scaling-aware data selection for end-to-end autonomous driving systems | arXiv: 2604.08366
- scaling4d pushing the frontier of video novel view synthesis through large-scale
- scan clusters not pixels a cluster-centric paradigm for efficient ultra-high-def
- scapo self-supervised category-level articulated pose estimation from a single 3
- sce-depth a spherical compound eye framework for wide fov depth estimation
- sce-slam scale-consistent monocular slam via scene coordinate embeddings
- scemos scene-aware 3d human motion synthesis by planning with geometry-grounded
- scene grounding in the wild | arXiv: 2603.26584
- scene-vlm multimodal video scene segmentation via vision-language models | arXiv: 2512.21778
- scenemaker open-set 3d scene generation with decoupled de-occlusion and pose est
- scenes as tokens multi-scale normal distributions transform tokenizer for genera
- scenescribe-1m a large-scale video dataset with comprehensive geometric and sema | arXiv: 2604.07990
- scenetok a compressed diffusable token space for 3d scenes
- scieducator scientific video understanding and educating via deming-cycle multi-
- scieval evaluating and benchmarking the faithfulness of scientific image generat
- scone bridging composition and distinction in subject-driven image generation vi | arXiv: 2512.12675
- score salience-coverage reduction for vision token pruning in vision-language mo
- score2instruct scaling up video quality-centric instructions via automated dimen | arXiv: 2506.21011
- sd fsmis adapting stable diffusion for few shot medical image segmentation | arXiv: 2604.03134
- sddf specificity-driven dynamic focusing for open-vocabulary camouflaged object | arXiv: 2603.26109
- sdgs spatial difference guided gaussian splatting for simultaneous localization
- sduie semi-supervised diffusion for underwater image enhancement with quant-text
- se3-equivariance with geometric and topological guidance for category-level obje
- sea evaluating sketch abstraction efficiency via element-level commonsense visua
- sea-flow3d simplified efficient and accurate scene flow via spatial vector sampl
- sea-vision a multilingual benchmark for comprehensive document and scene text un | arXiv: 2603.15409
- searchad large-scale rare image retrieval dataset for autonomous driving | arXiv: 2604.08008
- season mitigating temporal hallucination in video large language models via self
- seatrack multimodal tracker | arXiv: 2604.12502
- secos semantic capture for rigorous classification in open-world semi-supervised | arXiv: 2604.27596
- sed-ud an influence-driven and hierarchically-decoupled information bottleneck f
- see and fix the flaws enabling vlms and diffusion models to comprehend visual ar
- see further think deeper advancing vlms reasoning ability with low-level visual
- see it say it sorted an iterative training-free framework for visually-grounded | arXiv: 2602.21497
- see less see right bi-directional perceptual shaping for multimodal reasoning
- see think act teaching multimodal agents to effectively interact with gui by ide | arXiv: 2509.13615
- see through the noise improving domain generalization in gaze estimation | arXiv: 2604.16562
- see what i mean aligning vision and language representations for video fine-grai
- see what we cannot see a geo-guided reasoning benchmark for object counting unde
- seegroup multi-layer depth estimation of transparent surfaces via self-determine
- seeing as experts do a knowledge-augmented agent for open-set fine-grained visua
- seeing beyond 8bits subjective and objective quality assessment of hdr-ugc video
- seeing beyond extrapolative domain adaptive panoramic segmentation | arXiv: 2603.15475
- seeing both sides towards bidirectional semantic alignment for open-vocabulary c
- seeing clearly reasoning confidently plug-and-play remedies for vision language | arXiv: 2602.19615
- seeing depth through frequency and motion a progressive training paradigm for mo
- seeing is improving visual feedback for iterative text layout refinement | arXiv: 2603.22187
- seeing motion through polarity for event-based action recognition
- seeing the scene matters revealing forgetting in video understanding models with | arXiv: 2603.27259
- seeing through blur tackling defocus in spike-based imaging
- seeing through boxes non-line-of-sight 3d reconstruction from radar signals
- seeing through light and darkness sensor-physics grounded deblurring hdr nerf fr
- seeing through the noise improving infrared small target detection and segmentat
- seeing through the shift causality-inspired robust generalized category discover
- seeing through touch tactile localization | arXiv: 2604.11579
- seeing what matters a training-free self-guided framework for multimodal detail
- seeing without pixels perception from camera trajectories | arXiv: 2511.21681
- seele a unified acceleration framework for real-time gaussian splatting on mobil
- seethrough3d occlusion aware 3d control in text-to-image generation | arXiv: 2602.23359
- seeu seeing the unseen world via 4d dynamics-aware generation | arXiv: 2512.03350
- segcompass exploring interpretable alignment with sparse autoencoders for enhanc | arXiv: 2605.22658
- segearth-r2 towards comprehensive language-guided segmentation for remote sensin
- seggbc justifiable coarse-to-fine granular-ball computing for enhancing clusteri
- segmo co-designing content-aware sparsity and locally-cohesive segment paralleli
- segmote token-level mixture of experts for medical image segmentation
- segquant a semantics-aware and generalizable quantization framework for diffusio | arXiv: 2507.14811
- select hypothesize and verify towards verified neuron concept interpretation | arXiv: 2603.24953
- select less reason more prioritizing evidence purity for video reasoning
- selection-as-nonlinearity bridging attention and activation via a joint game-dec
- selective amnesia using contrastive subnet erasure for class level unlearning in
- selective regularized and calibrated harnessing vision foundation models for cro | arXiv: 2605.19340
- Selectively Extracting and Injecting Visual Attributes into Text-to-Image Models
- selectkd selective token-weighted knowledge distillation for llms
- self-consistency for llm-based motion trajectory generation and verification | arXiv: 2603.29301
- self-corrected image generation with explainable latent rewards | arXiv: 2603.24965
- self-diffusion driven blind imaging
- self-evaluation unlocks any-step text-to-image generation
- self-supervised dynamic heterogeneous degradation modeling for unified zero-shot
- selfhvd self-supervised handheld video deblurring | arXiv: 2508.08605
- selfi self-improving reconstruction engine via 3d geometric feature alignment
- semantic audio-visual navigation in continuous environments | arXiv: 2603.19660
- semantic derivative flow graph-guided diffusion for controllable instance intera
- semantic foam unifying spatial and semantic scene decomposition | arXiv: 2604.26262
- semantic noise reduction via teacher-guided dual-path audio-visual representatio
- semantic-guided global-local collaborative prompt learning for few-shot class in
- semantics lead the way harmonizing semantic and texture modeling with asynchrono
- semanticvla towards semantic reasoning over action memorization via synergistic
- semi-supervised conformal prediction with unlabeled nonconformity score | arXiv: 2505.21147
- semi-supervised echocardiography video segmentation via anchor semantic awarenes
- semigda generative dual-distribution alignment for semi-supervised medical image | arXiv: 2604.23274
- semlayer semantic-aware generative segmentation and layer construction for abstr | arXiv: 2603.24039
- semlt3d semantic-guided expert distillation for camera-only long-tailed 3d objec | arXiv: 2604.18476
- semvideo reconstructs what you watch from brain activity via hierarchical semant
- sensesearch empowering vision-language models with high-resolution agentic searc
- sensor2sensor cross-embodiment sensor conversion for autonomous driving | arXiv: 2605.22809
- sepatch3d revisiting token compression for accelerating vit based sparse 3d detectors | arXiv: 2604.14563
- sfr-net steering-fusion-refining network in multi-label zero-shot sewer defect d
- sg-lora semantic-guided lora parameters generation
- sgad-slam splatting gaussians at adjusted depth for better radiance fields in rg | arXiv: 2603.21055
- sgdrive scene-to-goal hierarchical world cognition for autonomous driving
- sgi structured 2d gaussians for efficient and compact large image representation | arXiv: 2603.07789
- sgnlf spectralgeometric neural fields for posefre | arXiv: 2603.12903
- sgs-intrinsic semantic-invariant gaussian splatting for sparse-view indoor invers | arXiv: 2603.27516
- sgsoft learning fused semantic-geometric features for 3d shape correspondence vi
- shands a multi-view dataset and benchmark for surgical hand-gesture and error re
- shape structure-aware hierarchical unsupervised domain adaptation with plausibil
- shape-of-you fused gromov-wasserstein optimal transport for semantic corresponde | arXiv: 2603.11618
- shaper robust conditional 3d shape generation from casual captures
- sharp short-window streaming for accurate and robust prediction in motion foreca | arXiv: 2603.28091
- sharptimegs sharp and stable dynamic gaussian splatting via lifespan modulation
- shedding light on vln robustness a black-box framework for indoor lighting-based
- shelfocc native 3d supervision beyond lidar for vision-based occupancy estimatio
- shiftlut spatial shift enhanced look-up tables for efficient image restoration | arXiv: 2603.00906
- shotdirector directorially controllable multi-shot video generation with cinemat
- showtable unlocking creative table visualization with collaborative reflection a | arXiv: 2512.13303
- showui-p flow-based generative models as gui dexterous hands
- siglino efficient multi-teacher distillation for agglomerative vision foundation
- sigma a physics-based benchmark for gas chimney understanding in seismic images
- signpr a progressive vector-quantized diffusion framework for sign language prod
- similarity-as-evidence calibrating overconfident vlms for interpretable and labe | arXiv: 2602.18867
- similarity-consistent likelihood diffusion enables hidden person detection from
- simlbr learning to detect fake images by learning to detect real images | arXiv: 2602.20412
- simpact simulation-enabled action planning using vision-language models | arXiv: 2512.05955
- simple but effective triplet-based compression strategies for compact visual loc
- simple-vilmedsam simple text prompts meet vision-language models for medical ima
- simpleposter a simple baseline for product poster generation | arXiv: 2605.08784
- simscale learning to drive via real-world simulation at scale | arXiv: 2511.23369
- simspine a biomechanics-aware simulation framework for 3d spine motion annotatio
- sineproject machine unlearning for stable vision language alignment | arXiv: 2511.18444
- singeo unlock single models potential for robust cross-view geo-localization
- sir structured image representations for explainable robot learning
- sjd-pac accelerating speculative jacobi decoding via proactive drafting and adap | arXiv: 2603.18599
- skeletoncontext skeleton-side context prompt learning for zero-shot skeleton-bas | arXiv: 2603.29692
- sketch2colab sketch-conditioned multi-human animation via controllable flow dist | arXiv: 2603.02190
- sketch2ct multimodal diffusion for structure-aware 3d medical volume generation
- sketchassist a practical assistant for semantic edits and precise local redrawin
- sketchdeco training-free latent composition for precise sketch colourisation | arXiv: 2405.18716
- sketchfacegs real-time sketch-driven face editing and generation with gaussian s | arXiv: 2604.19202
- sketchrevive fine-grained pixel-to-vector sketch completion with diffusion-prior
- sketchvl policy optimization via fine-grained credit assignment for chart unders
- skullptor high fidelity 3d head reconstruction in seconds with multi-view normal
- sky2ground a benchmark for site modeling under varying altitude | arXiv: 2603.13740
- skyreels-text fine-grained font-controllable text editing for poster design | arXiv: 2511.13285
- skysense-vita towards universal in-context segmentation of multi-modal remote se
- slideredit continuous image editing with fine-grained instruction control
- slvmeval synthetic meta evaluation benchmark for text-to-long video generation | arXiv: 2603.29186
- small object great challenge a benchmark for small object visual grounding
- smap semantic route planning with map-grounded multimodal alignment
- smart replay adaptive scheduling of memory rehearsal for computational resource-
- smoes soft modality-guided expert specialization in moe-vlms | arXiv: 2604.23996
- smokesvd smoke reconstruction from a single view via progressive novel view synt
- smoothing the score function to enhance generalization in diffusion models
- smrabooth subject and motion representation alignment for customized video gener
- smv-ear bring spatiotemporal multi-view representation learning into efficient e
- smvrt implicit human 3d modeling
- so-bench a structural output evaluation of multimodal llm
- so3-equivariant vit-adapter for data-efficient zero-shot sim-to-real indoor pano
- socialnav training human-inspired foundation model for socially-aware embodied n
- socratic-geo synthetic data generation and cross-modal geometric reasoning via m
- soft modality-guided expert specialization in moe-vlms
- solace self confidence rewards t2i | arXiv: 2603.00918
- solireward mitigating susceptibility to reward hacking and annotation noise in v
- solving a nonlinear blind inverse problem for tagged mri with physics and deep g | arXiv: 2603.00882
- sonoworld from one image to a 3d audio-visual scene | arXiv: 2603.28757
- sope spherical coordinate-based positional embedding for enhancing spatial perce | arXiv: 2602.22716
- sope spherical positional encoding 3d lvlm | arXiv: 2602.22716
- sota self-adaptive optimal transport for zero-shot classification with multiple
- soul breathe life into digital human for high-fidelity long-term multimodal anim
- souple enhancing audio-visual localization and segmentation with learnable promp | arXiv: 2603.22732
- space-time forecasting of dynamic scenes with motion-aware gaussian grouping
- spacetools tool-augmented spatial reasoning via double interactive rl | arXiv: 2512.04069
- span spatial-projection alignment for monocular 3d object detection | arXiv: 2511.06702
- spar single-pass any-resolution vit for open-vocabulary segmentation | arXiv: 2604.02252
- sparrow learning spatial precision and temporal referential consistency in pixel | arXiv: 2603.12382
- sparse spectral lora routed experts for medical vlms
- sparse task vector mixup wsi prognosis | arXiv: 2603.10526
- sparse-lavida sparse multimodal discrete diffusion language models
- sparsecam4d spatio-temporally consistent 4d reconstruction from sparse cameras | arXiv: 2603.26481
- sparsely timing the change a spiking temporal framework for remote sensing inter
- sparsesplat towards applicable feed-forward 3d gaussian splatting with pixel-una
- sparsity as a key unlocking new insights from latent structures for out-of-distr | arXiv: 2604.26409
- sparsity-aware voxel attention and foreground modulation for 3d semantic scene c | arXiv: 2604.05780
- sparvar exploring sparsity in visual autoregressive modeling for training-free a | arXiv: 2602.04361
- spatia video generation with updatable spatial memory
- spatial retrieval augmented autonomous driving
- spatial-aware vla pretraining through visual-physical alignment from human video
- spatial-frequency collaborative learning for occluded visible-infrared person re
- spatial-sam spatially consistent 3d electron microscopy segmentation with sdf me
- spatial-ssrl enhancing spatial understanding via self-supervised reinforcement l | arXiv: 2510.27606
- spatialqa a benchmark for evaluating spatial logical reasoning in vision-languag | arXiv: 2602.20901
- spatialreward verifiable spatial reward modeling for fine-grained spatial consis
- spatialscore towards comprehensive evaluation for spatial intelligence | arXiv: 2505.17012
- spatialstack layered geometry-language fusion for 3d vlm spatial reasoning | arXiv: 2603.27437
- spatialtree how spatial intelligence branches out in mllms
- spatio-temporal conditional denoising transformer for modality-missing rgbt trac
- spatio-temporal difference guided motion deblurring with the complementary visio
- spatiotemporal pyramid flow matching for climate emulation
- spdmark selective parameter displacement for robust video watermarking | arXiv: 2512.12090
- spe-bevhead rethinking the detection head design for birds-eye-view object detec
- spe-mvs spatial position encoding enhanced multi-view stereo with monocular dept
- spear-1 scaling beyond robot demonstrations via 3d understanding
- specificity-aware reinforcement learning for fine-grained open-world classificat | arXiv: 2603.03197
- spectral conformal risk control distribution-free tail guarantees via bayesian q
- spectral scalpel amplifying adjacent action discrepancy via frequency-selective
- spectral super-resolution via adversarial unfolding and data-driven spectrum reg | arXiv: 2603.00920
- spectrally distilled representations aligned with instruction-augmented llms for
- speede3dgs speedy deformable 3d gaussian splatting with temporal pruning and mot
- speediff scalable pixel-anchored end-to-end latent diffusion model
- speeding up the learning of 3d gaussians with much shorter gaussian lists | arXiv: 2603.09277
- spegc continual test-time adaptation via semantic-prompt-enhanced graph clusteri | arXiv: 2603.11492
- spherical voronoi directional appearance as a differentiable partition of the sp
- spike-driven discrete aggregation for event-based object detection
- spiketrack a spike-driven framework for efficient visual tracking | arXiv: 2602.23963
- spiketrack high-performance and energy-efficient event-based object tracking wit
- spiraldiff spiral diffusion with lora for rgb-to-raw conversion across cameras | arXiv: 2603.14885
- spk2vidnet a hierarchical recurrent architecture for high-fidelity video reconstr
- splat-based metal artifact reduction in cone-beam ct via compact attenuation mod
- splatent splatting diffusion latents for novel view synthesis | arXiv: 2512.09923
- splatsure selective super-resolution for multi-view consistent 3d gaussian splat
- spot spatiotemporal prompt optimization for motion-stabilized mllm-guided video
- spot spatiotemporal prompt optimization for motion-stabilized mllm-guided video
- spot the ball a benchmark for visual social inference
- sr3r rethinking super-resolution 3d reconstruction with feed-forward gaussian sp | arXiv: 2602.24020
- sra 2 variational autoencoder self-representation alignment for efficient diffus
- srpo self-referential policy optimization for vision-language-action models
- ssm-aware token-efficient vmamba via adaptive patch pruning and merging for pers
- st4r-splat spatio-temporal referring segmentation in 4d gaussian splatting
- stabilizing feature geometry in noisy pretrained models for robust downstream ta
- stable and efficient single-rollout rl for multimodal reasoning
- Stable Mean Flow: Lyapunov-Inspired One-Step Flow Matching
- stable spike dual consistency optimization via bitwise and operations for spikin | arXiv: 2603.11676
- stablemtl repurposing latent diffusion models for multi-task learning from parti | arXiv: 2506.08013
- stac plug-and-play spatio-temporal aware cache compression for streaming 3d reco | arXiv: 2603.20284
- stake the points structure-faithful instance unlearning | arXiv: 2603.12915
- stamo unsupervised learning of generalizable robot motion from compact state rep
- stand-in a lightweight and plug-and-play identity control for video generation
- star test-time adaptation can enhance universal prompt learning for vision-langu
- star-kvqa structured reasoning traces for implicit-knowledge visual question ans
- star-r1 multi-view spatial transformation reasoning by reinforcing multimodal ll
- starflow-v end-to-end video generative modeling with autoregressive normalizing
- statistical characteristic-guided denoising for rapid high-resolution transmissi | arXiv: 2603.18834
- stavatar soft binding and temporal density control for monocular 3d head avatars | arXiv: 2511.19854
- stay in your lane role specific queries with overlap suppression loss for dense | arXiv: 2603.11439
- stcast adaptive boundary alignment for global and regional weather forecasting | arXiv: 2509.25210
- stcdit spatio-temporally consistent diffusion transformer for high-quality video
- stcdit spatio-temporally consistent diffusion transformer for high-quality video
- steering where to diffuse generative modeling of phenotypic response simulation
- stepwise credit assignment for grpo on flow-matching models
- stitch semantic transition and transportation in collaboration for training-free
- Stochastic Ray Tracing for the Reconstruction of 3D Gaussian Splatting
- store3d sparse token relevance in vits for efficient multi-view 3d object detect | arXiv: 2605.14110
- storytailora zero-shot pipeline for action-rich multi-subject visual narratives | arXiv: 2602.21273
- streamavatar streaming diffusion models for real-time interactive human avatars | arXiv: 2512.22065
- streaming diffusion model for fast infrared and visible video fusion
- streaming video instruction tuning
- streamlined knowledge distillation
- streamlined open-vocabulary human-object interaction detection
- streamrag enhancing real-time video understanding with retrieval augmentation
- streamready learning what to answer and when in long streaming videos | arXiv: 2603.08620
- streamvlo streaming visual-lidar odometry with cumulative drift compensation
- strnet visual navigation with spatio-temporal representation through dynamic gra | arXiv: 2604.02829
- stronger normalization-free transformers | arXiv: 2512.10938
- structural action transformer for 3d dexterous manipulation
- structural graph probing of vision-language models
- structure-aware representation distillation for tiny-dense object segmentation
- structure-to-intensity diffusion for adverse-weather lidar generation
- structxlip enhancing vision-language models with multimodal structural cues | arXiv: 2602.20089
- style-grpo semantic-aware preference optimization for image style transfer guide
- stylegallery training-free and semantic-aware personalized style transfer from a
- subflot submodel extraction for efficient and personalized federated learning vi | arXiv: 2604.06631
- submodel extraction for efficient and personalized federated learning via optima
- subspace alignment for clip-based continual learning via canonical correlation a
- subspacead training-free few-shot anomaly detection via subspace modeling | arXiv: 2602.23013
- sunfaded illumination-aware gaussian splatting for dark scenes with camera-mount
- sup sub-cloud driven point cloud registration
- superman unifying skeleton and vision for human motion perception and generation
- suppressing non-semantic noise in masked image modeling representations | arXiv: 2604.00172
- surf signature-retained fast video generation
- surgcot advancing spatiotemporal reasoning in surgical videos through a chain-of
- sv-gs sparse view 4d reconstruction with skeleton-driven gaussian splatting
- SVBench: Evaluation of Video Generation Models on Social Reasoning
- svhalluc benchmarking speech-vision hallucination in audio-visual large language | arXiv: 2606.02642
- swift sliding window reconstruction for few-shot training-free generated video a | arXiv: 2603.08536
- swifttailor efficient 3d garment generation with geometry image representation | arXiv: 2603.19053
- swiftvla unlocking spatiotemporal dynamics for lightweight vla models at minimal
- switchcraft training-free multi-event video generation with attention controls | arXiv: 2602.23956
- symphomotion joint control of camera motion and object dynamics for coherent vid | arXiv: 2604.03723
- syncdreamer controllable and expressive avatar generation beyond the talking hea
- synclip synonym-coherent language-image pretraining for robust open-vocabulary d
- synergistic bleeding region and point detection in laparoscopic surgical videos | arXiv: 2503.22174
- synmotion semantic-visual adaptation for motion customized video generation
- synthesizing visual concepts as vision-language programs
- synthetic curriculum reinforces compositional text-to-image generation
- synthetic object compositions for scalable and accurate learning in detection se
- t2sgrid temporal-to-spatial gridification for video temporal grounding
- tablemix enhancing multimodal table reasoning in mllms from a data-centric persp
- tackling alignment ambiguity in person retrieval through conversational attribut
- tackling model bias via game-theoretic multi-agent collaboration framework for h
- taco task-aware contrastive learning for joint lidar localization and 3d object
- tacsim a dataset and benchmark for football tactical style imitation | arXiv: 2603.25199
- tag-moe task-aware gating for unified generative mixture-of-experts | arXiv: 2601.08881
- tagsplat topology-aware gaussian splatting for dynamic mesh modeling and trackin | arXiv: 2512.01329
- taligndiff automatic tooth alignment assisted by diffusion-based transformation
- talk2move reinforcement learning for text-instructed object-level geometric tran
- talking together synthesizing co-located 3d conversations from audio | arXiv: 2603.08674
- talo pushing 3d vision foundation models towards globally consistent online reco | arXiv: 2512.02341
- talon test-time adaptive learning for on-the-fly category discovery | arXiv: 2603.08075
- tamer a tri-modal contrastive alignment and multi-scale embedding refinement fra
- taming generative diffusion model for task-oriented infrared imaging
- taming noise-induced prototype degradation for privacy-preserving personalized f | arXiv: 2604.27833
- taming preference mode collapse via directional decoupling alignment in diffusio | arXiv: 2512.24146
- taming sampling perturbations with variance expansion loss for latent diffusion | arXiv: 2603.21085
- taming the long tail rebalancing adversarial training via adaptive perturbation | arXiv: 2605.13395
- taming video models for 3d and 4d generation via zero-shot camera control | arXiv: 2509.15130
- tango learning distribution-wise foundation prior consistency and instance-wise
- tango text-anchored guided optimization for robust fine-tuning vision-language m
- tap a token-adaptive predictor framework for training-free diffusion acceleratio | arXiv: 2603.03792
- tape task-adaptive prototype evolution in audio-language models for fully few-sh
- tar token-aware refinement for fine-grained generalized category discovery
- target-aware invertible encoder with reconstruction guidance for infrared small
- tas-lora transformer architecture search with mixture-of-lora experts | arXiv: 2605.07256
- task-oriented data synthesis and control-rectify sampling for remote sensing sem | arXiv: 2512.16740
- taskforce cooperative multi-agent reinforcement learning for multi-task optimiza
- taskit memory-efficient fine-tuning of multi-lora llms via cross-task importance
- tavatar topology-aware gaussian attribute derivation for animatable human avatar
- taxonomy-aware representation alignment for hierarchical visual recognition with | arXiv: 2603.00431
- tc-pade trajectory-consistent pade approximation for diffusion acceleration
- tcei test time calibration experience intuition mot | arXiv: 2603.21629
- tco learning 3d reconstruction with priors in test time | arXiv: 2604.03878
- tdatr improving end-to-end table recognition via table detail-aware learning and | arXiv: 2603.22819
- teaching dinov3 about partial 3d geometry a self-supervised geometry-aware appro
- teamhoi learning a unified policy for cooperative human-object interactions with | arXiv: 2603.07988
- tear temporal-aware automated red-teaming for text-to-video models | arXiv: 2511.21145
- teflow enabling multi-frame supervision for self-supervised feed-forward scene f | arXiv: 2602.19053
- tehor text-guided 3d human and object reconstruction with textures | arXiv: 2602.19679
- tell model where to look mitigating hallucinations in mllms by vision-guided att | arXiv: 2511.20032
- tell2adapt a unified framework for source free unsupervised domain adaptation vi | arXiv: 2603.05012
- tempocontrol temporal attention guidance for text-to-video models
- tempomaster efficient long video generation via next-frame-rate prediction
- temporal imbalance of positive and negative supervision in class-incremental lea | arXiv: 2603.02280
- temporal inversion for learning interval change in chest x-rays
- temporal representation enhancement tre learning to forget dominant patterns for
- tempr1 improving temporal understanding of mllms via temporal-aware multi-task r
- terrascope pixel-grounded visual reasoning for earth observation
- terraseg self-supervised ground segmentation for any lidar | arXiv: 2603.27344
- tessera temporal embeddings of surface spectra for earth representation and anal
- test-time 3d occupancy prediction | arXiv: 2503.08485
- test-time alignment of text-to-image diffusion models via null-text embedding op
- test-time ego-exo-centric adaptation for action anticipation via multi-label pro | arXiv: 2603.09798
- test-time instance-specific parameter composition a new paradigm for adaptive ge | arXiv: 2603.27665
- test-time multi-prompt adaptation for open-vocabulary remote sensing image segme
- test-time perturbation tuning with delayed feedback for vision-language-action m
- test-time training for lidar semantic segmentation under corruption via geometri
- text-image conditioned 3d generation | arXiv: 2603.21295
- text-phase synergy network with dual priors for unsupervised cross-domain image | arXiv: 2603.12711
- text-printed image bridging the image-text modality gap for text-centric trainin
- textit4dsurf high-fidelity dynamic scene surface reconstruction | arXiv: 2603.28064
- textpecker rewarding structural anomaly quantification for enhancing visual text | arXiv: 2602.20903
- tf-cade foreground-concentrated text-video alignment for zero-shot temporal acti
- tf-ssd a strong pipeline via synergic mask filter for training-free co-salient o
- tgsformer scalable temporal gaussian splatting for embodied semantic scene compl
- tgt text-grounded trajectories for locally controlled video generation
- tgtrack temporal generative learning for unified single object tracking
- the coherence trap when mllm-crafted narratives exploit manipulated visual conte | arXiv: 2505.17476
- the consistency critic correcting inconsistencies in generated images via refere
- the devil is in attention sharing improving complex non-rigid image editing fait
- the devil is in gradient entanglement energy-aware gradient coordinator for robu
- the devil is in the details enhancing video virtual try-on via keyframe-driven d | arXiv: 2512.20340
- the drift kernel why diffusion models change even when told not to
- the geometry of robustness optimizing loss landscape curvature and feature manif | 📄 paper_cache/CVPR2026/cvf-the_geometry_of_robustness_optimizing_lo.txt
- the golden subspace where efficiency meets generalization in continual test-time | arXiv: 2603.21928
- the image as its own reward reinforcement learning with adversarial reward for i
- the invisible gorilla effect in out-of-distribution detection | arXiv: 2602.20068
- the llm bottleneck why open-source vision llms struggle with hierarchical visual | arXiv: 2505.24840
- the missing gap from solving square jigsaw puzzles to handling real world archae
- the missing point in vision transformers for universal image segmentation
- the more the merrier contrastive fusion for higher-order multimodal alignment | arXiv: 2511.21331
- the power of decaying steps enhancing attack stability and transferability for s | arXiv: 2602.19096
- the power of prior training-free open-vocabulary semantic segmentation with llav
- the road less seen segment exploration for weakly supervised video anomaly detec
- the sa-fari dataset segment anything in footage of animals for recognition and i
- the surprising effectiveness of noise pretraining for implicit neural representa | arXiv: 2603.29034
- the universal normal embedding | arXiv: 2603.21786
- thera thermal-aware visual-language prompting for controllable rgb-to-thermal in
- thermal diffusion matters infrared spatial-temporal video super-resolution throu
- thermal is always wild characterizing and addressing challenges in thermal-only
- thermal-det language-guided cross-modal distillation for open-vocabulary thermal
- thermally activated dual-modal adversarial clothing against ai surveillance syst | arXiv: 2511.09829
- think 360 evaluating the width-centric reasoning capability of mllms beyond dept | arXiv: 2603.22689
- think 360deg beyond depth evaluating the width-centric reasoning capability of m
- think then verify a hypothesis-verification multi-agent framework for long video | arXiv: 2603.04977
- think visually reason textually vision-language synergy in abstract reasoning
- think with 3d geometric imagination grounded spatial reasoning from limited view
- think-as-you-see streaming chain-of-thought reasoning for large vision-language
- think-then-generate structural chain-of-thought reasoning for consistent 3d gene
- thinking beyond labels vocabulary-free fine-grained recognition using reasoning-
- thinking diffusion penalize and guide visual-grounded reasoning in diffusion mul | arXiv: 2604.05497
- Thinking in 360°: Humanoid Visual Search in the Wild
- thinking in dynamics how multimodal large language models perceive track and rea | arXiv: 2603.12746
- thinking in uncertainty mitigating hallucinations in mlrms with latent entropy-a
- thinking with drafts speculative temporal reasoning for efficient long video und
- thinking with frames generative video distortion evaluation via frame reward mod
- thinking with frames generative video distortion evaluation via frame reward mod
- thinking with programming vision towards a unified view for thinking with images
- thinking with video video generation as a promising multimodal reasoning paradig
- thinking with videos multimodal tool-augmented reinforcement learning for long v
- thinking-while-generating interleaving textual reasoning throughout visual gener
- through the frequency lens cross-domain generalisable gaze estimation with adapt
- tiacam text-anchored invariant feature learning with auto-augmentation for camer | arXiv: 2602.18863
- tiger a unified framework for time images and geo-location retrieval | arXiv: 2603.24749
- tim temporal decoupling with iterative mutual-refinement model for longitudinal
- time blindness why video-language models cant see what humans can
- time without time pseudo-temporal representation for space-time super-resolution
- time-aware one step diffusion network for real-world image super-resolution
- timebridge self-supervised video representation learning via start-end joint emb
- timelens rethinking video temporal grounding with multimodal llms | arXiv: 2512.14698
- timeripples accelerating vdits by understanding the spatio-temporal correlations
- timeviper a hybrid mamba-transformer vision-language model for efficient long vi
- tina text-free inversion attack for unlearned text-to-image diffusion models | arXiv: 2603.17828
- tipsv2 patch text alignment | arXiv: 2604.12012
- tivibench benchmarking think-in-video reasoning for video generation
- tlma mitigating the impact of weakly labeled information for video anomaly detec
- tm-bsn triangular-masked blind-spot network for real-world self-supervised image | arXiv: 2604.04484
- token reduction via local and global contexts optimization for efficient video l | arXiv: 2603.01400
- token warping helps mllms look from nearby viewpoints | arXiv: 2604.02870
- tokengs decoupling 3d gaussian prediction from pixels with learnable tokens
- tokenhand discrete token representation for efficient hand mesh reconstruction
- tokenlight precise lighting control in images using attribute tokens | arXiv: 2604.15310
- tokensplat token-aligned 3d gaussian splatting for feed-forward pose-free recons
- tokentrace multi-concept attribution through watermarked token recovery
- too vivid to be real benchmarking and calibrating generative color fidelity | arXiv: 2603.10990
- topocl topological contrastive learning for medical imaging
- topohr hierarchical centerline representation for cyclic topology reasoning in d | arXiv: 2604.24119
- topology-aware feature propagation for unsupervised non-rigid point cloud corres
- topoma topology-guided multi-agent dense rgb 3d reconstruction via distributed i
- topomesh high-fidelity mesh autoencoding via topological unification | arXiv: 2603.24278
- toposlide topologically-informed histopathology whole slide image representation
- touchdream 3d object completion through imagined touch
- toward early quality assessment of text-to-image diffusion models
- toward generalizable whole brain representations with high-resolution light-shee | arXiv: 2603.29842
- towards an incremental unified multimodal anomaly detection augmenting multimoda
- towards balanced multi-modal learning in 3d human pose estimation | arXiv: 2501.05264
- towards calibrating prompt tuning of vision-language models | arXiv: 2602.19024
- towards cross-modal preservation consistency and alignment for privacy-preservin
- towards decompositional human motion generation with energy-based diffusion mode
- towards dynamic modality alignment in multimodal continual learning
- towards efficient medical reasoning with minimal fine-tuning data | arXiv: 2508.01450
- towards fine-grained attribution instance-aware preference optimization for alig
- towards foundation models for 3d scene understanding instance-aware self-supervi
- towards generalizable ai-generated image detection via image-adaptive prompt lea | arXiv: 2508.01603
- towards generalized representations for low-light understanding when signal cons
- towards gui agents vision-language diffusion models for gui grounding | arXiv: 2603.26211
- towards high-quality image segmentation improving topology accuracy by penalizin | arXiv: 2603.18671
- towards high-resolution and disentangled reference-based sketch colorization
- towards highly transferable vision-language attack via semantic-augmented dynami | arXiv: 2603.04839
- towards highly-constrained human motion generation with retrieval-guided diffusi
- towards holistic modeling for video frame interpolation with auto-regressive dif
- towards human-like robot handwriting via contour-aware generation
- towards intrinsic-aware monocular 3d object detection | arXiv: 2603.27059
- towards knowledge-augmented bayesian deep learning for computer vision
- towards motion turing test evaluating human-likeness in humanoid robots
- towards multimodal domain generalization with few labels | arXiv: 2602.22917
- towards open environments and instructions general vision-language navigation vi | arXiv: 2601.09111
- towards open-vocabulary industrial defect understanding with a large-scale multi
- towards persistence learning topological constraints for event-based small objec
- towards real-world document parsing via realistic scene synthesis and document-a | arXiv: 2603.23885
- towards reliable evaluation of adversarial robustness for spiking neural network
- towards robust multi-modal semantic segmentation with teacher-student framework
- towards robust sequential decomposition for complex image editing | arXiv: 2605.09233
- towards robust vision transformers path dependency analysis and a simple two-sta
- towards stable self-supervised object representations in unconstrained egocentri
- towards stealthy and effective backdoor attacks on lane detection a naturalistic
- towards uncertainty-aware unsupervised domain adaptation for videos and time-ser
- towards unified human perception and machine understanding token flow guided com
- towards visual query localization in the 3d world | arXiv: 2605.01498
- tr2m transferring monocular relative depth to metric depth with language descrip | arXiv: 2506.13387
- tracegen world modeling in 3d trace space enables learning from cross-embodiment
- tracking-guided 4d generation foundation-tracker motion priors for 3d model anim
- trackmae video representation learning via track mask and predict | arXiv: 2603.27268
- trafficalign aligning large language models for traffic scenario generation
- training high-level schedulers with execution-feedback reinforcement learning fo | arXiv: 2511.22235
- training one model to master cross-level agentic actions via reinforcement learn | arXiv: 2512.09706
- training-free detection of generated videos via spatial-temporal likelihoods | arXiv: 2603.15026
- training-free mixed-resolution latent upsampling for spatially accelerated diffu
- training-free open-vocabulary camouflaged object segmentation via fine-grained o
- training-free perceptually consistent low-resolution previews
- training-only heterogeneous image-patch-text graph supervision for advancing few
- trajrag retrieving geometric-semantic experience for zero-shot object navigation | arXiv: 2605.01700
- trajtok learning trajectory tokens enables better video understanding | arXiv: 2602.22779
- transform to transfer boosting adversarial attack transferability on vision-lang
- transition matching distillation for fast video generation
- transprune token transition pruning for efficient large vision-language model
- trcorsurg temporal-relational co-reasoning for surgical video triplet recognitio
- treeteaming autonomous red-teaming of vision-language models via hierarchical s | arXiv: 2603.22882
- tri-modal fusion transformers for uav-based object detection
- tri-subspaces disentanglement for multimodal sentiment analysis | arXiv: 2602.19585
- trident a trimodal cascade generative framework for drug and rna-conditioned cel
- tridf evaluating perception detection and hallucination for interpretable deepfa | arXiv: 2512.10652
- trilite efficient weakly supervised object localization with universal visual fe | arXiv: 2602.23120
- trisim tri-dimensional similarity modeling with extreme value theory for false-n
- trivia self-supervised fine-tuning of vision-language models for table recogniti | arXiv: 2512.01248
- trm-vla temporal-aware chain-of-thought reasoning and memorization for vision-la
- trophies temporal reconstruction of places humans and cameras from multi-view vi
- truckdrive long-range autonomous highway driving dataset
- trust-calibrated collaborative learning for long-tailed visual recognition
- tstm temporal segmentation for task-relevant mask in visual reinforcement learni
- ttapformer robust arbitrary point tracking via transient asynchronous fusion of
- ttl test-time textual learning for ood detection with pretrained vision-language | arXiv: 2604.15756
- ttp test-time padding for adversarial detection and robust adaptation on vision-
- ttrv test-time reinforcement learning for vision language models
- tttlrm test-time training for long context and autoregressive 3d reconstruction | arXiv: 2602.20160
- tudsr twice upsampling-diffusion for higher super-resolution
- tuna taming unified visual representations for native unified multimodal models
- turbo-gs accelerating 3d gaussian fitting for high-quality radiance fields | arXiv: 2412.13547
- turning pre-trained vision transformers into end-to-end histopathology whole sli
- tutor-student reinforcement learning a dynamic curriculum for robust deepfake de | arXiv: 2603.24139
- tv2tv a unified framework for interleaved language and video generation
- tvhighlights llm-guided human-free collaborative training for video highlight de
- tweo transformers without extreme outliers enables fp8 training and quantization
- twin-t twintvqa a reliable structure-detail separating vlm and a comprehensive b
- twings thin plate splines warp-aligned initialization for sparse-view gaussian s | arXiv: 2605.22069
- u-mind a unified framework for real-time multimodal interaction with audiovisual | arXiv: 2602.23739
- u2flow uncertainty aware unsupervised optical flow estimation | arXiv: 2604.10056
- u4d uncertainty-aware 4d world modeling from lidar sequences | arXiv: 2512.02982
- uare a unified vision-language model for image quality assessment restoration an
- uav-cb a complex-background rgb-t dataset and local frequency bridge network for
- uavgen visual prototype conditioned focal region generation for uav based object detection | arXiv: 2604.02966
- UAVLight: A Benchmark for Illumination-Robust 3D Reconstruction in Unmanned Aerial Vehicle (UAV) Scenes
- ucan unified convolutional attention lightweight sr | arXiv: 2603.11680
- ucmnet uncertainty-aware context memory network for under-display camera image r
- udapose unsupervised domain adaptation for low light human pose estimation | arXiv: 2604.10485
- UETrack: A Unified and Efficient Framework for Single Object Tracking | arXiv: 2603.01412
- UFO: Unifying Feed-Forward and Optimization-based Methods for Large Driving Scene Modeling
- ufvideo towards unified fine-grained video cooperative understanding with large | arXiv: 2512.11336
- ui-lens assessing general mllms potential to automate ui display quality assuran
- ulf-loc unbiased landmark feature for robust visual localization with 3d gaussia
- ultra diffusion poser diffusion-based human motion tracking from sparse inertial | arXiv: 2606.02153
- ultra-fast neural video compression | arXiv: 2606.04410
- ultra-low bitrate perceptual image compression with shallow encoder
- UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios
- ultrasound-clip semantic-aware contrastive pre-training for ultrasound image-tex | arXiv: 2604.01749
- unblur-slam dense neural slam for blurry inputs | arXiv: 2603.26810
- uncertainty-aware exploratory direct preference optimization for multimodal larg
- uncertainty-aware knowledge distillation for multimodal large language models | arXiv: 2603.21426
- uncertainty-aware modality fusion for unaligned rgb-t salient object detection
- uncertainty-driven 3d gaussian splatting active mapping via anisotropic visibili | arXiv: 2605.30342
- uncertainty-guided compositional alignment with part-to-whole semantic represent | arXiv: 2603.22042
- underground plant exploration non-destructive 3d root assessment with gpr based
- understanding and mitigating hallucinations in multimodal chain-of-thought model | arXiv: 2603.27201
- understanding task transfer in vision-language models | arXiv: 2511.18787
- understanding temporal logic consistency in video-language models through cross- | arXiv: 2510.08138
- understanding the role of hallucination in reinforcement post-training of multim | arXiv: 2604.03179
- uni-dad unified distillation and adaptation of diffusion models for few-step few | arXiv: 2511.18281
- uni-encoder meets multi-encoders representation before fusion for brain tumor se | arXiv: 2604.22177
- uni-ood unified object- and image-level out-of-distribution detection via cross-
- uni3r unified 3d reconstruction and semantic understanding via generalizable gau
- uniavgen unified audio and video generation with asymmetric cross-modal interact | arXiv: 2511.03334
- unicac universal computational aberration correction benchmark | arXiv: 2603.12083
- unicbench unified counting benchmark for mllm | arXiv: 2603.00595
- unicomp rethinking video compression through informational uniqueness | arXiv: 2512.03575
- unicompress token compression for unified vision-language understanding and gene
- unicorrn unified correspondence transformer across 2d and 3d | arXiv: 2605.04044
- unidac universal metric depth estimation for any camera
- unidex a robot foundation suite for universal dexterous hand control from egocen | arXiv: 2603.22264
- uniedit-i training-free image editing for unified vlm via iterative understandin
- unified camera positional encoding for controlled video generation | arXiv: 2512.07237
- unified generation and self-verification for vision-language models via advantag
- unified number-free text-to-motion generation via flow matching
- unified primitive proxies for structured shape completion | arXiv: 2601.00759
- unified spherical frontend learning rotation-equivariant representations of sphe | arXiv: 2511.18174
- unified vector floorplan generation via markup representation | arXiv: 2604.04859
- unifusion a unified image fusion framework with robust representation and source | arXiv: 2603.14214
- unifying language-action understanding and generation for autonomous driving
- unifying perception and action a hybrid-modality pipeline with implicit visual c
- UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in RL
- unigendet a unified generative-discriminative framework for co-evolutionary imag | arXiv: 2604.21904
- unildiff unlocking the power of diffusion priors for all-in-one image restoratio
- unils end-to-end audio-driven avatars for unified listening and speaking | arXiv: 2512.09327
- unim a unified any-to-any interleaved multimodal benchmark | arXiv: 2603.05075
- unimernet a universal network for real-world mathematical expression recognition
- unimmad unified multi-modal and multi-class anomaly detection via moe-driven fea | arXiv: 2509.25934
- unipart part-level 3d generation with unified 3d geom-seg latents
- unipercept a unified diffusion model for generalizable visual perception
- unipr unified object-level real-to-sim perception and reconstruction from a sing
- unirain unified image deraining rag dataset distillation | arXiv: 2603.03967
- unirefiner teaching pre-trained vits to self-dispose dross via contrastive regis | arXiv: 2605.19622
- unispector towards universal open-set defect recognition via spectral-contrastiv | arXiv: 2604.02905
- unisplat 3d representations unposed | arXiv: 2604.10573
- unit unified multimodal chain-of-thought test-time scaling
- unityvideo unified multi-modal multi-task learning for enhancing world-aware vid
- universal 3d shape matching via coarse-to-fine language guidance | arXiv: 2602.19112
- universal guideline-driven image clustering via a hybrid llm agent
- universal-to-specific dynamic knowledge-guided multiple instance learning for fe
- universe a unified modulation framework for segmentation-free disentangled multi
- unlearning without forgetting securely removing targeted concepts from large-sca
- unleashing the intrinsic visual representation capability of multimodal large la
- unleashing the power of chain-of-prediction for monocular 3d object detection
- unleashing vision-language semantics for deepfake video detection | arXiv: 2603.24454
- unleashing vla potentials in autonomous driving via explicit learning from failu
- unlocking motion from large vision models with a semantic and kinematic duality
- unlocking positive transfer in incrementally learning surgical instruments a sel | arXiv: 2604.02877
- unlocking pre-trained weights parameter inheritance for zero-shot initialization
- unlocking strong supervision a data-centric study of general-purpose audio pre-t | arXiv: 2603.25767
- unlocking the power of critical factors for 3d visual geometry estimation | arXiv: 2604.21713
- unlocking token rewards via training-free reward attribution
- unpaired image deraining using reward-guided self-reinforcement strategy
- unposed-to-3d learning simulation-ready vehicles from real-world images | arXiv: 2604.19257
- unreflectanything rgb-only highlight removal by rendering synthetic specular sup
- unstitching the chimera frame-level risk and train-free mitigation for video hal
- unsupervised monocular 3d keypoint discovery from multi-view diffusion priors | arXiv: 2507.12336
- unsupervised multi-scale segmentation of 3d subcellular world with stable diffus
- upsample anything a simple and hard to beat baseline for feature upsampling
- urban-gs a unified 3d gaussian splatting framework for compact and high-fidelity
- urica a uniformity region affine identifier capture algorithm for arbitrary regi
- ust-hand an uncertainty-aware spatiotemporal point cloud interaction network for | arXiv: 2605.17742
- uvu improving multimodal understanding via vision-language unified autoregressiv
- v-attack targeting disentangled value features for controllable adversarial atta | arXiv: 2511.20223
- v-rgbx video editing with accurate controls over intrinsic properties
- v2-sam marrying sam2 with multi-prompt experts for cross-view object corresponde
- v2u4real a real-world large-scale dataset for vehicle-to-uav cooperative percept
- va-p variational policy alignment for pixel-aware autoregressive generation
- vabench a comprehensive benchmark for audio-video generation
- vad-gs visibility-aware densification for 3d gaussian splatting in dynamic urban
- vanast virtual try-on with human image animation via synthetic triplet supervisi | arXiv: 2604.04934
- var rl done right tackling asynchronous policy conflicts in visual autoregressiv
- variation-aware vision token dropping for faster large vision-language models | arXiv: 2509.01552
- varsplat uncertainty-aware 3d gaussian splatting for robust rgb-d slam | arXiv: 2603.09673
- vcp-attack visual-contrastive projection for transferable black-box targeted att
- vcu-bridge hierarchical visual connotation understanding via semantic bridging
- vde training-free accelerating rectified flow model via velocity decomposition a | arXiv: 2605.23381
- vdfe difference-aware 3d scene editing with non-intrusive video diffusion priors
- vecattention vector-wise sparse attention for accelerating long context inferenc | arXiv: 2603.29494
- vecglypher unified vector glyph generation with language models | arXiv: 2602.21461
- vectorark learning practical image vectorization with rounded polygon representa | arXiv: 2605.24398
- velox learning representations of 4d geometry and appearance | arXiv: 2605.04527
- vemamba efficient isotropic reconstruction of volume electron microscopy with ax
- venus benchmarking and empowering multimodal large language models for aesthetic | arXiv: 2602.23980
- versecrafter dynamic realistic video world model with 4d geometric control | arXiv: 2601.05138
- ves-rft rewarding visual evidence sensitivity to mitigate hallucinations in larg
- vesmamba 3d pulmonary vessel segmentation from ct images via mamba with structur
- vfm-vae vision foundation models can be good tokenizers for latent diffusion mod | arXiv: 2510.18457
- vga bench unified benchmark for video aesthetics and generation quality | arXiv: 2604.10127
- vga empowering aerial-ground localization by visual geometry alignment
- vga-bench a unified benchmark and multi-model framework for video aesthetics and
- vgent visual grounding via modular design for disentangling reasoning and predic
- vgg-t3 offline feed-forward 3d reconstruction at scale | arXiv: 2602.23361
- vggdrive empowering vision-language models with cross-view geometric grounding f | arXiv: 2602.20794
- vggt-360 geometry-consistent zero-shot panoramic depth estimation
- vggt-det mining vggt internal priors for sensor-geometry-free multi-view indoor | arXiv: 2603.00912
- vggt-segmentor geometry-enhanced cross-view segmentation
- vggt-ω | arXiv: 2605.15195
- viaformer voxel-image alignment transformer for high-fidelity voxel refinement
- vibes a conversational agent with behaviorally intelligent 3d virtual body | arXiv: 2512.14234
- vibetoken scaling 1d image tokenizers and autoregressive models for dynamic reso
- video generation with stable transparency via shiftable rgb-a distribution learn
- video panels for long video understanding | arXiv: 2509.23724
- video-only tom enhancing theory of mind in multimodal large language models | arXiv: 2603.24484
- video2robo 3dgs-based synthetic data from one video enables scalable robot learn
- videoarm agentic reasoning over hierarchical memory for long-form video understa | arXiv: 2512.12360
- videochatm1 collaborative policy planning for vide | arXiv: 2511.19524
- videofusion a spatio-temporal collaborative network for multi-modal video fusion | arXiv: 2503.23359
- videoitg multimodal video understanding with instructed temporal grounding
- videomt your vit is secretly also a video segmentation model | arXiv: 2602.17807
- videonet a large-scale dataset for domain-specific action recognition | arXiv: 2605.02834
- videorealbench a chain-of-thought realism evaluation benchmark for generated hum
- videoseek long-horizon video agent with tool-guided seeking | arXiv: 2603.20185
- videoweaver multimodal multi-view video-to-video transfer for embodied agents
- vidtag temporally aligned video to gps geolocalization with denoising sequence p
- vidtag video gps geolocalization | arXiv: 2604.12159
- vihoi human-object interaction synthesis with visual priors | arXiv: 2603.24383
- vikey enhancing temporal understanding in videos via visual prompting | arXiv: 2603.23186
- vilearn accelerating training convergence of image-to-3d generation via visibili
- vimcan visual-inertial 3d human pose estimation with hybrid mamba-cross-attentio | arXiv: 2605.07552
- vinedresser3d agentic text-guided 3d editing | arXiv: 2602.19542
- vinedresser3d towards agentic text-guided 3d editing
- vinqa visual elements interleaved long-form answer generation for real-world mul
- vins-120k ultra high-resolution image editing with a large-scale dataset
- viral visual sim-to-real at scale for humanoid loco-manipulation
- virc enhancing visual interleaved mathematical cot with reason chunking | arXiv: 2512.14654
- vird view-invariant representation through dual-axis transformation for cross-vi | arXiv: 2603.12918
- viro robust and efficient neuro-symbolic reasoning with verification for referri | arXiv: 2601.12781
- virtual full-stack scanning of brain mri via imputing any quantised code | arXiv: 2501.18328
- virtual immunohistochemistry staining with dual-aligned multi-task feature guida
- virtual nodes guided dynamic graph neural network for brain tumor segmentation w
- virtuebench evaluating trustworthiness under uncertainty in long video understan | arXiv: 2603.07071
- visilock authorizing instruction-based image editing with dual score distillatio
- vision foundation models can be good tokenizers for latent diffusion models
- vision on request enhanced vllm efficiency with sparse dynamically selected visi | arXiv: 2603.23495
- vision transformers need more than registers | arXiv: 2602.22394
- vision-language attribute disentanglement and reinforcement for lifelong person | arXiv: 2603.19678
- vision-language model guided source-free domain adaptation via optimal transport
- vision-oriented lightweight neural architecture search with budget-adaptive eval
- visiondirector vision-language guided closed-loop refinement for generative imag
- visionleaf entropy-guided leaf-first reasoning for efficient and accurate think-
- vismem latent vision memory unlocks potential of vision-language models
- visplay self-evolving vision-language models
- visref visual refocusing test time scaling | arXiv: 2603.00207
- visres bench on evaluating the visual reasoning capabilities of vlms
- vista a test-time self-improving video generation agent
- vista4d video reshooting with 4d point clouds | arXiv: 2604.21915
- vistorybench comprehensive benchmark suite for story visualization | arXiv: 2505.24862
- visual diffusion models are geometric solvers
- visual document understanding and reasoning a multi-agent collaboration framewor
- visual grounding for object questions
- visual prototype conditioned focal region generation for uav-based object detect
- visual-aware cot achieving high-fidelity visual consistency in unified models
- visualad language-free zero-shot anomaly detection via vision transformer | arXiv: 2603.07952
- visualoverload probing visual understanding of vlms in really dense scenes | arXiv: 2509.25339
- vit3 unlocking test time training in vision | arXiv: 2512.01643
- viterbiplannet injecting procedural knowledge via differentiable viterbi for pla | arXiv: 2603.04265
- vitprompt training-free prompt refinement with visual tokens for open-vocabulary
- viva vlm-guided instruction-based video editing with reward optimization
- vl-routerbench a benchmark for vision-language model routing | arXiv: 2512.23562
- vla models are more generalizable than you think revisiting physical and spatial
- vlic vision-language models as perceptual judges for human-aligned image compres
- vlm model inversion adaptive token weight | arXiv: 2508.04097
- vlm-3r vision-language models augmented with instruction-aligned 3d reconstructi
- vlm-guided group preference alignment for diffusion-based human mesh recovery | arXiv: 2602.19180
- vlm-loc localization in point cloud maps via vision-language models | arXiv: 2603.09826
- vlm-pruner buffering for spatial sparsity in an efficient vlm centrifugal token | arXiv: 2512.02700
- vlm4rsdet collaborative optimization with vision-language model for enhancing re
- vmd-fact a new video dataset and mllm-based method for detecting realistic ai-ge
- vmonarch efficient video diffusion transformers with structured attention
- vodasure a large-scale dataset revealing domain shift in volumetric super-resolu
- vold reasoning transfer from llms to vision-language models via on-policy distil
- vosr a vision only generative model for image super resolution | arXiv: 2604.03225
- voxify3d pixel art meets volumetric rendering | arXiv: 2512.07834
- voxtell free-text promptable universal 3d medical image segmentation
- vq-va world towards high-quality visual question-visual answering
- vqrae representation quantization autoencoders for multimodal understanding gene
- vrclip multimodal canonical correlation alignment for clip-driven vision-radio p
- vs bench evaluating vlms for strategic abilities in multi agent environments | arXiv: 2506.02387
- vsrell a simple baseline for video super-resolution and enhancement in low-light
- vt-intrinsic physics-based decomposition of reflectance and shading using a sing | arXiv: 2509.10388
- vulcan tool-augmented multi agents for iterative 3d object arrangement
- vvs accelerating speculative decoding for visual autoregressive generation via p | arXiv: 2511.13587
- w2w language-model-based trajectory prediction with reinforcement learning
- wadi weight direction-aware distillation for one-step image synthesis | arXiv: 2603.08258
- walkgpt grounded vision-language conversation with depth-aware segmentation for | arXiv: 2603.10703
- wam-flow parallel coarse-to-fine motion planning via discrete flow matching for
- wan-weaver interleaved multi-modal generation via decoupled training | arXiv: 2603.25706
- wanderland geometrically grounded simulation for open-world embodied ai | arXiv: 2511.20620
- watch and learn learning to use computers from online videos | arXiv: 2510.04673
- waterflow watermark temporal robustness via flow consistency
- wave-former through-occlusion 3d reconstruction via wireless shape completion
- wavelet-based frame selection by detecting semantic boundary for long video unde | arXiv: 2603.00512
- weakly supervised video anomaly detection with anomaly-connected components and | arXiv: 2603.00550
- weathercity urban scene reconstruction with controllable multi-weather transform
- weave unleashing and benchmarking the in-context interleaved comprehension and g
- weavetime streaming from earlier frames into emergent memory in videollms
- weavetime streaming video llm memory | arXiv: 2602.22142
- webchain a large-scale human-annotated dataset of real-world web interaction tra
- webgym scaling training environments for long-horizon visual web agents with rea
- wedetect fast open-vocabulary object detection as retrieval
- wemmu enhanced bridging of vision-language models and diffusion models via noisy
- what are you doing a closer look at controllable human video generation
- what do visual tokens really encode uncovering sparsity and redundancy in multim | arXiv: 2603.00510
- what is it like to be a noise an entropy-based gaussian noise regularization for
- what is the optimal ranking score between precision and recall we can always fin | arXiv: 2511.22442
- what is wrong with synthetic data for scene text recognition a strong synthetic | arXiv: 2602.06450
- what makes good synthetic training data for zero-shot stereo matching | arXiv: 2504.16930
- What Matters in Practical Learned Image Compression
- what your features reveal data-efficient black-box feature inversion attack for
- whats wrong with synthetic data for scene text recognition a strong synthetic en
- when anonymity breaks identifying models behind text-to-image leaderboards
- when avsr meets video conferencing dataset degradation and the hidden mechanism
- when clip sees more it fights back harder multi-view guided adaptive counteratta
- when do models actually decide mapping the layer-wise decision timeline in pretr
- when lines meet textures spatial-frequency aligned diffusion features for cross-
- when local rules create global order self-organized representation learning for
- when lora betrays backdooring text-to-image models by masquerading as benign ada | arXiv: 2602.21977
- when numbers speak aligning textual numerals and visual instances in text-to-vid | arXiv: 2604.08546
- when pretty isnt useful investigating why modern text-to-image models fail as re
- when robots obey the patch universal transferable patch attacks on vision-langua | arXiv: 2511.21192
- When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering
- when safety collides resolving multi-category harmful conflicts in text-to-image | arXiv: 2602.20880
- when to think and when to look uncertainty-guided lookback | arXiv: 2511.15613
- when token pruning is worse than random understanding visual token information i | arXiv: 2512.07580
- when transformers meet mamba a hybrid transformer-mamba network for video object
- when understanding becomes a risk authenticity and safety risks in the emerging | arXiv: 2603.24079
- when visualizing is the first step to reasoning mira a benchmark for visual chai
- where does vision meet language understanding and refining visual fusion in mllm
- where mllms attend and what they rely on explaining autoregressive token generat | arXiv: 2509.22496
- where what why toward explainable 3d-gs watermarking | arXiv: 2603.08809
- which concepts to forget and how to refuse decomposing concepts for continual un | arXiv: 2603.21484
- white-balance first adjust later cross-camera color constancy via vision-languag | arXiv: 2605.19613
- whu-mars a multispectral aerial-ground benchmark towards any-scenario person re-
- why does rl generalize better than sft a data-centric perspective on vlm post-tr
- why not hyperparameter-friendly optimisation a monotonic adaptive norm rescaling
- widget2code from visual widgets to ui code via multimodal llms | arXiv: 2512.19918
- wikiclip an efficient contrastive baseline for open-domain visual entity recogni
- wildcap facial albedo capture in the wild via hybrid inverse rendering | arXiv: 2512.11237
- wildpose a unified framework for robust pose estimation in the wild
- wildrayzer self-supervised large view synthesis in dynamic environments
- will multimodal models be dazzled by multi-image visual puzzles
- wiseedit benchmarking cognition- and creativity-informed image editing
- wiser wider search deeper thinking and adaptive fusion for training-free zero-sh | arXiv: 2602.23029
- witta-bench benchmarking test-time adaptation for wifi sensing
- wod-e2e waymo open dataset for end-to-end driving in challenging long-tail scena
- world in a frame understanding culture mixing as a new challenge for vision-lang
- worldreel 4d video generation with consistent geometry and motion modeling
- WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memories
- wpt world-to-policy transfer via online world model distillation | arXiv: 2511.20095
- Write Where It Matters: Policy-Guided Watermarks for 3D Gaussian Splatting
- x-avdt audio-visual cross-attention for robust deepfake detection
- x-pcr a benchmark for cross-modality progressive clinical reasoning in ophthalmi
- x-win building chest radiograph world model via predictive sensing | arXiv: 2511.14918
- x2-fusion cross-modality and cross-dimension flow estimation in event edge space | arXiv: 2603.16671
- xseg a large-scale x-ray contraband segmentation benchmark for real-world securi | arXiv: 2604.03706
- yieldsat a multimodal benchmark dataset for high-resolution crop yield predictio
- yocity personalized and boundless 3d realistic city scene generation via self-cr | arXiv: 2511.18734
- yoeo you only erase once erasing anything without bringing unexpected content | arXiv: 2603.27599
- yolo-master moe-accelerated with specialized transformers for enhanced real-time
- yolo-ulm ultra-lightweight models for real-time object detection
- yose you only select essential tokens for efficient dit-based video object remov
- your classifier can do more towards balancing the | arXiv: 2505.19459
- your dissimilarities define you complementary learning exploiting class diversit
- your latent mask is wrong pixel-equivalent latent compositing for diffusion mode
- yume15 a text-controlled interactive world generation model
- z-order transformer for feed-forward gaussian splatting | arXiv: 2605.13465
- zero-shot image denoising via hybrid prior-guided pseudo sample generation
- zero-shot reconstruction of animatable 3d avatars with cloth dynamics from a sin | arXiv: 2603.14772
- zeroidir zero-reference illumination degradation image restoration with perturbe | arXiv: 2605.11435
- zina multimodal fine-grained hallucination detection and editing | arXiv: 2506.13130
- zipmap linear-time stateful 3d reconstruction via test-time training
- zoo-prune training-free token pruning via zeroth-order gradient estimation in vi
- zoo3d zero-shot 3d object detection at scene level
- zoomearth active perception for ultra-high-resolution geospatial vision-language
- δynamics language-based representation for inferring rigid-body dynamics from vi | arXiv: 2605.20576
- φ-dpo fairness direct preference optimization approach to continual learning in | arXiv: 2602.22601