ECCV2026 论文笔记 TODO¶
总计: 422 篇 | 已完成: 411 | 待更新: 11
- \(S^{2}\)-FracMix: Label-Preserving Self-Saliency Mixup Augmentation | arXiv: 2606.25784
- 2K Retrofit: Entropy-Guided Efficient Sparse Refinement for High-Resolution 3D Geometry Prediction | arXiv: 2603.19964
- 360Anything: Geometry-Free Lifting of Images and Videos to 360° | arXiv: 2601.16192
- 3D Field of Junctions: A Noise-Robust, Training-Free Structural Prior for Volumetric Inverse Problems | arXiv: 2603.02149
- A Classifier-Agnostic Zero-Shot Adversarial Attack Detection via CLIP | arXiv: 2606.30342
- A Mechanism-Driven Theory of Phase Transitions in Active Learning | arXiv: 2607.00144
- A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world | arXiv: 2606.21216
- AC3S: Adaptive Conditioning for 3D-Aware Synthetic Data Generation | arXiv: 2606.31204
- Accelerated Likelihood Maximization for Diffusion-based Versatile Content Generation | arXiv: 2606.31323
- Accelerating Multimodal Large Language Models with Prior-Corrected Token Reduction | arXiv: 2606.24156
- AdaBoosting Text Prompts for Vision-Language Models | arXiv: 2607.00684
- ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs | arXiv: 2606.31054
- Adaptive Spectrum-Aware Feature Disentangled Network for Small Object Detection | arXiv: 2606.29029
- Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods | arXiv: 2606.24484
- AerialMetric: Benchmarking and Adapting UAV Monocular Metric Depth Estimation in the Real World | arXiv: 2606.29716
- AFFMAE: Scalable Vision Pre-Training for High-Resolution Microscopy Segmentation on Desktop Hardware | arXiv: 2602.16249
- Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale | arXiv: 2506.12009
- agentic collaborative cognition for zero-shot 3d understanding | arXiv: 2606.24649
- AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors | arXiv: 2603.17975
- AirZoo: A Unified Large-Scale Dataset for Grounding Aerial Geometric 3D Vision | arXiv: 2604.26567
- Anchored, Not Graded: Vision-Language Models Fail at Slant-from-Texture Perception | arXiv: 2606.06714
- Anchoring on Reality: Breaking the Pseudo-Target Ceiling in Makeup Transfer | arXiv: 2606.31089
- AnyMatch: Supercharging Universal Multi-Modal Image Matching with Large-Scale Single-View Images | arXiv: 2606.31077
- AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization | arXiv: 2603.17461
- Articulat3D: Reconstructing Articulated Digital Twins From Monocular Videos with Geometric and Motion Constraints | arXiv: 2603.11606
- ASTAD: Asymmetric Style Transfer for Synthetic-to-Real Adaptation in Autonomous Driving | arXiv: 2606.29286 | 📄 paper_cache/ECCV2026/2606.29286.txt
- Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video | arXiv: 2512.12165
- AutoSpeed: Annotation-Free Stage-Adaptive Motion Speed Learning for Robot Manipulation | arXiv: 2607.01051
- AVSR-Diff: Scale-Agnostic Diffusion Priors for Temporally Consistent Arbitrary-Scale Video Super-Resolution | arXiv: 2607.00987
- AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation | arXiv: 2606.30811
- BackTranslation2.0 -- A Linguistically Motivated Metric to Assess Sign Language Production | arXiv: 2606.28673
- Belief Contraction in Dynamic Epistemic Logic | arXiv: 2606.31861
- Beyond IID: How General Are Tabular Foundation Models, Really? | arXiv: 2606.30410
- BioMedVR: Confusion-Aware Mixture-of-Prompt Experts for Biomedical Visual Reprogramming | arXiv: 2606.24740 | 📄 paper_cache/ECCV2026/2606.24740.txt
- BrainFIBRE: A Foundation Model via Information Decomposition for Brain Microstructure | arXiv: 2607.00573
- BrainRiem: Riemannian Prototype Learning for Source-Free Cross-Site Brain Network Diagnosis | arXiv: 2606.29200 | 📄 paper_cache/ECCV2026/2606.29200.txt
- BrepLLM: Enabling Large Language Models to Understand Boundary Representations | arXiv: 2512.16413
- Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction | arXiv: 2606.29445
- C3-Bench: A Context-Aware Change Captioning Benchmark | arXiv: 2606.25445 | 📄 paper_cache/ECCV2026/2606.25445.txt
- Capacity-Controlled Multi-View Stylization of 3D Gaussian Splatting | arXiv: 2606.26754
- caption bottleneck models
- CMDS-AD: Cross-Modal Dual-Stream Decoupling for Few-Shot Anomaly Detection | arXiv: 2606.20300
- Condensing Large-Scale Datasets Directly with Minimal Information Loss | arXiv: 2607.00916
- Continuous Speculative Decoding for Autoregressive Image Generation | arXiv: 2411.11925
- Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints | arXiv: 2603.11755 | 📄 paper_cache/ECCV2026/2603.11755.txt
- CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization | arXiv: 2606.31219
- cross space distillation teaching one step students with modern diffusion teache
- Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning | arXiv: 2606.22394
- Curvature-Guided Mixing for MLLM Adaptation | arXiv: 2606.24963 | 📄 paper_cache/ECCV2026/2606.24963.txt
- CustomX: Unified Character, Action, and Scene Customization in Video World Models | arXiv: 2512.17796
- Deconfounded Lifelong Learning for Autonomous Driving via Dynamic Knowledge Spaces | arXiv: 2603.14354 | 📄 paper_cache/ECCV2026/2603.14354.txt
- DefenseSplat: Enhancing the Robustness of 3D Gaussian Splatting via Frequency-Aware Filtering | arXiv: 2602.19323
- Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation | arXiv: 2512.20117 | 📄 paper_cache/ECCV2026/2512.20117.txt
- Denoising-Enhanced Coarse-to-Fine Infrared Small Target Detection with Attention Prior-Guided Knowledge Distillation | arXiv: 2606.21956
- Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views | arXiv: 2606.23557 | 📄 paper_cache/ECCV2026/2606.23557.txt
- DeWorldSG: Depth-Aware 3D Semantic Scene Graph Generation via World-Model Priors | arXiv: 2607.00889
- DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues | arXiv: 2606.26602
- Diffusion Integrated Gradients: Controllable Path Generation for Flexible Feature Attribution | arXiv: 2606.22314
- Diffusion-Based Material Regularization for Physics-Based Inverse Rendering | arXiv: 2606.31065
- Diffusion-Based Material Regularization for Physics-Based Inverse Rendering | arXiv: 2606.31065
- Distill on a Diet: Efficient Knowledge Distillation via Learnable Data Pruning | arXiv: 2606.25488
- Distill Once, Adapt Life-Long: Exploring Dataset Distillation for Continual Test-Time Adaptation | arXiv: 2606.20196
- Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models | arXiv: 2605.01896
- DivRL: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation | arXiv: 2606.23950 | 📄 paper_cache/ECCV2026/2606.23950.txt
- DLGStream: Dynamic Language-embedded Gaussian Splatting for Open-vocabulary Enabled Free-viewpoint Video Streaming | arXiv: 2606.28840
- Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts | arXiv: 2607.00666
- Don't Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance | arXiv: 2606.27371
- Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off | arXiv: 2603.22607
- DriftScope: Measuring The Hidden Effects of Diffusion Model Adaptation | arXiv: 2607.00183
- Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout | arXiv: 2605.05092
- DriveVA: Video Action Models are Zero-Shot Drivers | arXiv: 2604.04198
- DriveWeaver: Point-Conditioned Video Inpainting for Controllable Vehicle Insertion in Autonomous Driving Simulation | arXiv: 2606.31918 | 📄 paper_cache/ECCV2026/2606.31918.txt
- DTI: Dynamic Trajectory Initialization for Generative Face Video Super-Resolution | arXiv: 2606.29198 | 📄 paper_cache/ECCV2026/2606.29198.txt
- Dual-End Consistency Model | arXiv: 2602.10764
- Dual-Prior Guided Null-Space Learning with Mixture-of-Splines for Arbitrary Medical Slice Super-Resolution | arXiv: 2606.26716
- E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation | arXiv: 2606.27268
- E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes | arXiv: 2604.04834
- EcoVideo: Entropy-Orchestrated Video Generation Paradigm in Cloud-Edge Dynamics | arXiv: 2606.30557 | 📄 paper_cache/ECCV2026/2606.30557.txt
- Edges Before Embeddings: A Confidence-Aware Blur Gate for Vision-Language Pipelines | arXiv: 2606.25838
- Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing | arXiv: 2603.03143
- editing everything everywhere all at once
- Efficient Document Tampering Localization with Multi-Level Discrepancy Features and Unified DCT-Quantization Embedding | arXiv: 2606.22285
- Efficient RGB-T Object Detection via Sparse Cross-Modality Fusion | arXiv: 2606.30215 | 📄 paper_cache/ECCV2026/2606.30215.txt
- EgoExo-Con: Exploring View-Invariant Video Temporal Understanding | arXiv: 2510.26113
- EgoExo-Con: 探索视角不变性的视频时序理解 | arXiv: 2510.26113
- EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding | arXiv: 2606.24422
- EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning | arXiv: 2511.18242
- ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval | arXiv: 2606.20280
- EMOSH: Expressive Motion and Shape Disentanglement for Human Animation | arXiv: 2606.28026
- Entropy-Controlled Flow Matching | arXiv: 2602.22265
- Tune-A-Video+: Enhanced Text-Guided Video Editing and Generation Framework
- EPO: Boosting 3D Foundation Models with Edge-based Pose Optimization | arXiv: 2607.00579
- EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal | arXiv: 2512.21545
- ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis | arXiv: 2603.25168
- Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs | arXiv: 2606.20177
- Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations | arXiv: 2606.24716
- event-driven video generation | arXiv: 2603.13402
- Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks | arXiv: 2603.11689
- Exploiting Local Flatness for Efficient Out-of-Distribution Detection | arXiv: 2606.29952
- ExPLoRe: Expert Patch-Level Loss Routing for Multi-Objective Masked Image Modeling | arXiv: 2606.31201
- ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving | arXiv: 2604.02714
- Fabric Image Demoiréing Benchmark from Synthesis to Restoration | arXiv: 2606.24072
- Face Anything: 4D Face Reconstruction from Any Image Sequence | arXiv: 2604.19702
- FaceMoE: Mixture of Experts for Low-Resolution Face Recognition | arXiv: 2606.32040
- FD\(^2\): A Dedicated Framework for Fine-Grained Dataset Distillation | arXiv: 2603.25144
- FedLAS: Feature-Modulated Bidirectional Label Smoothing for Neural Network Calibration | arXiv: 2606.28654
- FedOT: Ownership Verification and Leakage Tracing via Watermarks for Federated LDMs | arXiv: 2606.22875
- FeVOS: Foresight Expression Video Object Segmentation | arXiv: 2606.25585
- Few-Shot Synthetic Image Attribution: Identifying Unseen Generators with Limited Samples | arXiv: 2509.25682
- FiCA: Feed-forward Instant Gaussian Codec Avatars from a Single Portrait Image | arXiv: 2606.24232
- fidelity- and perception-aware local implicit attention for arbitrary-scale imag | arXiv: 2606.21910 | 📄 paper_cache/ECCV2026/2606.21910.txt
- FLM-Occ: Feed-forward Likelihood Maximization for Efficient Indoor Occupancy Prediction | arXiv: 2606.21373
- FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation | arXiv: 2606.22424
- FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation | arXiv: 2511.21029 | 📄 paper_cache/ECCV2026/2511.21029.txt
- FMA-Net++: Motion- and Exposure-Aware Joint Video Super-Resolution and Deblurring | arXiv: 2512.04390
- Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE | arXiv: 2606.26938
- Following the Flow: Advection-Consistent Modeling for Event-based Small Object Detection | arXiv: 2606.22378
- Forget, Anticipate and Adapt: Test Time Training for Long Videos | arXiv: 2606.26515
- FreqOrtho-SR: Frequency-Guided Orthogonal Expert Learning for Real-World Image Super-Resolution | arXiv: 2606.28745
- from hallucination to grounding diagnosing visual spatial intelligence via crisp | arXiv: 2606.26535 | 📄 paper_cache/ECCV2026/2606.26535.txt
- From Phase to Phenomenon: Self-Supervised Learning of Subsurface Scattering with Minimal Phase-shift Inputs | arXiv: 2606.29461 | 📄 paper_cache/ECCV2026/2606.29461.txt
- From Reconstruction to Decision: A Post-Encoder Plug-in Adapter for Curvilinear Segmentation | arXiv: 2606.23486
- FrozenDrive: Zero-Shot Text-Guided Driving Scene Generation and Data Augmentation | arXiv: 2606.20110
- FUMO: Prior-Modulated Diffusion for Single Image Reflection Removal | arXiv: 2603.19036
- G2P: Gaussian-to-Point Attribute Alignment for Boundary-Aware 3D Segmentation | arXiv: 2601.03510
- GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models | arXiv: 2601.18197 | 📄 paper_cache/ECCV2026/2601.18197.txt
- Gaussian Belief Propagation Network for Depth Completion | arXiv: 2601.21291
- GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation | arXiv: 2603.26661
- GaussLite: Online Task-Conditioned 3D Gaussian Splatting for Real-Time Robotic Mapping | arXiv: 2606.30809
- GaussLite:在线任务条件化的 3D 高斯泼溅实时机器人建图 | arXiv: 2606.30809
- Gen2Balance: Generative Balancing for Long-Tailed Video Action Recognition | arXiv: 2606.22416
- GENA3D: Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence | arXiv: 2511.21945
- Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval | arXiv: 2602.00813
- Generative Lane Topology Reasoning via Autoregressive Model with Geometry Prior | arXiv: 2606.31814 | 📄 paper_cache/ECCV2026/2606.31814.txt
- GenSP: Consistent Spherical Parameterization via Learning Shape Generative Models | arXiv: 2607.00492
- genvideolens where lvlms fall short in ai generated video detection
- Geo-ID: Test-Time Geometric Consensus for Cross-View Consistent Intrinsics | arXiv: 2603.13859
- GeoEdit: Geometry-Aware Object Editing via Dual-Branch Denoising | arXiv: 2606.30003
- Geometric Gradient Rectification for Safe Open-Set Semi-Supervised Learning | arXiv: 2606.26973
- geometry-anchored transport framework for exemplar-free class-incremental learni | arXiv: 2606.25347 | 📄 paper_cache/ECCV2026/2606.25347.txt
- Geometry-Aware Style Transfer in 3D Gaussian Splatting | arXiv: 2606.24144
- Geometry-Preserving in 3D Gaussian Splatting for LiDAR-Camera Extrinsic Calibration | arXiv: 2606.20103
- GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis | arXiv: 2603.14965 | 📄 paper_cache/ECCV2026/2603.14965.txt
- GKDT: General Keypoint Detection Transformer | arXiv: 2607.00752
- GRAFT: Geometric Refinement and Fitting Transformer for Human Scene Reconstruction | arXiv: 2604.19624
- Graph Coloring for Multi-Task Learning | arXiv: 2509.16959 | 📄 paper_cache/ECCV2026/2509.16959.txt
- GryphOne: Symbol-Aware Masked Diffusion for Structural Refinement in Offline Handwritten Mathematical Expression Recognition | arXiv: 2602.03370
- GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation | arXiv: 2603.26266
- H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks | arXiv: 2606.25578 | 📄 paper_cache/ECCV2026/2606.25578.txt
- HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration | arXiv: 2606.28215
- HieDG: A Hierarchical Discrete Geometry-Guided Framework for Multi-Animal Tracking | arXiv: 2607.00494
- Hierarchical 3D Scene Graph Construction and Belief-based Planning for Semantic Navigation | arXiv: 2606.31071 | 📄 paper_cache/ECCV2026/2606.31071.txt
- Hierarchical Spatial and Channel Aggregation for Cross-domain Few-shot Segmentation | arXiv: 2606.24296 | 📄 paper_cache/ECCV2026/2606.24296.txt
- HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-training | arXiv: 2606.20189 | 📄 paper_cache/ECCV2026/2606.20189.txt
- Histogram-constrained Image Generation | arXiv: 2606.31683
- Histopathology Multi-modal Embedding for Pathology Composed Retrieval | arXiv: 2502.07221
- HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis | arXiv: 2606.09098
- Horizon3D: Sparse Radar-Camera Fusion for Long-Range 3D Perception in Autonomous Driving | arXiv: 2606.31096
- HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding | arXiv: 2602.12957
- HSDF-Lane: Height-Aligned Signed Distance Field with Semantic Lane Prior for 3D Lane Detection | arXiv: 2606.31172
- Hybrid Event–Frame Sensors: Modeling, Calibration, and Simulation | arXiv: 2511.18037
- HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding | arXiv: 2607.00428
- HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization | arXiv: 2604.20328
- Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting | arXiv: 2607.00159
- Improving Adversarial Robustness via Activation Amplification and Attenuation | arXiv: 2606.27784
- Improving Sparse-View 3DGS Generalization via Flat Minima Optimization | arXiv: 2607.00885
- In-context Region-based Drag: Drag Any Region to Any Shape | arXiv: 2606.25907
- Information-Regularized Attention for Visual-Centric Reasoning | arXiv: 2607.00434
- Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models | arXiv: 2607.01222
- Interact3D: Compositional 3D Generation of Interactive Objects | arXiv: 2603.16085
- InterEdit: Navigating Text-Guided 3D Dyadic Human Motion Editing | arXiv: 2603.13082
- Intermediate Text Representation Guided Text-to-Image Generation for Enhancing One-and-Only Alignment | arXiv: 2606.30262
- Intrinsically Stable Spiking Neural Networks: Overcoming the Performance Barrier in the Absence of Batch Normalization | arXiv: 2606.31695
- Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity | arXiv: 2606.25343
- IREU: Identity-Related Encoder-Only Unlearning for Customized Portrait Generation | arXiv: 2606.29880 | 📄 paper_cache/ECCV2026/2606.29880.txt
- ISAC: Training-Free Instance-to-Semantic Attention Control for Multi-Instance Generation | arXiv: 2505.20935
- JanusMesh: Fast and Zero-Shot 3D Visual Illusion Generation via Cross-Space Denoising | arXiv: 2606.20563
- LaGen: Towards Autoregressive LiDAR Scene Generation | arXiv: 2511.21256
- Large-Scale High-Quality 3D Gaussian Head Reconstruction from Multi-View Captures | arXiv: 2605.04035
- Latent Visual Diffusion Reasoning with Monte Carlo Tree Search | arXiv: 2606.27988 | 📄 paper_cache/ECCV2026/2606.27988.txt
- LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion | arXiv: 2603.14526
- Layer-Specific Prompt Fusion Discovery via Differentiable Search in Vision Foundation Models | arXiv: 2606.26379
- Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking | arXiv: 2509.12046
- LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection | arXiv: 2512.05663 | 📄 paper_cache/ECCV2026/2512.05663.txt
- Learn Once, Edit Anywhere: Visual Direction Transfer for Diffusion Models | arXiv: 2403.19645
- Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents | arXiv: 2606.31270
- Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval | arXiv: 2607.00374
- Learning to Deny: Action Denial in Multimodal Large Language Models | arXiv: 2606.31187
- Learning Transferable Dynamics Priors from Action to World Modeling | arXiv: 2606.29501
- Learning Video Dynamics with Predictive Differentiable Rendering | arXiv: 2606.31050
- LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models | arXiv: 2606.23686
- Lie Group Diffusion Models for Hardware-Aware Quantum Circuit Synthesis | arXiv: 2606.29636
- lightstar efficient visual document retrieval via lightweight selection with vis | arXiv: 2606.23539
- LipsFlow:神经形态 OT-CFM 在多说话人视觉语音识别中的首次探索 | arXiv: 2606.31225
- LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing | arXiv: 2606.26740
- LogicIR: Logic Gate Networks for Image Restoration | arXiv: 2606.26609
- LogiCo: A Unified Framework for Logical and Structural Anomaly Detection | arXiv: 2606.28688
- Long-term Traffic Simulation via Structured Autoregressive Modeling | arXiv: 2606.31209
- Lost in the Tail: Addressing Geographic Imbalance in Urban Visual Place Recognition | arXiv: 2607.00090
- LoT-Pass: Long-term-robust Image Watermarking for Image to Video Generation | arXiv: 2509.17773
- Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Generation | arXiv: 2605.31603
- LUNA: Learning Universal 3D Human Animation Beyond Skinning | arXiv: 2606.31981 | 📄 paper_cache/ECCV2026/2606.31981.txt
- M4-SAR: A Multi-Resolution, Multi-Polarization, Multi-Scene, Multi-Source Dataset and Benchmark for optical-SAR Object Detection | arXiv: 2505.10931
- Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs | arXiv: 2602.05275
- MambaRaw: Selective State Space Modeling for Efficient 4K Raw Image Reconstruction | arXiv: 2606.24479
- MASS: Motion-Aligned Selective Scan for Refinement in Flow-Based Video Frame Interpolation | arXiv: 2606.27718
- Match-Any-Events: Zero-Shot Motion-Robust Feature Matching Across Wide Baselines for Event Cameras | arXiv: 2604.18744 | 📄 paper_cache/ECCV2026/2604.18744.txt
- MATCH: Flow Matching for Multi-View Anomaly Detection | arXiv: 2606.24375
- MAVFusion: Efficient Infrared and Visible Video Fusion via Motion-Aware Sparse Interaction | arXiv: 2604.01958 | 📄 paper_cache/ECCV2026/2604.01958.txt
- MedCAGD: Context-Aware Gated Decoder for Efficient Medical Image Segmentation | arXiv: 2607.00409
- MedRegion-CT: Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation | arXiv: 2506.23102
- MeGAS: Thermomechanical Dynamic Gaussian Splatting for Thermophysical Scene Editing | arXiv: 2606.23455 | 📄 paper_cache/ECCV2026/2606.23455.txt
- MemLearner: Learning to Query Context Memory for Video World Models | arXiv: 2606.31734
- MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts | arXiv: 2607.00371
- MG-RWKV: Multi-Grained Context-Aware RWKV for Temporal Forgery Localization | arXiv: 2607.00902
- MGI: Member vs Generated Inference | arXiv: 2606.23872
- MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation | arXiv: 2606.26016
- MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations | arXiv: 2606.27779 | 📄 paper_cache/ECCV2026/2606.27779.txt
- MIRROR: Aligning Semantic Relations from Language to Image via Gromov--Wasserstein | arXiv: 2606.29462 | 📄 paper_cache/ECCV2026/2606.29462.txt
- MirrorPPR: Exemplar-Based Portrait Photo Retouching | arXiv: 2606.29308
- MixTTA: Low-Rank Cross-Channel Mixing for Reliable Test-Time Adaptation | arXiv: 2606.28142
- MLVC: 面向真实部署的多平台学习式视频编解码器 | arXiv: 2606.28027
- MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation | arXiv: 2604.19679
- MobileManiBench: Simplifying Model Verification for Mobile Manipulation | arXiv: 2602.05233
- ModuSeg: Decoupling Object Discovery and Semantic Retrieval for Training-Free Weakly Supervised Segmentation | arXiv: 2604.07021 | 📄 paper_cache/ECCV2026/2604.07021.txt
- Moiré Video Authentication: A Physical Signature Against AI Video Generation | arXiv: 2604.01654
- MonoSR: Open-Vocabulary Spatial Reasoning on Monocular Images | arXiv: 2511.19119
- Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splatting | arXiv: 2606.30017
- MotionAtlas: Detailed Region Captioning for Motion-Centric Videos | arXiv: 2606.29531 | 📄 paper_cache/ECCV2026/2606.29531.txt
- MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment | arXiv: 2607.00858 | 📄 paper_cache/ECCV2026/2607.00858.txt
- MoVA:非对称双投影实现模块化长视频-文本对齐 | arXiv: 2607.00858
- MSPL: Multi-Step Pseudo-Labeling for Open-Vocabulary Object Detection | arXiv: 2510.14792
- Multi-modality Image Fusion under Adverse Weather: Mask-Guided Feature Restoration and Interaction | arXiv: 2606.26812
- Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments | arXiv: 2607.00457 | 📄 paper_cache/ECCV2026/2607.00457.txt
- Multi-scale Object-Aware Gaze Estimation via Geometric Reasoning | arXiv: 2606.29334 | 📄 paper_cache/ECCV2026/2606.29334.txt
- Multi-Task Bayesian In-Context Learning | arXiv: 2606.20538
- Multi4D: High-Fidelity Dynamic Gaussian Splatting via Multi-Level Competitive Allocation | arXiv: 2606.22197
- NaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene Generation | arXiv: 2606.29395
- NavWM: A Unified Navigation World Model for Foresight-Driven Planning | arXiv: 2606.24101
- NegAS: Negative Label Guided Attention and Scoring for Out-of-Distribution Object Detection with Vision-Language Models | arXiv: 2606.22537 | 📄 paper_cache/ECCV2026/2606.22537.txt
- Neural Gate: Mitigating Privacy Risks in LVLMs via Neuron-Level Gradient Gating | arXiv: 2603.12598
- Next-Frame Decoding for Ultra-Low-Bitrate Image Compression with Video Diffusion Priors | arXiv: 2603.15129
- NGPS: Structure-Preserving Self-Supervised Denoising via Neighbor-Guided Patch Sampling | arXiv: 2606.23200
- No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs | arXiv: 2606.31933 | 📄 paper_cache/ECCV2026/2606.31933.txt
- NoiseTilt: Noise-Tilted Reverse Kernels for Diffusion Reward Alignment | arXiv: 2606.18066
- Not All Prediction Targets Keep Training-Free Diffusion Guidance on the Manifold | arXiv: 2607.00647
- NURBS Splatting: A Unified Differentiable Rendering Framework for Vector Graphics | arXiv: 2606.31764
- Obliviate: Erasing Concepts from Autoregressive Image Generation Models | arXiv: 2606.28643
- Odoriko: A Shape-Aware Multimodal Diffusion Framework for Human Motion | arXiv: 2606.21135
- OLIVE: View-Augmented Latent Prediction with Waveform Reconstruction for Speech SSL | arXiv: 2606.30356
- OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data | arXiv: 2606.30019 | 📄 paper_cache/ECCV2026/2606.30019.txt
- OmniNWM: Omniscient Driving Navigation World Models | arXiv: 2510.18313 | 📄 paper_cache/ECCV2026/2510.18313.txt
- On Test-Time Scaling for Vision-Language Models | arXiv: 2606.28864
- On the Faithfulness of Post-Hoc Concept Bottleneck Models | arXiv: 2606.30498
- On the Vulnerability of Parameter-Level Defenses to Model Merging | arXiv: 2606.30360 | 📄 paper_cache/ECCV2026/2606.30360.txt
- One Video, One World: Turning Monocular Video into Physical 4D Scenes | arXiv: 2606.31388
- OnPoint: Offline-to-Online Multi-Level Distillation for Point-Supervised Online Temporal Action Localization | arXiv: 2607.00289
- OP3DSG: Open-Vocabulary Part-Aware 3D Scene Graph Generation for Real-World Environments | arXiv: 2606.29786 | 📄 paper_cache/ECCV2026/2606.29786.txt
- Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints | arXiv: 2606.24353
- OrthoTrack: Continuous 6-DoF UAV Trajectory Estimation Anchored in Public Orthophotos | arXiv: 2606.25245
- OTCache: Optimal Transport for Geometry-Aware Caching in Diffusion Models | arXiv: 2606.31026
- PA-VAD: Diffusion-Based Pseudo-Only Video Anomaly Detection via Domain-Aligned Memory Updates | arXiv: 2512.06845 | 📄 paper_cache/ECCV2026/2512.06845.txt
- Pano3D: Unified 3D Reconstruction and Panoptic Segmentation | arXiv: 2606.14307
- PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding | arXiv: 2512.20907
- Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models | arXiv: 2606.27373
- Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising | arXiv: 2607.00407
- Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer | arXiv: 2511.19778
- PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing | arXiv: 2606.26551
- PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation | arXiv: 2512.24551
- Physically Grounded 3D Generative Reconstruction under Hand Occlusion using Proprioception and Multi-Contact Touch | arXiv: 2604.09100
- Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation | arXiv: 2606.25306
- PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation | arXiv: 2606.26916
- PIAvatar: Physically Interactive Avatars via Deformation Gradient Decoupling | arXiv: 2606.21162
- PLOT: Pseudo-Labeling via Object Tracking for Monocular 3D Object Detection | arXiv: 2507.02393
- Pointer-CAD v2: Plan-Then-Construct CAD Generation with Dimension-Aware Parametric Precision | arXiv: 2606.29301
- PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models | arXiv: 2606.22540
- Pondering the Way: Spatial-perceiving World Action Model for Embodied Navigation | arXiv: 2606.29908 | 📄 paper_cache/ECCV2026/2606.29908.txt
- Pose Anything Anywhere:Model-free Object Poses from Arbitrary References | arXiv: 2606.23634 | 📄 paper_cache/ECCV2026/2606.23634.txt
- PoseShield: Neural Collision Fields for Human Self-Collision Resolution | arXiv: 2606.29686
- PriorEye: Geospatial Visual Priors for End-to-End Autonomous Driving | arXiv: 2606.31830 | 📄 paper_cache/ECCV2026/2606.31830.txt
- ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models | arXiv: 2603.19466 | 📄 paper_cache/ECCV2026/2603.19466.txt
- Progressive Pose-Guided 4D Animal Reconstruction from Monocular Video | arXiv: 2607.00157 | 📄 paper_cache/ECCV2026/2607.00157.txt
- Prompt2Effect: Training-Free Image-to-Video Model Specialization via LoRA Generation | arXiv: 2606.13971
- ProtoFair: Fair Self-Supervised Contrastive Learning via Pseudo-Counterfactual Pairs | arXiv: 2605.01971
- PS-MOT: Cultivating Instance Awareness from Point Seeds for Multi-Object Tracking | arXiv: 2606.30476
- PS-MOT: Cultivating Instance Awareness from Point Seeds for Multi-Object Tracking | arXiv: 2606.30476
- QuantV2X: A Fully Quantized Multi-Agent System for Cooperative Perception | arXiv: 2509.03704
- RAGA: Real Time Ray Traced Gaussian Shadow Casting for 3DGS Avatar-Scene Interaction | arXiv: 2606.29329
- Rank-Aware Hyperbolic Alignment for Vision–Language Dataset Distillation | arXiv: 2606.29464
- RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation | arXiv: 2606.22749
- RBE-Flow: Recurrent Bayesian Estimation on Feature Manifolds for Cross-Modal Registration | arXiv: 2606.30492
- Real-Time Source-Free Object Detection | arXiv: 2606.31834
- Recovery operators in quasi-Nelson logic: the prelinear case | arXiv: 2606.31277
- RefAlign: Representation Alignment for Reference-to-Video Generation | arXiv: 2603.25743
- Render-FM: Feedforward Model for Real-time Photorealistic Volumetric Rendering | arXiv: 2505.17338
- RePer-360: Releasing Perspective Priors for 360\(^\circ\) Depth Estimation via Self-Modulation | arXiv: 2603.05999
- ReShift: Aha-Moment-Driven Reasoning-Level Backdoor Attacks on Vision-Language Models | arXiv: 2607.00361
- Residual-Guided Expert Specialization for Incomplete Multimodal Learning | arXiv: 2606.30355
- ResilPhase: Plug-and-Play Phase Mapping and Noise-Resilient Macro-Trajectory Extrapolation for Diffusion Acceleration | arXiv: 2606.26769
- RESOLVE: A Multi-Resolution and Multi-Modal Dataset for Roadside Cooperative Perception | arXiv: 2606.31895
- Rethinking Continual Anomaly Detection on the Edge: Benchmarking Under Realistic Industrial Conditions | arXiv: 2605.24251
- Rethinking Garment Conditioning in Diffusion-based Virtual Try-On: Decouple, Don't Denoise | arXiv: 2511.18775
- Rethinking Prototype-based Similarity Learning for Few-Shot Object Detection | arXiv: 2606.23069
- Rethinking Training & Inference for Forecasting: Linking Winner-Take-All back to GMMs | arXiv: 2606.26424 | 📄 paper_cache/ECCV2026/2606.26424.txt
- Revisiting Autoregressive Models for Generative Image Classification | arXiv: 2603.19122
- Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation | arXiv: 2606.31382
- Reweighting Framewise Attention in Video Transformers for Facial Expression Understanding | arXiv: 2606.30611
- RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios | arXiv: 2511.18011
- RoboAtlas: Contextual Active SLAM | arXiv: 2606.26046
- Robust Zero-shot Anomaly Detection under Limited Auxiliary Anomaly Priors | arXiv: 2606.29428
- ROVA: Are Video Reasoning Models Ready to Go Outside? | arXiv: 2603.10652
- RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning | arXiv: 2606.28266
- Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation | arXiv: 2604.03118
- SAM2Matting: Generalized Image and Video Matting | arXiv: 2606.27339
- SARIF: Segment Anything for Robust Image Forensics | arXiv: 2606.21108
- ScAle: Attention Head Scaling as a Minimal Adapter for Spatial Reasoning in Vision–Language Models | arXiv: 2606.29579
- SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation | arXiv: 2606.30124
- see sniff learning visuo-olfactory representations | arXiv: 2606.27307 | 📄 paper_cache/ECCV2026/2606.27307.txt
- Seeing Touch from Motion: A Unified Modality-Aware Visuo-Tactile Policy with Tactile Motion Correlation | arXiv: 2606.29941 | 📄 paper_cache/ECCV2026/2606.29941.txt
- Segmenting, Fast and Slow: Real-Time Open-Vocabulary Video Instance Segmentation with Dual-Path Processing | arXiv: 2607.00124
- Self-supervised Garment Dynamics with Persistent Wrinkles | arXiv: 2606.25065
- Semantic Browsing: Controllable Diversity for Image Generation | arXiv: 2606.23679
- SemCityLoc: Aerial 6DoF Localization Using Semantic 3D City Models | arXiv: 2606.27444
- SENTRY: SAM2-Enhanced Neighbor-Aware and Temporally Reasoned Memory for Visual Tracking | arXiv: 2606.24449
- SFDATrack: Generalized Source-Free Domain Adaptive Tracking Under Adverse Weather Conditions | arXiv: 2607.00369 | 📄 paper_cache/ECCV2026/2607.00369.txt
- ShellMaker: Language-Guided Exterior Completion under Structural Constraints | arXiv: 2606.31680
- SICAGE: Speaker-Independent Culture-Aware Gesture Generation using TED4C-L Dataset | arXiv: 2606.30001
- SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models | arXiv: 2606.27741
- SIGNER: Temporally Grounded Sign Language Generation via Time-Resolved Conditioning | arXiv: 2506.07460
- SIGNET: Motion-Level Knowledge Transfer for Cross-Language Sign Language Translation | arXiv: 2606.28626
- SignNet-1M: Large-Scale Multilingual Sign Language Video Dataset with Downstream Benchmarks | arXiv: 2606.24361
- Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs | arXiv: 2604.20937
- SK-Adapter: Skeleton-Based Structural Control for Native 3D Generation | arXiv: 2603.14152 | 📄 paper_cache/ECCV2026/2603.14152.txt
- SKEL-CF: Coarse-to-Fine Biomechanical Skeleton and Surface Mesh Recovery | arXiv: 2511.20157
- SkelEM: Training-Signal Decoupling of Skeleton and Diffusion for Self-supervised Axial Super-Resolution in Volume Microscopy | arXiv: 2606.30012
- Skin-R1: Clinical Knowledge-Guided Dermatological Diagnosis Using Vision-Language Models | arXiv: 2511.14900
- Skin-R1: Clinical Knowledge-Guided Dermatological Diagnosis Using Vision-Language Models | arXiv: 2511.14900
- Skin-R1:教材知识引导的皮肤病诊断视觉语言模型 | arXiv: 2511.14900 | 📄 paper_cache/ECCV2026/2511.14900.txt
- SlowBA: An efficiency backdoor attack towards VLM-based GUI agents | arXiv: 2603.08316 | 📄 paper_cache/ECCV2026/2603.08316.txt
- SMART: When is it Actually Worth Expanding a Speculative Tree? | arXiv: 2604.09731
- Solving Semi-Supervised Few-Shot Learning from an Auto-Annotation Perspective | arXiv: 2512.10244
- SON-GOKU:图着色实现多任务学习的无冲突调度 | arXiv: 2509.16959
- Sonar-MASt3R: Real-Time Opti-Acoustic Fusion in Turbid, Unstructured Environments | arXiv: 2603.13585
- SONIC: Spectral Optimization of Noise for Inpainting with Consistency | arXiv: 2511.19985
- Spanning the Visual Analogy Space with a Weight Basis of LoRAs | arXiv: 2602.15727
- Sparsity-Inducing Divergence Losses for Biometric Verification | arXiv: 2606.31664 | 📄 paper_cache/ECCV2026/2606.31664.txt
- SPECSIA: Stylization Dataset for Novel-View Enhancement in Drawing-based 3D Animation | arXiv: 2607.00525
- Spectral and Trajectory Regularization for Diffusion Transformer Super-Resolution | arXiv: 2603.06275 | 📄 paper_cache/ECCV2026/2603.06275.txt
- Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models | arXiv: 2606.24165
- Spectral Gating via Damped Oscillations for Adaptive Implicit Neural Representations | arXiv: 2606.23129
- SpectralSplats: Robust Differentiable Tracking via Spectral Moment Supervision | arXiv: 2603.24036 | 📄 paper_cache/ECCV2026/2603.24036.txt
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models | arXiv: 2510.12784 | 📄 paper_cache/ECCV2026/2510.12784.txt
- StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley | arXiv: 2507.07445 | 📄 paper_cache/ECCV2026/2507.07445.txt
- Staying VIGILant: Mitigating Visual Laziness via Counterfactual Visual Alignment in MLLMs | arXiv: 2606.26387 | 📄 paper_cache/ECCV2026/2606.26387.txt
- Steerable Visual Representations | arXiv: 2604.02327 | 📄 paper_cache/ECCV2026/2604.02327.txt
- StereoGS: Sparse-View 3D Gaussian Splatting via Stereo Priors | arXiv: 2606.30545 | 📄 paper_cache/ECCV2026/2606.30545.txt
- StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning | arXiv: 2607.00465 | 📄 paper_cache/ECCV2026/2607.00465.txt
- StreamEdit: Training-Free Video Editing via Few-Step Streaming Video Generation | arXiv: 2605.21466 | 📄 paper_cache/ECCV2026/2605.21466.txt
- Streaming Dense Voxel Representations for 3D Occupancy Prediction | arXiv: 2503.22087 | 📄 paper_cache/ECCV2026/2503.22087.txt
- Structural Assessment for Understanding and Guiding Dataset Distillation in Discrete Token Space | arXiv: 2606.21705 | 📄 paper_cache/ECCV2026/2606.21705.txt
- Structured Hyperedge Adaptation for Parameter-Efficient Fine-Tuning of Vision Transformers | arXiv: 2606.22383 | 📄 paper_cache/ECCV2026/2606.22383.txt
- SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance | arXiv: 2603.12703 | 📄 paper_cache/ECCV2026/2603.12703.txt
- SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation | arXiv: 2606.31259
- Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding | arXiv: 2604.07753 | 📄 paper_cache/ECCV2026/2604.07753.txt
- SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation | arXiv: 2606.30849 | 📄 paper_cache/ECCV2026/2606.30849.txt
- Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction | arXiv: 2510.03117 | 📄 paper_cache/ECCV2026/2510.03117.txt
- TaxoMIL: Taxonomy-Constrained Learning for Hierarchical Whole Slide Image Analysis | arXiv: 2606.31100 | 📄 paper_cache/ECCV2026/2606.31100.txt
- Tessellating The Earth | arXiv: 2606.27514 | 📄 paper_cache/ECCV2026/2606.27514.txt
- Text-Conditioned Background Generation for Editable Multi-Layer Documents | arXiv: 2512.17151 | 📄 paper_cache/ECCV2026/2512.17151.txt
- TextDS: Parameter-Efficient Representation Alignment for Scene Text Detection under Distribution Shifts | arXiv: 2606.28077 | 📄 paper_cache/ECCV2026/2606.28077.txt
- The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models | arXiv: 2607.00402 | 📄 paper_cache/ECCV2026/2607.00402.txt
- There and Back Again: A Flexible-Frame Transformer for Multi-Exposure Fusion | arXiv: 2606.27905 | 📄 paper_cache/ECCV2026/2606.27905.txt
- Think While You Map: Asynchronous Vision-Language Agents for Incremental 3D Scene Graphs | arXiv: 2606.31471 | 📄 paper_cache/ECCV2026/2606.31471.txt
- Toward Robust In-Context Segmentation via Concept Guidance | arXiv: 2606.28149 | 📄 paper_cache/ECCV2026/2606.28149.txt
- Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning | arXiv: 2511.20196 | 📄 paper_cache/ECCV2026/2511.20196.txt
- Towards Consistent and Efficient Dataset Distillation via Diffusion-Driven Selection | arXiv: 2412.09959 | 📄 paper_cache/ECCV2026/2412.09959.txt
- Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation | arXiv: 2606.30598 | 📄 paper_cache/ECCV2026/2606.30598.txt
- Towards Long-Form Spatio-Temporal Video Grounding | arXiv: 2602.23294 | 📄 paper_cache/ECCV2026/2602.23294.txt
- Towards Memory-Efficient Autoregressive Video Generation via Instance-Specific Parametric Absorption | arXiv: 2607.00712 | 📄 paper_cache/ECCV2026/2607.00712.txt
- Towards Metric-Agnostic Trajectory Forecasting | arXiv: 2607.01133 | 📄 paper_cache/ECCV2026/2607.01133.txt
- Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics | arXiv: 2512.13660 | 📄 paper_cache/ECCV2026/2512.13660.txt
- Training-free Cross-domain Few-shot Segmentation via Robust Semantic Representation and Matching | arXiv: 2606.24297 | 📄 paper_cache/ECCV2026/2606.24297.txt
- Triangular Consistency as a Universal Constraint for Learning Optical Flow | arXiv: 2606.19938 | 📄 paper_cache/ECCV2026/2606.19938.txt
- Understanding Cross-Rig Generalization in Automotive Perception: a Multi-Rig Benchmark and Rig Variation Metrics | arXiv: 2606.27554 | 📄 paper_cache/ECCV2026/2606.27554.txt
- UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving | arXiv: 2601.04453 | 📄 paper_cache/ECCV2026/2601.04453.txt
- UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation | arXiv: 2603.22282 | 📄 paper_cache/ECCV2026/2603.22282.txt
- UniPR-3D: Towards Universal Visual Place Recognition with Visual Geometry Grounded Transformer | arXiv: 2512.21078 | 📄 paper_cache/ECCV2026/2512.21078.txt
- UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation | arXiv: 2606.31451 | 📄 paper_cache/ECCV2026/2606.31451.txt
- UniTeD: Unified Temporal Diffusion for Joint Perception and Planning in Autonomous Driving | arXiv: 2606.25736 | 📄 paper_cache/ECCV2026/2606.25736.txt
- UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation | arXiv: 2606.24333 | 📄 paper_cache/ECCV2026/2606.24333.txt
- UniTriSplat: A Unified 3D Gaussian Splatting Framework with Uniform Spherical Rasterization for Universal Cameras | arXiv: 2606.29794 | 📄 paper_cache/ECCV2026/2606.29794.txt
- UNITY: Attention Flow Networks for Adaptive Conditioning in Diffusion | arXiv: 2606.20971 | 📄 paper_cache/ECCV2026/2606.20971.txt
- Universal Image Immunization against Diffusion-based Image Editing via Semantic Injection | arXiv: 2602.14679 | 📄 paper_cache/ECCV2026/2602.14679.txt
- Unveiling Transferability in Trajectory Prediction via Latent Scene Embeddings | arXiv: 2606.30777 | 📄 paper_cache/ECCV2026/2606.30777.txt
- URoPE: Universal Relative Position Embedding across Geometric Spaces | arXiv: 2604.18747 | 📄 paper_cache/ECCV2026/2604.18747.txt
- Vector Scaffolding: Inter-Scale Orchestration for Differentiable Image Vectorization | arXiv: 2605.11913 | 📄 paper_cache/ECCV2026/2605.11913.txt
- VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement | arXiv: 2607.00446 | 📄 paper_cache/ECCV2026/2607.00446.txt
- ViDiT:从图像编辑对中学习可迁移的连续视觉方向 | arXiv: 2403.19645
- ViewSplat: View-Adaptive 3D Gaussian Splatting for Feed-Forward Synthesis | arXiv: 2603.25265 | 📄 paper_cache/ECCV2026/2603.25265.txt
- ViQ: Text-Aligned Visual Quantized Representations at Any Resolution | arXiv: 2606.27313 | 📄 paper_cache/ECCV2026/2606.27313.txt
- VisCritic: Visual State Comparison as Process Reward for GUI Agents | arXiv: 2606.24525 | 📄 paper_cache/ECCV2026/2606.24525.txt
- VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning | arXiv: 2603.01195 | 📄 paper_cache/ECCV2026/2603.01195.txt
- Visual Prompt Discovery via Semantic Exploration | arXiv: 2603.16250 | 📄 paper_cache/ECCV2026/2603.16250.txt
- Vitality-Aware Compression for Efficient Image-to-Shape Diffusion Transformers | arXiv: 2607.00382 | 📄 paper_cache/ECCV2026/2607.00382.txt
- VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors | arXiv: 2510.00458 | 📄 paper_cache/ECCV2026/2510.00458.txt
- VOCA: Visual Odometry with Codec Awareness | arXiv: 2607.00189 | 📄 paper_cache/ECCV2026/2607.00189.txt
- VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction | arXiv: 2509.19297 | 📄 paper_cache/ECCV2026/2509.19297.txt
- VTEdit-Bench: A Comprehensive Benchmark for Multi-Reference Image Editing Models in Virtual Try-On | arXiv: 2603.11734 | 📄 paper_cache/ECCV2026/2603.11734.txt
- Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs | arXiv: 2607.00302 | 📄 paper_cache/ECCV2026/2607.00302.txt
- Walking in the Implicit: Interactive World Exploration via Neural Scene Representation | arXiv: 2606.30045 | 📄 paper_cache/ECCV2026/2606.30045.txt
- WarpI2I: Image Warping for Image-to-Image Translation | arXiv: 2606.31018 | 📄 paper_cache/ECCV2026/2606.31018.txt
- What if? Emulative Simulation with World Models for Situated Reasoning | arXiv: 2603.06445 | 📄 paper_cache/ECCV2026/2603.06445.txt
- When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models | arXiv: 2604.03316 | 📄 paper_cache/ECCV2026/2604.03316.txt
- Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation | arXiv: 2606.30248 | 📄 paper_cache/ECCV2026/2606.30248.txt
- Zero-Shot Depth from Defocus | arXiv: 2603.26658 | 📄 paper_cache/ECCV2026/2603.26658.txt
- Zero-Shot Quantization for Object Detectors using Off-the-Shelf Generative Models | arXiv: 2606.31456 | 📄 paper_cache/ECCV2026/2606.31456.txt
- ZR-0: Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision | arXiv: 2606.30552
- 光密纳米线超材料:多重散射下的偏振保持 | arXiv: 2606.29019
- 单调最小调节时间PI整定:基于切触恒等式的FOTD系统无超调控制 | arXiv: 2606.22217
- 基于扰动高斯集合的层析成像主动视角选择 | arXiv: 2603.06852
- 无滤波快照高光谱成像:基于引导Patch扩散模型 | arXiv: 2412.02798
- 混合事件-帧传感器:统一噪声建模、标定与仿真 | arXiv: 2511.18037