ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills Paper • 2610.12403 • Published 3 days ago • 14
GeoSET: Generalist Foundation Model for SAR-to-EO Image Translation Paper • 2609.37496 • Published 15 days ago • 7
Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features Paper • 2609.33463 • Published 14 days ago • 12
AutoGUIWorld: Image Generators as Visual World Models for GUI Agent Paper • 2610.01215 • Published 10 days ago • 63
VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control Paper • 2609.19554 • Published 24 days ago • 43
Generative Late-Interaction Embeddings For Visual Document Retrieval Paper • 2609.11808 • Published Sep 10 • 27
Cadence: Error-Bounded Lossy Compression of Demand Time Series with a Time-Series Foundation Model Paper • 2609.06008 • Published Sep 5 • 19
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published Aug 18 • 161
Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis Paper • 2608.18940 • Published Aug 19 • 35
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published Aug 14 • 287