RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 7 days ago • 274
Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory Paper • 2610.02521 • Published 7 days ago • 52
VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks Paper • 2610.00972 • Published 7 days ago • 51
EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling Paper • 2610.02298 • Published 7 days ago • 54
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 12 days ago • 324
When Does Dense Retrieval Need Asymmetric Geometry? A Bias-Variance Theory of Shared and Dual Projections Paper • 2609.32488 • Published 12 days ago • 26
Chinese-Jev: Bringing System One Model to Chinese-Language Tasks Paper • 2609.36965 • Published 9 days ago • 25
Marathoner: Ultra-Long-Horizon Autonomous Intelligence Paper • 2609.34378 • Published 10 days ago • 42
Think Before You Score: Thinking Reward Model for Visual Generation Paper • 2609.37372 • Published 9 days ago • 101