3

Are Frontier VLM Agents Ready to Be Robot Generalists? An Empirical Study with the Embodied Agent Arena

Oct 1, 2026

PIVOT: Pivot-Aware On Policy Self Distillation for Multi-Turn VLM Agents

Sep 28, 2026

In-Context Learning for Robots: Methods and Applications

Sep 28, 2026

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

Sep 24, 2026

Robo-Harness K1: Harnessing Robot-Use Agents via Perception Augmentation

Sep 24, 2026

RACaP: Agentic Reasoning, Acting, and Coding as Policies for Evolvable Robot Learning

Sep 24, 2026

The Past Frames the Future: Memory for Autoregressive Video Generation

Sep 23, 2026

AffordanceWAM: Affordance-Aware Joint World-Action Modeling for Robot Manipulation

Sep 16, 2026

ReWorld: An Interactive World Model with Long-Horizon Memory

Aug 24, 2026

Platonic Representation Hypothesis on World Models

Aug 24, 2026

GenRouter: Unified Workflow Routing for Agentic Image Generation

Aug 17, 2026

The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering

Aug 5, 2026

UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation

Jul 7, 2026

Sat2City v2: Native 3D City Asset Generation from a Single Satellite Image

Jun 23, 2026

Intend, Reflect, Refine: An Adaptive Multimodal Reflection Framework for Autonomous Driving

Jun 22, 2026

The Right Inference Strategy Is All You Need: Nearly Training-Free Domain-Wise Inference for EgoCross Challenge

May 30, 2026

RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes

May 30, 2026

LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation

May 18, 2026

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data

May 13, 2026

Focusable Monocular Depth Estimation

May 12, 2026

AnimationBench: Are Video Models Good at Character-Centric Animation?

Apr 16, 2026

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators

Mar 31, 2026

DVD: Deterministic Video Depth Estimation with Generative Priors

Mar 12, 2026

CalliMaster: Mastering Page-level Chinese Calligraphy via Layout-guided Spatial Planning

Mar 12, 2026

SPOT-Occ: Sparse Prototype-guided Transformer for Camera-based 3D Occupancy Prediction

Feb 4, 2026

A Mechanistic View on Video Generation as World Models: State and Dynamics

Jan 22, 2026

VideoMemory: Toward Consistent Video Generation via Memory Integration

Jan 7, 2026

Spatial4D-Bench: A Versatile 4D Spatial Intelligence Benchmark

Dec 31, 2025

SDPose: Exploiting Diffusion Priors for Out-of-Domain and Robust Pose Estimation

Sep 29, 2025

OmniBooth: Learning Latent Control for Image Synthesis with Multi-modal Instruction

Oct 1, 2024