Ying-Cong Chen
  • Bio
  • Students
  • News
  • Publications
  • Playground
  • Projects
    • Envision Paper Audit
    • Pandas
    • PyTorch
    • scikit-learn
  • Advising Statement for Prospective Students
  • Projects
  • Experience
  • Blog
    • 🎉 Easily create your own simple yet highly customizable blog
    • 🧠 Sharpen your thinking with a second brain
    • 📈 Communicate your results effectively with the best data visualizations
    • 👩🏼‍🏫 Teach academic courses
    • ✅ Manage your projects
  • News
    • Two papers selected as ICCV 2025 Highlight papers
    • Our open-source research ecosystem reaches 10K+ GitHub stars
    • Our CVPR paper "Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation" is getting noticed
    • Our CVPR paper "TransPixeler: Advancing Text-to-Video Generation with Transparency" is getting noticed
    • Our arxiv paper "Lotus Diffusion-based Visual Foundation Model for High-quality Dense Prediction" is getting noticed
    • Our paper Luciddreamer is selected as Spotlight at CVPR 2024 and featured on Hugging Face Daily
    • Invited Talk: International conference on Artificial Intelligence & Machine Learning (AIM-2024)
    • Invited Talk: International Conference on Applied Mathematics 2024
    • We win the 1st Place in RobotDrive Challenge (ICRA 2024)
    • I am invited to give a Annual Progressive Report in China3DV
    • Invited Talk: 2024年广州市国资并购联合会会员大会
    • 关于生成式模型及其社会影响的采访(广州日报)
    • Our Paper “Ref-NeuS” Nominated for Best Paper at ICCV 2023
  • Playground
    • Transpixar
    • Depth Prediction
    • Motion Inversion
    • Normal Prediction
  • Recent & Upcoming Talks
    • Example Talk
  • Selected Publications
    • In-Context Learning for Robots: Methods and Applications
    • World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal
    • The Past Frames the Future: Memory for Autoregressive Video Generation
    • Platonic Representation Hypothesis on World Models
    • ReWorld: An Interactive World Model with Long-Horizon Memory
    • GenRouter: Unified Workflow Routing for Agentic Image Generation
    • The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering
    • UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation
    • Less is More: Improving LLM Reasoning with Minimal Test-Time Intervention
    • SceneLM: 3D-Aware Language Models for Editable 3D Scene Synthesis
    • InfiniVerse: Occupancy Guided Unbounded Scene Generation for Autonomous Driving
    • No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs
    • WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents
    • Streaming Communication in Multi-Agent Reasoning
    • Policy and World Modeling Co-Training for Language Agents
    • RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes
    • LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
    • RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data
    • Focusable Monocular Depth Estimation
    • AnimationBench: Are Video Models Good at Character-Centric Animation?
    • DualCoT-VLA: Visual-Linguistic Chain of Thought via Parallel Reasoning for Vision-Language-Action Models
    • S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight
    • ZeRCP: Towards Communication-Efficient Collaborative Perception and Future Scene Prediction via Request-Free Spatial Filtering
    • CalliMaster: Mastering Page-level Chinese Calligraphy via Layout-guided Spatial Planning
    • DVD: Deterministic Video Depth Estimation with Generative Priors
    • SPOT-Occ: Sparse Prototype-guided Transformer for Camera-based 3D Occupancy Prediction
    • TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
    • A Mechanistic View on Video Generation as World Models: State and Dynamics
    • VideoMemory: Toward Consistent Video Generation via Memory Integration
    • Dual-balancing for multi-task learning
    • Find, Fix, Reason: Context Repair for Video Reasoning
    • ImpText: A Benchmark and Tool-Augmented Framework for Implicit Text Reasoning
    • RectifiedHR: Enable Efficient High-Resolution Synthesis via Energy Rectification
    • ScalingAR: Scaling Confidence for Autoregressive Image Generation
    • Show, Don't Tell: Morphing Latent Reasoning into Image Generation
    • T-Rex-Omni: Integrating Negative Visual Prompt in Generic Object Detection
    • UniCalli: A Unified Diffusion Framework for Column-Level Generation and Recognition of Chinese Calligraphy
    • Spatial4D-Bench: A Versatile 4D Spatial Intelligence Benchmark
    • STANCE: Motion Coherent Video Generation Via Sparse-to-Dense Anchored Encoding
    • SDPose: Exploiting Diffusion Priors for Out-of-Domain and Robust Pose Estimation
    • Co-Painter: Fine-Grained Controllable Image Stylization via Implicit Decoupling and Adaptive Injection
    • ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback
    • DisEnvisioner: Disentangled and Enriched Visual Prompt for Customized Image Generation
    • DivPro: diverse protein sequence design with direct structure recovery guidance
    • DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving
    • Event-Guided Consistent Video Enhancement with Modality-Adaptive Diffusion Pipeline
    • exGen: Flexible Multi-View Generation from Text and Image Inputs
    • GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMs
    • Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation
    • Large Language Models for Transforming Healthcare: A Perspective on DeepSeek‐R1
    • Lotus: Diffusion-based visual foundation model for high-quality dense prediction
    • LucidFusion: Reconstructing 3D Gaussians with Arbitrary Unposed Images
    • Motion inversion for video customization
    • MTMamba++: Enhancing Multi-Task Dense Scene Understanding via Mamba-Based Decoders
    • Orchestrating Audio: Multi-Agent Framework for Long-Video Audio Synthesis
    • PARM: Multi-Objective Test-Time Alignment via Preference-Aware Autoregressive Reward Model
    • POSTA: A Go-to Framework for Customized Artistic Poster Generation
    • PreGenie: An Agentic Framework for High-quality Visual Presentation Generation
    • PRM: Photometric Stereo based Large Reconstruction Model
    • RhythmGuassian: Repurposing Generalizable Gaussian Model For Remote Physiological Measurement
    • Sat2City: 3D City Generation from A Single Satellite Image with Cascaded Latent Diffusion
    • Scene Graph Guided Generation: Enable Accurate Relations Generation in Text-to-Image Models via Textural Rectification
    • SEED-Story: Multimodal Long Story Generation with Large Language Model
    • StreamGS: Online Generalizable Gaussian Splatting Reconstruction for Unposed Image Streams
    • SURGEON: Memory-Adaptive Fully Test-Time Adaptation via Dynamic Activation Sparsity
    • TransPixeler: Advancing Text-to-Video Generation with Transparency
    • Uni-IR: One Stage is Enough for Ambiguity-Reduced Inverse Rendering
    • Uni-Renderer: Unifying Rendering and Inverse Rendering Via Dual Stream Diffusion
    • OmniBooth: Learning Latent Control for Image Synthesis with Multi-modal Instruction
    • An Incremental Unified Framework for Small Defect Inspection
    • Bi-TTA: Bidirectional Test-Time Adapter for Remote Physiological Measurement
    • Defect Spectrum: A Granular Look of Large-scale Defect Datasets with Rich Semantics
    • MTMamba: Enhancing Multi-Task Dense Scene Understanding by Mamba-Based Decoders
    • Text-Anchored Score Composition: Tackling Condition Misalignment in Text-to-Image Diffusion Models
    • Bridging Data Gaps in Diffusion Models with Adversarial Noise-Based Transfer Learning
    • Learning to Remove Wrinkled Transparent Film with Polarized Prior
    • Low-Rank Approximation for Sparse Attention in Multi-Modal LLMs
    • LucidDreamer: Towards High-Fidelity Text-to-3D Generation via Interval Score Matching
    • Backdoor Contrastive Learning via Bi-level Trigger Optimization
    • Denoising Diffusion Step-aware Models
    • GNeRP: Gaussian-guided Neural Reconstruction of Reflective Objects with Noisy Polarization Priors
    • Adv3D: Generating 3D Adversarial Examples for 3D Object Detection in Driving Scenarios with NeRF
    • MantraNet: Label Name is Mantra: Unifying Point Cloud Segmentation across Heterogeneous Datasets
    • Rethinking Rendering in Generalizable Neural Surface Reconstruction: A Learning-based Solution
    • Ref-NeuS: Ambiguity-Reduced Neural Implicit Surface Learning for Multi-View Reconstruction with Reflection
    • Not All Steps are Created Equal: Selective Diffusion Distillation for Image Manipulation
    • Photo-Realistic Out-of-domain GAN inversion via Invertibility Decomposition
    • Lift3D: Synthesize 3D Training Data by Lifting 2D GAN to 3D Generative Radiance Field
    • Neuron Structure Modeling for Generalizable Remote Physiological Measurement
    • Real-time 6K Image Rescaling with Rate-distortion Optimization
    • CP-NeRF: Conditionally Parameterized Neural Radiance Fields for Cross-scene Novel View Synthesis
    • Artificial intelligence-enabled detection and assessment of Parkinson’s disease using nocturnal breathing signals
    • DecoupleNet: Decoupled Network for Domain Adaptive Semantic Segmentation
    • RC-MVSNet: Unsupervised Multi-View Stereo with Neural Rendering
    • Semi-supervised Monocular 3D Object Detection by Multi-view Consistency
    • Representation Compensation Networks for Continual Semantic Segmentation
    • Learning to Know Where to See: A Visibility-Aware Approach for Occluded Person Re-identification
    • SC-GAN: Image Synthesis via Semantic Composition
    • Delving into Deep Imbalanced Regression
    • PointINS: Point-based instance segmentation
    • Text-Guided Human Image Manipulation via Image-Text Shared Space
    • Attentive normalization for conditional image generation
    • Domain Adaptive Image-to-image Translation
    • Homomorphic Interpolation Network for Unpaired Image-to-image Translation
    • VCNet: A Robust Approach to Blind Image Inpainting
    • Homomorphic latent space interpolation for unpaired image-to-image translation
    • Semantic component decomposition for face attribute manipulation
    • View Independent Generative Adversarial Network for Novel View Synthesis
    • Facelet-bank for fast portrait manipulation
    • Person Re-Identification by Camera Correlation Aware Feature Augmentation
    • Makeup-go: Blind reversion of portrait edit
    • An asymmetric distance model for cross-view feature mapping in person reidentification
    • An enhanced deep feature representation for person re-identification
    • Mirror representation for modeling view-specific transform in person re-identification
  • Teaching
    • Learn JavaScript
    • Learn Python

Policy and World Modeling Co-Training for Language Agents

Jun 1, 2026·
Ning Lu
,
Baijiong Lin
,
Shengcai Liu
,
Jiahao Wu
,
Haoze Lv
,
Yanbin Wei
,
Lingting Zhu
,
Shengju Qian
,
Xin Wang
Ying-Cong Chen
Ying-Cong Chen
,
Qi Wang
,
Ke Tang
· 0 min read
PDF Cite arXiv
Type
1
Publication
EMNLP 2026 (main)
Last updated on Sep 30, 2026
Ying-Cong Chen
Authors
Ying-Cong Chen
Assistant Professor

← Streaming Communication in Multi-Agent Reasoning Jun 3, 2026
RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes May 30, 2026 →

© 2026 Me. This work is licensed under CC BY NC ND 4.0

Published with Hugo Blox Builder — the free, open source website builder that empowers creators.