


default search action
Wenxuan Song
Person information
SPARQL queries 
Refine list

refinements active!
zoomed in on ?? of ?? records
view refined list in
2020 – today
- 2026
[c15]Wenxuan Song, Ziyang Zhou, Han Zhao, Jiayi Chen, Pengxiang Ding, Haodong Yan, Yuxin Huang, Feilong Tang, Donglin Wang, Haoang Li:
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver. AAAI 2026: 18549-18557
[c14]Yihao Wang, Pengxiang Ding, Lingxiao Li, Can Cui, Zirui Ge, Xinyang Tong, Wenxuan Song, Han Zhao, Wei Zhao, Pengxu Hou, Siteng Huang, Yifan Tang, Wenhui Wang, Ru Zhang, Jianyi Liu, Donglin Wang:
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model. AAAI 2026: 18638-18646
[c13]Duo Wang
, Kyrian Liang, Wenxuan Song
, Qingxiao Zheng, Zyra Sheikh, Jen Whiting, Victor Lu, Mike Yao, Caroline Cao:
EmSim: AI-Driven VR Simulation for De-Escalation Training in Law Enforcement. AIxVR 2026: 39-47
[c12]Duo Wang
, Jianwei Ni, Yuhan Zhou, Weiyu Ding, Wenxuan Song
, Shandra Jamison, Mike Yao, Caroline Cao:
SpatialTutor: Object-Aware Mixed Reality Training for Procedural Medical Skills Training with AI-Driven Support. AIxVR 2026: 57-66
[c11]Karthik S. Bhat
, Jiayue Melissa Shi
, Wenxuan Song
, Dong Whi Yoo
, Koustuv Saha
:
"In my defense, only three hours on Instagram": Designing Toward Digital Self-Awareness and Wellbeing. CHI 2026: 622:1-622:20
[c10]Duo Wang
, Kyrian Liang
, Qingxiao Zheng
, Wenxuan Song
, Jen Whiting
, Mike Yao
, Caroline G. L. Cao
:
Does Sequencing Matter? Evaluating AI and Human Simulations for High-Stakes Communication Training in Law Enforcement. CHI 2026: 1586:1-1586:19
[c9]Duo Wang
, Wenxuan Song
, Jianwei Ni
, Qingxiao Zheng
, Yuhan Zhou
, Kyrian Liang
, Mike Yao
, Caroline G. L. Cao
:
Should the AI Speak First? Evaluating Proactive vs. Reactive Facilitation in Mixed-Reality Medical Training. CHI 2026: 1631:1-1631:19
[i37]Shanshan Zhu, Wenxuan Song, Jiayue Melissa Shi, Dong Whi Yoo, Karthik S. Bhat, Koustuv Saha:
Designing KRIYA: An AI Companion for Wellbeing Self-Reflection. CoRR abs/2601.14589 (2026)
[i36]Han Zhao, Jingbo Wang, Wenxuan Song, Shuai Chen, Yang Liu, Yan Wang, Haoang Li, Donglin Wang:
FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment. CoRR abs/2602.17259 (2026)
[i35]Wenxuan Song, Jiayi Chen, Xiaoquan Sun, Huashuo Lei, Yikai Qin, Wei Zhao, Pengxiang Ding, Han Zhao, Tongxin Wang, Pengxu Hou, Zhide Zhong, Haodong Yan, Donglin Wang, Jun Ma, Haoang Li:
Rethinking the Practicality of Vision-language-action Model: A Comprehensive Benchmark and An Improved Baseline. CoRR abs/2602.22663 (2026)
[i34]Zehua Fan, Wenqi Lyu, Wenxuan Song, Linge Zhao, Yifei Yang, Xi Wang, Junjie He, Lida Huang, Haiyan Liu, Bingchuan Sun, Guangjun Bao, Xuanyao Mao, Liang Xu, Yan Wang, Feng Gao:
PROSPECT: Unified Streaming Vision-Language Navigation via Semantic-Spatial Fusion and Latent Predictive Representation. CoRR abs/2603.03739 (2026)
[i33]Haodong Yan, Zhide Zhong, Jiaguan Zhu, Junjie He, Weilin Yuan, Wenxuan Song, Xin Gong, Yingjie Cai, Guanyi Zhao, Xu Yan, Bingbing Liu, Ying-Cong Chen, Haoang Li:
S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight. CoRR abs/2603.16195 (2026)
[i32]Zirui Ge, Pengxiang Ding, Baohua Yin, Qishen Wang, Zhiyong Xie, Yemin Wang, Jinbo Wang, Hengtao Li, Runze Suo, Wenxuan Song, Han Zhao, Shangke Lyu, Zhaoxin Fan, Haoang Li, Ran Cheng, Cheng Chi, Huibin Ge, Yaozhi Luo, Donglin Wang:
VAMPO: Policy Optimization for Improving Visual Dynamics in Video Action Models. CoRR abs/2603.19370 (2026)
[i31]Yang Liu, Pengxiang Ding, Tengyue Jiang, Xudong Wang, Wenxuan Song, Minghui Lin, Han Zhao, Hongyin Zhang, Zifeng Zhuang, Wei Zhao, Siteng Huang, Jinkui Shi, Donglin Wang:
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation. CoRR abs/2603.25406 (2026)
[i30]Wenxuan Song, Jiayi Chen, Shuai Chen, Jingbo Wang, Pengxiang Ding, Han Zhao, Yikai Qin, Xinhu Zheng, Donglin Wang, Yan Wang, Haoang Li:
Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance. CoRR abs/2603.25661 (2026)
[i29]Jiayi Chen, Wenxuan Song, Shuai Chen, Jingbo Wang, Zhijun Li, Haoang Li:
DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching. CoRR abs/2603.26320 (2026)
[i28]Wenxuan Song, Han Zhao, Fuhao Li, Ziyang Zhou, Xi Wang, Jing Lyu, Pengxiang Ding, Yan Wang, Donglin Wang, Haoang Li:
CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models. CoRR abs/2605.10903 (2026)
[i27]Huashuo Lei, Wenxuan Song, Huarui Zhang, Jieyuan Pei, Jiayi Chen, Haodong Yan, Han Zhao, Pengxiang Ding, Zhipeng Zhang, Lida Huang, Donglin Wang, Yan Wang, Haoang Li:
RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark. CoRR abs/2605.10921 (2026)
[i26]Jingzhi Huang, Junkai Huang, Wenxuan Song, Haoyang Yang, Hailong Huang, Haoang Li, Yi Wang:
SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation. CoRR abs/2605.17249 (2026)
[i25]Jinzhao Li, Yinuo Chen, Wenxuan Song, Yijia Lei, Yichi Zhang, Honglei Yan, Panwang Pan, Miao Liu:
IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams. CoRR abs/2605.27074 (2026)
[i24]Nan Sun, Yuan Zhang, Yongkun Yang, Wentao Zhao, Peiyan Li, Jun Guo, Wenxuan Song, Pengxiang Ding, Runze Suo, Yifei Su, Xin Xiao, Xinghang Li, Huaping Liu:
Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation. CoRR abs/2606.03784 (2026)
[i23]Shuxiang Zhang, Yiting Yin, Wenxuan Song, Yuhang Wu, Miao Liu:
PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models. CoRR abs/2606.23092 (2026)
[i22]Jun Guo, Piaopiao Jin, Jason Li, Peiyan Li, Yingyan Li, Futeng Liu, Wanli Peng, Optimus Qin, Yifei Su, Nan Sun, Qiao Sun, Runze Suo, Heyun Wang, Yunhong Wang, Rujie Wu, Caoyu Xia, Lina Zhang, Jack Zhao, Guoliang Chen, Wenlong Chen, Xinze He, Bin Li, Qing Li, Zhuorong Li, Heng Qu, Wenxuan Song, Diyun Xiang, Yifan Xie, Peiran Xu, Hangjun Ye, Wen Ye, Han Zhao, Quanyun Zhou:
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories. CoRR abs/2607.15330 (2026)- 2025
[j3]Qianwen Lv, Zhixin Qi, Haibo Li, Wenxuan Song, Di Huang, Zihao Ding, Qian Shi:
Evaluating The potential of multi-source remote sensing and urban big data for urban land use classification. Int. J. Digit. Earth 18(1) (2025)
[j2]Wenxuan Song, Jiaming Zhang, Lihao Hua, Zhihua Xiong
, Wenlei Zhao
:
Emergence of Classical Random Walk from Non-Hermitian Effects in Quantum Kicked Rotor. Entropy 27(3): 288 (2025)
[c8]Huapeng Li, Wenxuan Song, Tianao Xu, Alexandre Elsig, Jonas Kulhanek:
WaterSplatting: Fast Underwater 3D Scene Reconstruction Using Gaussian Splatting. 3DV 2025: 969-978
[c7]Feilong Tang, Chengzhi Liu, Zhongxing Xu, Ming Hu, Zile Huang, Haochen Xue, Ziyang Chen, Zelin Peng, Zhiwei Yang, Sijin Zhou, Wenxue Li
, Yulong Li, Wenxuan Song, Shiyan Su, Wei Feng, Jionglong Su, Mingquan Lin, Yifan Peng, Xuelian Cheng, Imran Razzak, Zongyuan Ge:
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding. CVPR 2025: 26147-26159
[c6]Wenxue Li, Tian Ye, Xinyu Xiong, Jinbin Bai, Feilong Tang, Wenxuan Song, Zhaohu Xing, Lie Ju, Guanbin Li, Lei Zhu:
GlassWizard: Harvesting Diffusion Priors for Glass Surface Detection. ICCV 2025: 17848-17858
[c5]Han Zhao, Wenxuan Song, Donglin Wang, Xinyang Tong, Pengxiang Ding, Xuelian Cheng, Zongyuan Ge:
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models. ICRA 2025: 11212-11218
[c4]Wenxuan Song, Jiayi Chen, Pengxiang Ding, Han Zhao, Wei Zhao, Zhide Zhong, Zongyuan Ge, Zhijun Li, Donglin Wang, Lujia Wang, Jun Ma, Haoang Li:
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding. IROS 2025: 13162-13169
[i21]Wenxuan Song, Jiayi Chen, Pengxiang Ding, Han Zhao, Wei Zhao, Zhide Zhong, Zongyuan Ge, Jun Ma, Haoang Li:
Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding. CoRR abs/2503.02310 (2025)
[i20]Han Zhao, Wenxuan Song, Donglin Wang, Xinyang Tong, Pengxiang Ding, Xuelian Cheng, Zongyuan Ge
:
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models. CoRR abs/2503.08007 (2025)
[i19]Can Cui, Pengxiang Ding, Wenxuan Song, Shuanghao Bai, Xinyang Tong, Zirui Ge, Runze Suo, Wanqi Zhou, Yang Liu
, Bofang Jia, Han Zhao, Siteng Huang, Donglin Wang:
OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation. CoRR abs/2505.03912 (2025)
[i18]Wenxuan Song, Jiayi Chen, Wenxue Li, Xu He, Han Zhao, Can Cui, Pengxiang Ding, Shiyan Su, Feilong Tang, Xuelian Cheng, Donglin Wang, Zongyuan Ge, Xinhu Zheng, Zhe Liu, Hesheng Wang, Haoang Li:
RationalVLA: A Rational Vision-Language-Action Model with Dual System. CoRR abs/2506.10826 (2025)
[i17]Wenxuan Song, Jiayi Chen, Pengxiang Ding, Yuxin Huang, Han Zhao, Donglin Wang, Haoang Li:
CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding. CoRR abs/2506.13725 (2025)
[i16]Wenxuan Song, Ziyang Zhou
, Han Zhao, Jiayi Chen, Pengxiang Ding, Haodong Yan, Yuxin Huang, Feilong Tang, Donglin Wang, Haoang Li:
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver. CoRR abs/2508.10333 (2025)
[i15]Zhide Zhong, Haodong Yan, Junfeng Li, Xiangchen Liu, Xin Gong, Wenxuan Song, Jiayi Chen, Haoang Li:
FlowVLA: Thinking in Motion with a Visual Chain of Thought. CoRR abs/2508.18269 (2025)
[i14]Yihao Wang, Pengxiang Ding, Lingxiao Li, Can Cui, Zirui Ge, Xinyang Tong, Wenxuan Song, Han Zhao, Wei Zhao, Pengxu Hou, Siteng Huang, Yifan Tang, Wenhui Wang, Ru Zhang, Jianyi Liu, Donglin Wang:
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model. CoRR abs/2509.09372 (2025)
[i13]Karthik S. Bhat, Jiayue Melissa Shi, Wenxuan Song, Dong Whi Yoo, Koustuv Saha:
"In my defense, only three hours on Instagram": Designing Toward Digital Self-Awareness and Wellbeing. CoRR abs/2509.21860 (2025)
[i12]Shuanghao Bai, Wenxuan Song, Jiayi Chen, Yuheng Ji, Zhide Zhong, Jin Yang, Han Zhao, Wanqi Zhou, Wei Zhao, Zhe Li, Pengxiang Ding, Cheng Chi, Haoang Li, Chang Xu, Xiaolong Zheng, Donglin Wang, Shanghang Zhang, Badong Chen:
Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey. CoRR abs/2510.10903 (2025)
[i11]Fuhao Li, Wenxuan Song, Han Zhao, Jingbo Wang, Pengxiang Ding, Donglin Wang, Long Zeng, Haoang Li:
Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model. CoRR abs/2510.12276 (2025)
[i10]Han Zhao, Jiaxuan Zhang, Wenxuan Song, Pengxiang Ding, Donglin Wang:
VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation. CoRR abs/2510.14902 (2025)
[i9]Jiayi Chen, Wenxuan Song, Pengxiang Ding, Ziyang Zhou, Han Zhao, Feilong Tang, Donglin Wang, Haoang Li:
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process. CoRR abs/2511.01718 (2025)
[i8]Wenzhuo Sun, Mingjian Liang, Wenxuan Song, Xuelian Cheng, Zongyuan Ge:
RoomPlanner: Explicit Layout Planner for Easier LLM-Driven 3D Room Generation. CoRR abs/2511.17048 (2025)
[i7]Haodong Yan, Hang Yu, Zhide Zhong, Weilin Yuan, Xin Gong, Zehang Luo, Chengxi Heyu, Junfeng Li, Wenxuan Song, Shunbo Zhou, Haoang Li:
Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation. CoRR abs/2512.01677 (2025)
[i6]Minghui Lin, Pengxiang Ding, Shu Wang, Zifeng Zhuang, Yang Liu, Xinyang Tong, Wenxuan Song, Shangke Lyu, Siteng Huang, Donglin Wang:
HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models. CoRR abs/2512.09928 (2025)
[i5]Shuanghao Bai, Wenxuan Song, Jiayi Chen, Yuheng Ji, Zhide Zhong, Jin Yang, Han Zhao, Wanqi Zhou, Zhe Li, Pengxiang Ding, Cheng Chi, Chang Xu, Xiaolong Zheng, Donglin Wang, Haoang Li, Shanghang Zhang, Badong Chen:
Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives. CoRR abs/2512.22983 (2025)- 2024
[j1]Fujia Sun, Wenxuan Song:
A Super-Resolution and 3D Reconstruction Method Based on OmDF Endoscopic Images. Sensors 24(15): 4890 (2024)
[c3]Pengxiang Ding, Han Zhao, Wenjie Zhang, Wenxuan Song, Min Zhang, Siteng Huang, Ningxi Yang, Donglin Wang:
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots. ECCV (5) 2024: 352-367
[c2]Wenxuan Song, Han Zhao, Pengxiang Ding, Can Cui, Shangke Lyu, Yaning Fan, Donglin Wang:
GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot. IROS 2024: 11879-11886
[c1]Can Cui
, Siteng Huang
, Wenxuan Song
, Pengxiang Ding
, Min Zhang
, Donglin Wang
:
ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification. ACM Multimedia 2024: 1583-1592
[i4]Wenxuan Song, Han Zhao, Pengxiang Ding, Can Cui, Shangke Lyu, Yaning Fan, Donglin Wang:
GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot. CoRR abs/2403.13358 (2024)
[i3]Huapeng Li, Wenxuan Song, Tianao Xu, Alexandre Elsig, Jonas Kulhanek:
WaterSplatting: Fast Underwater 3D Scene Reconstruction Using Gaussian Splatting. CoRR abs/2408.08206 (2024)
[i2]Can Cui, Siteng Huang, Wenxuan Song, Pengxiang Ding, Min Zhang, Donglin Wang:
ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification. CoRR abs/2409.20081 (2024)
[i1]Xue Xia
, Daiwei Zhang, Wenxuan Song, Wei Huang, Lorenz Hurni:
MapSAM: Adapting Segment Anything Model for Automated Feature Detection in Historical Maps. CoRR abs/2411.06971 (2024)
Coauthor Index

manage site settings
To protect your privacy, all features that rely on external API calls from your browser are turned off by default. You need to opt-in for them to become active. All settings here will be stored as cookies with your web browser. For more information see our F.A.Q.
Unpaywalled article links
Add open access links from
to the list of external document links (if available).
Privacy notice: By enabling the option above, your browser will contact the API of unpaywall.org to load hyperlinks to open access articles. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the Unpaywall privacy policy.
Archived links via Wayback Machine
For web page which are no longer available, try to retrieve content from the
of the Internet Archive (if available).
Privacy notice: By enabling the option above, your browser will contact the API of archive.org to check for archived content of web pages that are no longer available. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the Internet Archive privacy policy.
Reference lists
Add a list of references from
,
, and
to record detail pages.
load references from crossref.org and opencitations.net
Privacy notice: By enabling the option above, your browser will contact the APIs of crossref.org, opencitations.net, and semanticscholar.org to load article reference information. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the Crossref privacy policy and the OpenCitations privacy policy, as well as the AI2 Privacy Policy covering Semantic Scholar.
Citation data
Add a list of citing articles from
and
to record detail pages.
load citations from opencitations.net
Privacy notice: By enabling the option above, your browser will contact the API of opencitations.net and semanticscholar.org to load citation information. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the OpenCitations privacy policy as well as the AI2 Privacy Policy covering Semantic Scholar.
OpenAlex data
Load additional information about publications from
.
Privacy notice: By enabling the option above, your browser will contact the API of openalex.org to load additional information. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the information given by OpenAlex.
last updated on 2026-08-09 23:15 CEST by the dblp team
all metadata released as open data under CC0 1.0 license
see also: Terms of Use | Privacy Policy | Imprint


Google
Google Scholar
Semantic Scholar
Internet Archive Scholar
CiteSeerX
ORCID






