MEMORY MEETS RESPONSIVE CONTROL

D2-VLA

Dual-Memory Dual-Frequency
Vision-Language-Action Model for
Long Dynamic Manipulation

Remember Longer. React Faster.

Zijian Ye1,*Chengqi Wei2,*Wei Huang1,†Anlin Zheng1,†Chunyu Zou1Liangyu Wu2Zikang Zhao2Zhenjie Peng2Yushuo Yang2Shuman Zhao1Zhongrui Wang2,✉Xiaojuan Qi1,✉

1 The University of Hong Kong2 Southern University of Science and Technology

* Equal contribution · † Project lead · ✉ Corresponding author

LONG DYNAMIC REAL-WORLD DEMONSTRATIONS4 tasks · 8 videos

Match and Pick

REAL ROBOT · D²-VLA
Demonstration 01 Open video ↗
Demonstration 02 Open video ↗
PERSISTENT MEMORY · RESPONSIVE CONTROLOverview figure to be added
Long-horizon context and fresh visual feedback, connected through the KV-cache interface.
29.3%DOMINODynamic manipulation
60.0%DOMINO-LongMemory + dynamics
97.5%LIBERO-LongLong-horizon manipulation
74.3%RoboTwin 2.050 clean bimanual tasks

01 / THE IDEA

Remember the past.
Act on the present.

Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-action policies often rely on the latest observation, and refreshing their visual context typically requires another costly vision-language model pass.

D²-VLA combines dual memory and dual-frequency control at the KV-cache interface of a pretrained VLA. Block-wise causal KV caching encodes observations incrementally; separate historical read views provide the VLM and action expert with the context each needs. Between VLM updates, a gated adapter incorporates fresh visual features, while a short fast-memory queue supports action replanning.

We introduce DOMINO-Long, a ten-task benchmark in which earlier visual cues guide later manipulation. Evaluation spans dynamic simulation, static long-horizon benchmarks, and eight real-robot tasks.

02 / METHOD

Two memories.
Two update rates.

D²-VLA co-designs historical context and visual refresh at the native KV-cache interface of a pretrained π₀.₅ policy.

DUAL MEMORY × DUAL FREQUENCYMethod figure to be added
Separate historical read views for the VLM and action expert; fresh visual conditioning between VLM updates.

01

Causal temporal memory

Encode only the newest observation block and reuse historical KV states. Block-wise causal attention carries task context across observations without re-encoding the entire history.

02

Consumer-specific reads

The VLM and action expert select historical visual tokens using independent attention statistics. Stored blocks stay complete; instruction tokens and current blocks are preserved.

03

Fresh visual feedback

A gated adapter refines the latest history-conditioned KV anchor with current image features. A bounded fast-memory queue supports replanning between periodic VLM refreshes.

Main dynamic
update schedule
0 · VLM + action25 · Adapter + action50 · Adapter + action75 · VLM + action

Intervals are measured in environment steps, not wall-clock frequency. The action expert replans every 25 steps; the VLM refreshes every 75.

03 / STATIC REAL-WORLD DEMONSTRATIONS

Long-horizon static tasks.

Four tasks on a Unitree G1D robot evaluate ordered execution, object assignments, and spatial memory.

Ordered Cup StackingVideo to be added

Ordered Cup Stacking

80% SR

Stack cups in a specified order across successive manipulations.

Cross-Plate Object SwapVideo to be added

Cross-Plate Object Swap

75% SR

Exchange the objects between two plates, preserving their assignments.

Ordered Object RetrievalVideo to be added

Ordered Object Retrieval

65% SR

Retrieve objects from a plate in the prescribed sequence.

Original-Layout RestorationVideo to be added

Original-Layout Restoration

65% SR

Return rearranged blocks to their original locations.

04 / EXPERIMENTAL RESULTS

Across tasks.
Across time scales.

D²-VLA improves complete-task success across dynamic and static simulation benchmarks. All numbers below follow the manuscript and its evaluation settings.

DOMINO

Success rate
π₀.₅9.6%
PUMA17.2%
D²-VLA29.3%

35 dynamic tasks · Aloha-AgileX · clean L1. D²-VLA also achieves a Manipulation Score of 40.6.

DOMINO-Long

Success rate
PUMA20.6%
π₀.₅35.4%
D²-VLA60.0%

10 tasks · Earlier cues guide later actions. +24.6 percentage points over π₀.₅.

Success-rate comparison across simulation benchmarks
Benchmarkπ₀.₅ SRD²-VLA SRGain
DOMINO9.6%29.3%+19.7 pp
DOMINO-Long35.4%60.0%+24.6 pp
LIBERO-Long92.4%97.5%+5.1 pp
RoboTwin 2.0 · clean57.0%74.3%+17.3 pp

Evaluation setting: LIBERO-Long (10 tasks) and RoboTwin 2.0 (50 clean tasks) disable the high-rate pathway while retaining causal memory and online historical KV selection. Dynamic benchmarks use both pathways. SR = complete-task success rate; pp = percentage points.

Full DOMINO comparison
DOMINO results from manuscript Table 1
MethodSR (%) ↑MS ↑
OpenVLA1.56.1
RDT-1B5.317.7
π₀8.224.0
π₀.₅9.626.2
InternVLA-M15.427.6
VLA-Adapter4.424.3
π₀-FAST3.520.9
OpenVLA-OFT9.124.1
StarVLA-OFT10.930.5
PUMA17.235.0
D²-VLA29.340.6

MS is DOMINO’s Manipulation Score, which captures manipulation quality beyond binary task completion. Source: Table 1(a).

Selected static benchmark comparisons

LIBERO-Long

Selected LIBERO-Long results
MethodSR (%) ↑
OpenVLA53.7
CronusVLA68.7
π₀85.2
GR00T-N190.6
π₀.₅92.4
MemoryVLA93.4
D²-VLA97.5

RoboTwin 2.0 · clean

Selected RoboTwin clean results
MethodParameters (B)SR (%) ↑
Diffusion Policy0.128.0
ACT0.129.7
DP30.355.2
π₀3.246.4
π₀.₅3.457.0
StarVLA-α3.850.3
D²-VLA3.574.3

Selected comparisons from manuscript Table 2; not a live leaderboard.

05 / CITATION

Cite D²-VLA

If this work is useful for your research, please consider citing our paper.

BIBTEX
@misc{ye2026d2vla,
  title = {{D$^2$-VLA}: Dual-Memory Dual-Frequency Vision-Language-Action Model For Long Dynamic Manipulation},
  author = {Zijian Ye and Chengqi Wei and Wei Huang and Anlin Zheng and
            Chunyu Zou and Liangyu Wu and Zikang Zhao and Zhenjie Peng and
            Yushuo Yang and Shuman Zhao and Zhongrui Wang and Xiaojuan Qi},
  year = {2026},
  eprint = {2609.34792},
  archivePrefix = {arXiv},
  primaryClass = {cs.CV},
  doi = {10.48550/arXiv.2609.34792},
  url = {https://arxiv.org/abs/2609.34792}
}