01
Causal temporal memory
Encode only the newest observation block and reuse historical KV states. Block-wise causal attention carries task context across observations without re-encoding the entire history.
MEMORY MEETS RESPONSIVE CONTROL
Dual-Memory Dual-Frequency
Vision-Language-Action Model for
Long Dynamic Manipulation
Remember Longer. React Faster.
1 The University of Hong Kong2 Southern University of Science and Technology
* Equal contribution · † Project lead · ✉ Corresponding author
01 / THE IDEA
Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-action policies often rely on the latest observation, and refreshing their visual context typically requires another costly vision-language model pass.
D²-VLA combines dual memory and dual-frequency control at the KV-cache interface of a pretrained VLA. Block-wise causal KV caching encodes observations incrementally; separate historical read views provide the VLM and action expert with the context each needs. Between VLM updates, a gated adapter incorporates fresh visual features, while a short fast-memory queue supports action replanning.
We introduce DOMINO-Long, a ten-task benchmark in which earlier visual cues guide later manipulation. Evaluation spans dynamic simulation, static long-horizon benchmarks, and eight real-robot tasks.
02 / METHOD
D²-VLA co-designs historical context and visual refresh at the native KV-cache interface of a pretrained π₀.₅ policy.
01
Encode only the newest observation block and reuse historical KV states. Block-wise causal attention carries task context across observations without re-encoding the entire history.
02
The VLM and action expert select historical visual tokens using independent attention statistics. Stored blocks stay complete; instruction tokens and current blocks are preserved.
03
A gated adapter refines the latest history-conditioned KV anchor with current image features. A bounded fast-memory queue supports replanning between periodic VLM refreshes.
Intervals are measured in environment steps, not wall-clock frequency. The action expert replans every 25 steps; the VLM refreshes every 75.
03 / STATIC REAL-WORLD DEMONSTRATIONS
Four tasks on a Unitree G1D robot evaluate ordered execution, object assignments, and spatial memory.
Stack cups in a specified order across successive manipulations.
Exchange the objects between two plates, preserving their assignments.
Retrieve objects from a plate in the prescribed sequence.
Return rearranged blocks to their original locations.
04 / EXPERIMENTAL RESULTS
D²-VLA improves complete-task success across dynamic and static simulation benchmarks. All numbers below follow the manuscript and its evaluation settings.
35 dynamic tasks · Aloha-AgileX · clean L1. D²-VLA also achieves a Manipulation Score of 40.6.
10 tasks · Earlier cues guide later actions. +24.6 percentage points over π₀.₅.
| Benchmark | π₀.₅ SR | D²-VLA SR | Gain |
|---|---|---|---|
| DOMINO | 9.6% | 29.3% | +19.7 pp |
| DOMINO-Long | 35.4% | 60.0% | +24.6 pp |
| LIBERO-Long | 92.4% | 97.5% | +5.1 pp |
| RoboTwin 2.0 · clean | 57.0% | 74.3% | +17.3 pp |
Evaluation setting: LIBERO-Long (10 tasks) and RoboTwin 2.0 (50 clean tasks) disable the high-rate pathway while retaining causal memory and online historical KV selection. Dynamic benchmarks use both pathways. SR = complete-task success rate; pp = percentage points.
| Method | SR (%) ↑ | MS ↑ |
|---|---|---|
| OpenVLA | 1.5 | 6.1 |
| RDT-1B | 5.3 | 17.7 |
| π₀ | 8.2 | 24.0 |
| π₀.₅ | 9.6 | 26.2 |
| InternVLA-M1 | 5.4 | 27.6 |
| VLA-Adapter | 4.4 | 24.3 |
| π₀-FAST | 3.5 | 20.9 |
| OpenVLA-OFT | 9.1 | 24.1 |
| StarVLA-OFT | 10.9 | 30.5 |
| PUMA | 17.2 | 35.0 |
| D²-VLA | 29.3 | 40.6 |
MS is DOMINO’s Manipulation Score, which captures manipulation quality beyond binary task completion. Source: Table 1(a).
| Method | SR (%) ↑ |
|---|---|
| OpenVLA | 53.7 |
| CronusVLA | 68.7 |
| π₀ | 85.2 |
| GR00T-N1 | 90.6 |
| π₀.₅ | 92.4 |
| MemoryVLA | 93.4 |
| D²-VLA | 97.5 |
| Method | Parameters (B) | SR (%) ↑ |
|---|---|---|
| Diffusion Policy | 0.1 | 28.0 |
| ACT | 0.1 | 29.7 |
| DP3 | 0.3 | 55.2 |
| π₀ | 3.2 | 46.4 |
| π₀.₅ | 3.4 | 57.0 |
| StarVLA-α | 3.8 | 50.3 |
| D²-VLA | 3.5 | 74.3 |
Selected comparisons from manuscript Table 2; not a live leaderboard.
05 / CITATION
If this work is useful for your research, please consider citing our paper.
@misc{ye2026d2vla,
title = {{D$^2$-VLA}: Dual-Memory Dual-Frequency Vision-Language-Action Model For Long Dynamic Manipulation},
author = {Zijian Ye and Chengqi Wei and Wei Huang and Anlin Zheng and
Chunyu Zou and Liangyu Wu and Zikang Zhao and Zhenjie Peng and
Yushuo Yang and Shuman Zhao and Zhongrui Wang and Xiaojuan Qi},
year = {2026},
eprint = {2609.34792},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
doi = {10.48550/arXiv.2609.34792},
url = {https://arxiv.org/abs/2609.34792}
}