LiLo-VLA: Compositional Long-Horizon Manipulation via Linked Object-Centric Policies

Chaining atomic skills through a motion planner, so unseen skill orderings run zero-shot.

Best Paper Award, IROS 2026 Compositional and Modular Learning
Yue Yang 1,†
UNC Chapel Hill
Georgia Tech
UNC Chapel Hill
UNC Chapel Hill
UNC Chapel Hill
UNC Chapel Hill
1University of North Carolina at Chapel Hill, 2Georgia Institute of Technology, 3Carnegie Mellon University, †Corresponding author: yygx@cs.unc.edu

Abstract

General-purpose robots must master long-horizon manipulation, defined as tasks involving multiple kinematic structure changes (e.g., attaching or detaching objects) in unstructured environments. While Vision-Language-Action (VLA) models offer the potential to master diverse atomic skills, they struggle with the combinatorial complexity of sequencing them and are prone to cascading failures due to environmental sensitivity. To address these challenges, we propose LiLo-VLA (Linked Local VLA), a modular framework capable of zero-shot generalization to novel long-horizon tasks without ever being trained on them. Our approach decouples transport from interaction: a Reaching Module handles global motion, while an Interaction Module employs an object-centric VLA to process isolated objects of interest, ensuring robustness against irrelevant visual features and invariance to spatial configurations. Crucially, this modularity facilitates robust failure recovery through dynamic replanning and skill reuse, effectively mitigating the cascading errors common in end-to-end approaches. We introduce a 21-task simulation benchmark consisting of two challenging suites: LIBERO-Long++ and Ultra-Long. In these simulations, LiLo-VLA achieves a 69% average success rate, outperforming Pi0.5 by 41% and OpenVLA-OFT by 67%. Furthermore, real-world evaluations across 8 long-horizon tasks demonstrate an average success rate of 85%.

85% vs 0%
success on reordered skill sequences

LiLo-VLA vs Pi0.5 · same Suite-1 tasks, only the order changes

69% / 86%
overall success rate / average progress

Pi0.5 28% / 31% · OpenVLA-OFT 2% / 4%

16 skills
longest task in the benchmark

21 configurations over 9 scenarios · prior LIBERO-Long suites run 3–4

85%
on a real Franka Panda

8 configurations, up to 8 skills, trained on atomic skills only

Overview — sound on

Reordering should be free

A long-horizon task is a sequence of atomic skills, and the same skills usually admit many valid orderings. Put the mug away before the bowl or after it; the skills are identical either way. A policy that has learned the skills should not care which order it is asked for.

Monolithic VLA policies care a great deal. Fine-tuned on the original long-horizon demonstrations of six LIBERO-Long++ tasks, Pi0.5 solves 83% of them. Ask for the same tasks with the skills in a different order and it solves none of them.

Method Original order Reordered Change
Pi0.5
83%
0%
−83 pts
LiLo-VLA (ours)
78%
85%
+7 pts
Suite 1 (LIBERO-Long++), success rate over 10 trials per configuration. The tasks, the objects and the trial count are identical across the two columns; only the order in which the skills are requested changes. Pi0.5 was LoRA fine-tuned on the original long-horizon demonstrations and prompted zero-shot with the reordered instruction. The evaluation is unseeded, so LiLo-VLA's +7 points is within run-to-run variance and should be read as unchanged, not as a gain.

The failure has a specific cause rather than a mysterious one: Pi0.5 frequently ignores the re-ordered language instruction and replays the order it was trained on. It has memorised a trajectory, not grounded a command.

LiLo-VLA is never trained on a task-level ordering at all — only on individual atomic skills, each starting from a canonical approach pose that a motion planner is responsible for reaching. No ordering is more in-distribution than another, so the reordered sequences of Suite 1 score no worse than the original ones. Everything below follows from that one decision.

Linking object-centric policies

Execution is split in two. A Reaching Module moves the end-effector from wherever the last skill ended to a per-object approach pose, using a collision-aware motion planner over the scene point cloud (MPLib). An Interaction Module then performs the contact-rich part with an object-centric VLA policy. A geometric verifier checks each skill’s effect, and a closed-loop wrapper decides what to do when it fails.

LiLo-VLA architecture, built up one stage at a time. The Reaching Module plans a collision-free path to an approach pose above the target object, with initial-state perturbation applied during data generation. The Interaction Module runs an object-centric VLA on a distractor-masked wrist view. The two alternate along a skill chain of transport, pick, transport, place, transport, open, and a failed skill falls back to the Reaching Module.

Architecture. The Reaching Module plans a collision-free path to an approach pose defined in the reference object’s frame; the pose is perturbed during demonstration generation so the downstream policy tolerates planner and perception error. The Interaction Module then executes the atomic skill from a wrist view with non-target objects masked out. A failed skill falls back to the Reaching Module for a reset rather than a blind retry.

Reaching: transport, planned not learned

The approach pose is defined in the reference object’s frame — a fixed face-down orientation plus a skill-specific offset, such as a vertical clearance for a pick — so it is deterministic and reusable across orderings. MPLib plans a collision-free path to it from the environment point cloud and the robot’s kinematic chain, with nothing learned.

Planner convergence and pose estimation are never exact, so demonstrations are not generated from the canonical pose. Each episode starts from a perturbed pose sampled around it, and the policy absorbs the residual error. Aggregated over the 27 unique skills in both suites, the perturbation-trained policy shows only a marginal drop under initial-pose noise, while a policy trained on canonical trajectory endpoints degrades sharply.

Interaction: one object, one wrist view

The interaction policy sees the wrist camera only. Fixed third-person views are deliberately excluded: as the global layout and the end-effector position change from skill to skill, a static camera sees inconsistent features — the observation-space shift that long-horizon execution suffers from — while the wrist view keeps the target in frame and most distractors out of it.

Non-target objects are then masked with black rectangles derived from their segmentation masks. Because that masking is itself a large visual shift, training applies random erasing to background pixels only, so the masked observation at deployment stays inside the distribution the policy learned.

Recovery, not retry

Each skill ends with a geometric verifier: LIBERO’s own contact, position and height conditions for a place, displacement or rotation for an articulation, and an added predicate for a pick, which succeeds when the target rises by at least 3 cm.

When a skill fails, what happens next depends on whether the gripper is holding something. A failed pick or articulation is retried locally — the object pose is re-estimated and the Reaching Module returns the end-effector to the approach pose first, so the retry starts from a state the policy trained on. A failed place means the held object’s location is uncertain, so execution backtracks to the most recent pick and re-acquires it. Every skill position keeps its own retry counter that no other skill’s success resets, with a budget of K = 10, which bounds the whole episode at O(NK) failures.

Setup, backbones and what is assumed

A benchmark of 21 configurations

The LIBERO-Long tasks this builds on stop at three or four skills. The benchmark here keeps that regime and adds one an order of magnitude longer, and every scenario ships with permuted skill orderings — those permutations are what no policy is trained on.

Left, Suite 1 LIBERO-Long++: two renders of the same LIBERO scene, one with a clean background and one with extra mugs and cans boxed in red, beside a four-chip skill chain shown in its original order and in a permuted order. Right, Suite 2 Ultra-Long: a kitchen scene with baskets, cans and a cabinet, beside three rows of coloured skill chips labelled Original Sequence, Variant 1 and Variant 2 over a per-object colour key.
The two suites. Suite 1 — LIBERO-Long++ takes 6 LIBERO-Long tasks that admit reordering and adds randomized distractors (boxed in red), so the policy has to ignore objects that belong to other skills. Suite 2 — Ultra-Long adds 3 tasks of 9, 10 and 16 skills, where object and receptacle density pushes a fixed-base tabletop arm to its kinematic reachability limits. Both suites include variant configurations with permuted skill orders.
Suite Scenarios Orderings each Configurations Skills per task
LIBERO-Long++
6
2
12
3–4
Ultra-Long
3
3
9
9–16
Total
9
—
21
—
LIBERO-Long++ tests visual robustness to clutter; Ultra-Long tests temporal scalability. Every configuration is a skill ordering no policy is trained on, so all 21 are executed zero-shot.

The three Ultra-Long tasks

  • Kitchen Organization — 9 skills. Dense receptacles: three baskets and a cabinet in one workspace.
  • Cooking Preparation — 10 skills. Geometrically awkward objects, the Moka Pot above all, which is where picks go wrong.
  • Living Room Organization — 16 skills. The longest chain in the benchmark, and the one where reachability, not perception, is the binding constraint.

Evaluation protocol

  • 10 trials per configuration.
  • SR counts an episode only if every skill in the sequence executes in the required order.
  • AP (average progress) counts the uninterrupted prefix of correct skills and stops at the first out-of-order or failed one.
  • Overall macro-averages the 9 scenarios, not the 21 configurations.

Results in simulation

Suite 1 · LIBERO-Long++ Suite 2 · Ultra-Long
Method Original Reordered Avg. Original Reorder 1 Reorder 2 Avg. Overall
Baseline · Pi0.5
83%
AP 93%
0%
AP 0%
42%
AP 46%
0%
AP 1%
0%
AP 0%
0%
AP 0%
0%
AP 0.3%
28%
AP 31%
Baseline · OpenVLA-OFT
7%
AP 11%
0%
AP 0%
3%
AP 5%
0%
AP 0%
0%
AP 0%
0%
AP 0%
0%
AP 0%
2%
AP 4%
Ablation · w/o Reaching
0%
AP 0%
0%
AP 0%
0%
AP 0%
0%
AP 0%
0%
AP 0%
0%
AP 0%
0%
AP 0%
0%
AP 0%
Ablation · w/o Masking
67%
AP 80%
77%
AP 87%
72%
AP 83%
0%
AP 16%
0%
AP 32%
0%
AP 10%
0%
AP 20%
48%
AP 64%
Ablation · w/o Recovery
2%
AP 25%
23%
AP 59%
13%
AP 42%
0%
AP 16%
0%
AP 16%
0%
AP 21%
0%
AP 18%
8%
AP 33%
LiLo-VLA (ours)
78%
AP 88%
85%
AP 89%
82%
AP 89%
53%
AP 79%
37%
AP 84%
43%
AP 74%
44%
AP 79%
69%
AP 86%
Success rate (SR) with average progress (AP) underneath, 10 trials per configuration. Overall macro-averages the 9 scenarios, not the 21 configurations. Suite 1 tests robustness to visual clutter, Suite 2 scalability to 16 sequential skills. Bold marks the best SR in a column.

Two things in that table are worth reading slowly.

Suite 2 is not solved. Both baselines score 0% on every Ultra-Long configuration, and LiLo-VLA reaches 44%. The gap is attributable to coupling: as layouts evolve across stages, a policy that owns transport as well as interaction needs intractable amounts of transition data, whereas delegating transport to a planner leaves the interaction policy invariant to those shifts. But 44% means most 9-to-16-skill episodes still end short of the goal.

Read AP next to SR. Suite 2 averages 44% SR at 79% AP — a typical trial completes about four fifths of its sequence. At these horizons SR collapses each trial to a single bit and discards most of the signal; AP is the more informative, and the more stable, number.

What each module buys

Each ablation removes one component and leaves the rest of the system intact.

Knowing where the object is does not help. The w/o Reaching ablation hands the whole episode to the Interaction Module and is granted the same ground-truth object poses the full method uses. It scores 0% on all 21 configurations. A VLA trained on atomic-skill demonstrations does not learn long-horizon transport from pose knowledge alone, so the motion planner is a structural prerequisite rather than an accelerator — and the oracle poses LiLo-VLA is given in simulation are not what the results rest on.

Masking matters on top of viewpoint. Removing random erasing costs 21 points overall, 69% to 48%, and takes Suite 2 to 0%. Even a wrist camera admits enough of the surrounding scene that object-centricity has to be enforced explicitly.

Recovery is not just more tries. A naive-retry control keeps the same budget of K = 10 but re-runs a failed skill from wherever the arm ended up, without re-estimating the object pose, resetting to the approach pose, or backtracking to the last pick:

Retry policy Success rate Average progress
Naive retry (same K = 10)
23%
52%
LiLo-VLA closed-loop recovery
53%
79%
Suite 2 Original, pooled over the three Ultra-Long scenarios, 30 trials. Fisher exact test on SR, p = 0.033. The gain comes from the loop restoring a known state before each attempt, not from the extra attempts.

Across the three Ultra-Long scenarios recovery triggers on 26.8% of executed skills and resolves 79.2% of them, and the yield is bounded by the atomic policy underneath it: Cooking Preparation triggers most often, on 37.5% of skills, and resolves only 46.7%, because some picks on geometrically complex objects never succeed.

Why wrist-only. On BOSS-C1, a benchmark that stress-tests policies against observation-space shift:

Observation Relative drop under observation-space shift
Wrist only (ours)
15.9%
Wrist + 3rd person
18.8%
3rd person only
29.8%
Wrist-only also reaches the highest clean success rate of the three configurations, 0.88. Only the relative drops and that clean rate are published, so only those are shown.

Real-robot deployment

A Franka Emika Panda with a Robotiq 2F-85 gripper and the dual-camera setup of DROID: a static ZED 2 gives the Reaching Module global context, and a wrist-mounted ZED Mini feeds the Interaction Module. Two things differ from simulation, and both are the point of this section — the backbone is Pi0.5 rather than OpenVLA-OFT, so the framework does not bind to one architecture, and the training data is teleoperated atomic skills only, with no long-horizon sequence data at all.

Scene Standard Change Sequence Diverse Layout
Scene A · 4 skills
100%
AP 100%
80%
AP 80%
100%
AP 100%
Scene B · 4 skills
100%
AP 100%
60%
AP 90%
80%
AP 95%
Scene C · 8 skills
100%
AP 100%
60%
AP 85%
—
5 trials per configuration, 85% success rate averaged over the 8 configurations. Standard runs the canonical configuration; Diverse Layout drastically alters the workspace and introduces unseen distractors; Change Sequence permutes the skill order and is executed zero-shot. Diverse Layout is not run on Scene C. Perception is YOLOE for detection and segmentation and FoundationPose for 6D pose; a human operator judges each skill.

Every rollout below is real hardware. Clips play as they scroll into view; each carries its task number and protocol in the frame, with the measured success rate overlaid.

Scene A · 4 skills

100% SR
80% SR
100% SR

Scene B · 4 skills

100% SR
60% SR
80% SR

Scene C · 8 skills

100% SR
60% SR

Run your own policy on it

The benchmark ships separately from the method. Evaluating your own policy on the 21 configurations needs neither the LiLo-VLA framework nor a GPU-sized dependency tree.

git clone https://github.com/YY-GX/LiLo-VLA.git && cd LiLo-VLA
pip install -e ".[libero]" # 9 packages, not 39
import lilo_vla.benchmark # registers the suites AND patches LIBERO
from libero.libero.benchmark import get_benchmark
bench = get_benchmark("ultra_long")() # or "libero_long_plus_plus",
# "ultra_long_all", "libero_long_plus_plus_all"
for i in range(bench.get_num_tasks()):
task = bench.get_task(i)
env, obs = bench.make_env(i) # the published evaluation protocol
# ... roll out YOUR policy ...
success = env.env._check_success()
env.close()

make_env() builds the environment exactly as every number on this page was measured, so results are comparable without matching a protocol by hand. Importing lilo_vla also applies the LIBERO compatibility patches the benchmark depends on — three of them change success criteria, so an unpatched LIBERO does not crash, it silently scores the same rollouts differently. The checkpoint and the atomic-skill dataset are released alongside it.

Limitations

A natural extension integrates an LLM/VLM planner that produces the skill sequence, and evaluates planning and execution end to end.

Takeaway

  • Problem. A monolithic VLA memorises the ordering it was trained on: Pi0.5 goes from 83% to 0% when the same six tasks are reordered, and to 0% everywhere once a task runs past a handful of skills.
  • Method. Delegate transport to a collision-aware motion planner, confine the learned policy to one masked, object-centric wrist view, and close the loop with a verifier that resets to a known state before retrying.
  • Benchmark. 21 configurations over 9 scenarios, including three Ultra-Long tasks of 9, 10 and 16 skills, all with permuted orderings that nothing is trained on.
  • Result. 69% SR and 86% AP in simulation against 28% for Pi0.5 and 2% for OpenVLA-OFT, and 85% across 8 real-robot configurations with a different backbone and atomic-skill data only.

BibTeX

@article{yang2026lilo,
title={LiLo-VLA: Compositional Long-Horizon Manipulation via Linked Object-Centric Policies},
author={Yang, Yue and Cheng, Shuo and Fang, Yu and Bharadhwaj, Homanga and Ding, Mingyu and Bertasius, Gedas and Szafir, Daniel},
journal={arXiv preprint arXiv:2602.21531},
year={2026}
}