arXiv preprint · 2026

Same Scene, Different Task Skill Alignment for Compositional Generalization in VLAs

Taegeun Yang, Youngju Na, Yoonki Cho, Sung-Eui Yoon

KAIST

Demonstrations cover pick the yellow cube and place it on the yellow plate, and pick the green cube and place it on the green plate. For the undemonstrated instruction pick the yellow cube and place it on the green plate, CRAFT places the yellow cube on the green plate, whereas standard fine-tuning places it on the yellow plate, marked as a vision shortcut. A bar chart compares undemonstrated-combination success of the best standard fine-tuning and CRAFT across three VLAs and two benchmarks.
Fine-tuned VLAs struggle to execute new combinations of demonstrated skills.CRAFT uses demonstrated skill executions to train the policy to follow instructions for new combinations,improving success on undemonstrated combinations across two benchmarks and three VLA models over standard fine-tuning.

Abstract

Vision-language-action (VLA) models often struggle to generalize to skill combinations absent from their fine-tuning demonstrations, even when every constituent skill has been demonstrated. We focus on a vision shortcut as one failure mode: during fine-tuning, visual observations can serve as a proxy for the instruction, so a policy may execute a demonstrated combination associated with similar observations rather than the instructed combination. This motivates training with counterfactual pairs formed by holding a demonstration observation fixed while changing the instruction to specify an undemonstrated combination. These pairs, however, lack corresponding demonstrated action targets. Crucially, the currently required skill has already been demonstrated, but actions from those executions cannot serve as direct targets because the same skill can require different actions across observations. We propose CRAFT, which transfers supervision from demonstrated executions of the required skill to counterfactual pairs using skill representations that can be reused across executions of the same skill. Across three VLA models and two simulation benchmarks, CRAFT improves success on undemonstrated combinations while maintaining high success on demonstrated ones; it also improves compositional generalization on a real robot.

Interactive

Same scene, different task

Demonstrations cover every constituent skill but only same-color combinations. Select a combination to compare FT (Full) and CRAFT; cells show CRAFT’s success rate (%).

CRAFT success for
Undemonstrated

FT (Full)
Video coming soon
CRAFT (ours)
Video coming soon

Motivation

Fine-tuned VLAs struggle with new skill combinations

We define a skill as an operation (e.g., pick, place) paired with the entity specified for that operation (e.g., red cube, blue plate), such as “picking a red cube” or “placing on a blue plate.” Fine-tuned VLAs often struggle to generalize to skill combinations absent from their fine-tuning demonstrations, even when every constituent skill has been demonstrated.

The six pick-and-place tasks of LIBERO-Goal define the demonstrated combinations (✓). Of the remaining nine object–target pairs, we evaluate five undemonstrated combinations (○), such as “put the wine bottle on the stove” highlighted.

Object–target combinations in LIBERO-Goal

cabinet toprackplatestovebowl
wine bottle✓1✓4executed○○instructed
bowl✓2✓3✓6
cream cheese○○○✓5

✓ demonstrated (task index) · ○ undemonstrated, evaluatedblank: undemonstrated, not evaluated

Across four publicly released LIBERO-fine-tuned VLAs, mean success ranges from 88.8 to 97.4% on demonstrated combinations, compared with 2.0 to 50.9% on undemonstrated combinations.

Their failure behavior is consistent with a vision shortcut: visual observations can serve as a proxy for the instruction, so a policy may execute a demonstrated combination associated with similar observations rather than the instructed one.

LIBERO-Goal success (%), mean ± std over three seeds

ModelDemonstrated
6 tasks
Undemonstrated
5 combinations
π094.6±0.82.9±0.4
π0.597.4±0.650.9±2.4
π0-FAST88.8±1.019.5±0.7
GR00T N1.796.1±0.22.0±1.0
Four frames of a LIBERO-Goal rollout for the instruction put the wine bottle on the stove: the policy grasps the wine bottle and places it on the rack.
Instructed “put the wine bottle on the stove” ○ undemonstrated Executed “put the wine bottle on the rack” ✓ demonstrated · task 4
Two rollouts with an empty instruction: in trial 1 the policy turns on the stove; in trial 2 it puts the bowl on the plate.
Instructed (empty instruction) Executed Trial 1: “turn on the stove” ✓ task 8
Trial 2: “put the bowl on the plate” ✓ task 3

Rollouts of the fine-tuned π0.5 policy; task indices follow the ten original LIBERO-Goal tasks.

Method

CRAFT: Counterfactual Skill-Representation Alignment for Fine-Tuning

To mitigate the vision shortcut, we form counterfactual pairs during fine-tuning by keeping a demonstration observation fixed while changing the instruction to specify an undemonstrated combination. These pairs lack corresponding demonstrated action targets, and actions from demonstrated executions of the required skill cannot serve as direct targets because the same skill can require different actions across observations. CRAFT transfers supervision from demonstrated executions of the required skill to counterfactual pairs using skill representations that can be reused across executions of the same skill.

Overview of CRAFT. (1) Skill and state representation learning: a contrastive objective over flow-velocity predictions encourages skill representations to be reusable across executions of the same skill. (2) Counterfactual skill representation alignment: the skill representation of a counterfactual pair, with the same scene and a different skill, is combined with the state representation of a demonstration of the required skill and supervised with its target velocity.
Overview of CRAFT. Yellow instruction highlights mark the currently required skill.(1) We encourage reuse of skill representations across executions by contrasting prediction errors.(2) We combine the counterfactual skill representation with a state representation from a demonstration of the required skill,whose target velocity supervises the prediction.

01

Skill and state representations

Skill queries attend to both image and text tokens to produce a skill representation that reflects the currently required skill; state queries attend only to image tokens to produce a state representation. Together, they condition action prediction.

02 · ℒskill

Reusable skill representations

Holding the state representation fixed, we replace only the skill representation with one from another execution. A contrastive objective over the resulting prediction errors encourages skill representations to be reusable across same-skill executions and to preserve distinctions between skills.

03 · ℒcf

Counterfactual supervision

We use the counterfactual pair’s skill representation to condition action prediction for a demonstrated execution of the currently required skill. We supervise this prediction with that execution’s action target. Thus, ℒcf trains the counterfactual pair’s skill representation to reflect the skill required under the counterfactual instruction in action prediction.

ℒtotal = ℒfm + β ℒskill + γ ℒcfoptimized jointly in a single training stage

Benchmarks

Two compositional benchmarks

We introduce two compositional benchmarks in which the fine-tuning demonstrations cover every constituent skill but only a subset of the possible combinations. In both, the same-color combinations are demonstrated, and we evaluate the remaining 12 in Pick-Place and 6 in Pick-Place-Press. We collect 50 demonstrations per demonstrated combination and will release both benchmarks and their demonstration datasets.

Pick-Place has 4 demonstrated and 12 undemonstrated combinations of cube and plate colors; Pick-Place-Press has 2 demonstrated and 6 undemonstrated combinations of cube, plate, and button colors.
Demonstrations cover all constituent skills but only a subset of combinations (checked cells); empty cells denote undemonstrated combinations.Rows and columns indicate cube and plate colors; the two Pick-Place-Press matrices correspond to red and blue buttons.

Results

Generalizing to undemonstrated combinations

Standard fine-tuning often performs well on demonstrated combinations but achieves substantially lower success on undemonstrated ones, and changing the adaptation method alone does not close this gap. Across all VLA models and benchmarks, CRAFT achieves the highest success on undemonstrated combinations while maintaining high performance on demonstrated ones.

Success rates (%) on demonstrated (Demo.) and undemonstrated (Undem.) combinations; mean ± std over three evaluation seeds.

Methodπ0.5π0GR00T N1.7
Demo.Undem.Demo.Undem.Demo.Undem.
Pick-Place
FT (Full)100.0±0.08.4±0.3100.0±0.02.8±0.298.8±0.31.1±0.1
FT (LoRA)100.0±0.046.3±0.997.5±0.01.0±0.297.8±0.62.6±0.3
FT (Frozen)99.0±0.021.1±0.248.2±1.43.0±0.592.3±1.32.2±0.1
CAG-VA96.3±1.566.5±0.550.5±1.36.8±0.688.7±1.34.9±0.5
Skill Composition100.0±0.025.7±0.499.5±0.012.7±0.499.7±0.613.1±0.7
CRAFT (ours)99.0±0.984.2±0.394.0±0.964.3±0.599.0±0.572.1±0.9
Pick-Place-Press
FT (Full)100.0±0.00.7±0.099.7±0.60.3±0.097.7±2.31.1±0.2
FT (LoRA)100.0±0.015.4±0.899.3±0.60.7±0.099.3±1.21.0±0.0
FT (Frozen)98.0±1.00.9±0.282.0±3.51.5±0.497.7±0.61.1±0.2
CAG-VA92.7±3.521.8±1.969.0±2.61.1±0.288.7±2.52.1±0.2
Skill Composition99.7±0.65.8±0.299.3±0.60.3±0.099.3±0.62.0±0.0
CRAFT (ours)99.7±0.662.6±3.995.3±0.640.7±1.599.3±0.631.9±1.1
  • FT (Full / LoRA / Frozen) denote VLM adaptation; the action expert is fully trained in all cases. CRAFT uses the Full setting, and all methods use the same robot demonstrations.
  • CAG-VA (Fang et al., 2026) mitigates vision shortcuts; we report the selected configuration for each VLA model and benchmark.
  • Skill Composition fine-tunes the VLA on constituent skill segments and sequentially executes the required skills at deployment.

Analysis

Behavior and representation analysis

It responds to instruction changes during execution

When an instruction change leaves the currently required skill unchanged, the policy continues executing that skill. When the required skill changes, the policy begins executing the newly required skill, from a state reached while executing a different skill.

Six frames of a rollout where the instruction changes during execution: the robot keeps approaching the red cube when only the placement target changes, redirects to the blue cube when the picking target changes, and changes its destination during placing to the yellow plate.
Yellow highlights identify the entity assigned to the current operation.The robot approaches the red cube and maintains this behavior when only the future placement target changes (1→2→3).It redirects its approach when the picking target changes (3→4),then changes its destination during placement (5→6), placing the blue cube on the yellow plate.
Instruction-change video

Instruction

Its action predictions reflect the currently required skill

Holding the visual observation fixed, the predicted trajectories under CRAFT separate according to the instructed cube or plate, whereas predictions from FT (Full) largely overlap across instruction changes.

PickingFixed observation: picking the blue cube“pick the red/blue/green/yellow cube”

FT (Full)
FT (Full): the predicted action chunks for all four instructed cubes overlap, heading to the blue cube.
CRAFT (ours)
CRAFT: the predicted action chunks separate toward the red, blue, green, and yellow cubes.

PlacingFixed observation: placing on the red plate“place it on the red/blue/green/yellow plate”

FT (Full)
FT (Full): the predicted action chunks for all four instructed plates overlap near the held red cube.
CRAFT (ours)
CRAFT: the predicted action chunks fan out toward the red, blue, green, and yellow plates.

Trajectory colors indicate the instructed entity,with predicted action chunks shown for 30 random seeds (π0.5, Pick-Place).Insets enlarge the marked regions.

Skill representations are organized by skill

Without ℒskill, representations from different entities are more mixed within the pick and place groups. With ℒskill, they form entity-specific groups within each operation.

t-SNE projections of skill representations learned without and with the skill objective. With it, points form entity-specific groups within the pick and place groups.
Skill representations learned with and without ℒskill on Pick-Place,colored by the entity of the currently required skill.

…and reusable across executions

Replacing the skill representation with one from another execution of the same skill changes the prediction by only 4.0 ± 3.1% on average across recipients. In contrast, substitutions from different skills produce substantially larger changes.

Heatmaps of relative action change when substituting skill representations between recipient and donor targets for picking and placing. Diagonal same-skill substitutions change the prediction by 3 to 5 percent; off-diagonal substitutions by 50 to 101 percent.
Mean relative change (%) in the predicted action chunk;outlined diagonal cells use same-skill donors (π0.5, Pick-Place).

Counterfactual supervision benefits from both terms

ℒcf covers both instruction changes that alter the currently required skill (ℒchange) and those that leave it unchanged (ℒpreserve). Success on undemonstrated combinations remains low when ℒskill is added and improves with ℒchange. Combining ℒchange and ℒpreserve yields substantially higher success on undemonstrated combinations than using either counterfactual term separately. The combined objective also maintains high success on demonstrated combinations.

Pick-Place success (%), mean ± std over three seeds. All variants include ℒfm.

Auxiliary objectivesπ0.5GR00T N1.7
ℒskillℒpreserveℒchangeDemo.Undem.Demo.Undem.
88.8±0.61.9±0.499.3±0.33.8±0.1
✓99.2±0.60.3±0.497.2±1.06.1±0.5
✓✓99.8±0.30.4±0.297.3±0.65.3±0.3
✓✓97.7±0.314.7±0.894.5±1.312.1±0.4
CRAFT (all)99.0±0.984.2±0.399.0±0.572.1±0.9

Real robot

Compositional generalization on a real robot

A Piper 6-DoF robot performs the Pick-Place tasks, with 20 demonstrations for each of the four same-color combinations. Over five trials per combination on π0.5, CRAFT achieves higher success than FT (Full) on undemonstrated combinations while maintaining high success on demonstrated combinations.

Real-robot success (successful / total trials)

MethodDemo.Undem.
FT (Full)19/209/60
CRAFT (ours)19/2043/60
Real-robot rollouts for pick the blue cube and place it on the green plate. CRAFT succeeds; FT Full places the blue cube on the blue plate and fails.
Undemonstrated “pick the blue cube and place it on the green plate.”Both CRAFT and FT (Full) pick the instructed blue cube, but CRAFT places it on the instructed green plate,whereas FT (Full) places it on the blue plate, corresponding to the demonstrated blue-to-blue combination.
blue cube→green plate
Video
red cube→yellow plate
Video
yellow cube→blue plate
Video
yellow cube→red plate
Video
blue cube→red plate
Video

Additional real-robot rollouts of CRAFT (π0.5) on undemonstrated Pick-Place combinations; all videos play at real-time speed.

BibTeX

@article{yang2026craft,
  title   = {Same Scene, Different Task: Skill Alignment for Compositional Generalization in {VLAs}},
  author  = {Yang, Taegeun and Na, Youngju and Cho, Yoonki and Yoon, Sung-Eui},
  journal = {arXiv preprint arXiv:2610.00524},
  year    = {2026}
}