When Reasoning Helps Action: Monitoring and Steering Chain-of-Thought in Vision-Language-Action Policies

Sathwik Karnik*, Joseph JR. Lee*, Aryaman Gupta, Somil Bansal
Safe and Intelligent Autonomy Lab, Stanford University

Abstract

Reasoning-enabled VLA policies expose chain-of-thought (CoT) traces that appear to explain and guide their actions, creating a potential interface for runtime safety through reasoning monitoring and correction. In this work, we define and operationalize two evaluation axes for assessing when this interface can improve embodied behavior: correctability, which measures whether unreliable reasoning can be detected and improved during generation, and actionability, which measures whether reasoning corrections produce behaviorally meaningful changes in the intended direction. To enable correctability, we introduce Token-level Reward for Utility-Steered Chain-of-Thought (TRUST), an offline-trained value model that predicts eventual reasoning correctness from partial prefixes and uses these estimates to monitor and selectively steer reasoning generation in frozen VLA policies. On the Alpamayo 1.5 driving VLA, TRUST monitors correctness with 88.9% accuracy and improves reasoning correctness from 75.9% to 90.0%. On a baseline-defined challenging subset in AlpaSim, TRUST reduces collision rate by 30.4% and maximum trajectory error by 11.5% relative to the unsteered policy, outperforming a compute-matched Best-of-4 baseline. On the DeepThinkVLA manipulation VLA, TRUST improves the correctness of grasp-state claims from 69.3% to 90.2% and action-choice claims from 68.8% to 85.9%, yet closed-loop task performance on LIBERO-Plus remains largely unchanged. Empirical analysis reveals intent-consistent behavioral effects in Alpamayo 1.5 but limited effects in DeepThinkVLA, helping interpret these different task-level outcomes. Together, our results show that gains in reasoning correctness do not automatically imply gains in embodied performance, motivating evaluation of correctability and actionability when using CoT as a runtime safety interface.

Key Findings

  1. 1 In reasoning-based VLA policies, CoT provides an appealing interface for runtime safety.
  2. 2 With TRUST, CoT reasoning tokens can be steered reliably to be more correct.
  3. 3 Reasoning correctability does not always imply actionability.
  4. 4 Reasoning correctability and actionability should be measured separately.

Actionability of Reasoning Corrections

Alpamayo 1.5 driving example: under unsteered reasoning the vehicle proceeds through the turn; under corrected reasoning it decelerates to a stop at the yellow left-turn arrow. Unsteered Reasoning: Accelerate to proceed through the intersection since the straight traffic light is green [min ADE = 8.83 m]. Correct Reasoning: Decelerate to a stop at the stop line because the left-turn arrow is yellow [min ADE = 2.02 m].
(a) Alpamayo 1.5 continues through the turn under its own reasoning but decelerates under the corrected reasoning trace.
DeepThinkVLA manipulation example: correcting a hallucinated grasp claim leaves the gripper-to-object distance and end-effector path nearly unchanged. Unsteered Reasoning: The robot arm has moved and is now grasping the cream cheese (center-left, in gripper)...Place the cream cheese in the bowl. Correct Reasoning: The cream cheese (blue box, center-left) is now being approached....Pick up the cream cheese and move it above the bowl.
(b) Correcting DeepThinkVLA's mistaken grasp-state claim and associated action intent produces little change in its action chunk.

We substitute scene-consistent reasoning for erroneous reasoning traces and examine whether actions change in the intended direction. These two examples illustrate contrasting behavioral responses to corrected reasoning.