From Responses to Behaviors
Knowledge unlearning focuses on what the model says, whereas trajectory unlearning concerns what the agent actually does. This difference calls for different optimization objectives.
Motivation
Overview
LLM unlearning has focused on suppressing facts an agent knows. As LLMs act as autonomous agents, they must also stop reproducing undesired behaviors through their action trajectories.
Existing large language model (LLM) unlearning has focused primarily on removing specific knowledge, such as harmful facts, private data, or copyrighted content. However, as LLMs are increasingly deployed as autonomous agents, a fundamental yet overlooked problem emerges: beyond suppressing what an agent knows, an agent should not reproduce undesired behaviors through its action trajectories. In this work, we introduce trajectory-level unlearning, a new problem formulation that targets the removal of specific action trajectories in long-horizon agentic tasks, rather than factual knowledge.
We identify two fundamental challenges that distinguish trajectory unlearning from knowledge unlearning: (1) our unlearning target is what the agent does, not what it says; and (2) trajectories are sequentially dependent action sequences that cannot be decomposed into isolated prompt-response pairs without losing inter-step structure. To address these challenges, we propose Group-injected Relative Policy Optimization (GiRPO), which injects forget trajectories into the policy rollout group with penalized rewards and isolates the normalization statistics, yielding a stable and bounded unlearning signal that does not corrupt gradient updates for normal task trajectories.
We construct trajectory unlearning benchmarks from two application scenarios, household tasks (ALFWorld) and online shopping (WebShop), and design three complementary metrics for evaluating forgetting quality and model utility. Experiments on ALFWorld and WebShop demonstrate that GiRPO effectively unlearns target trajectories while preserving task success rates, outperforming existing knowledge-unlearning baselines on both forgetting quality and task utility.
Knowledge unlearning focuses on what the model says, whereas trajectory unlearning concerns what the agent actually does. This difference calls for different optimization objectives.
A trajectory consists of a sequence of actions that depend on the interaction history and on previous actions. In contrast, text-based unlearning methods typically treat trajectories as a collection of independent prompt-response pairs, thereby overlooking the dependencies between steps.
Method
Forget trajectories ride alongside real rollouts inside the same policy-gradient update, but with a penalized reward and normalization statistics isolated from the normal-task group.
Results
On both ALFWorld and WebShop, GiRPO drives down exact-match and LLM-as-judge similarity to the forgotten trajectory while matching or exceeding the base model's success rate on target and untargeted tasks.
Averaged over the four ALFWorld unlearning settings (Clean, Heat, Cool, Mixed) reported below; WebShop results follow.
Analysis
In the forget trajectory dataset for ‘Clean’ target tasks, there is one forget trajectory for each of the 650 tasks. Among these, 595 forget trajectories successfully complete the corresponding task (complete) and 55 do not (incomplete). We analyze the effect of trajectory completeness on unlearning quality and task utility separately.
Incomplete trajectories correspond to behaviors the model already struggles to reproduce, whereas complete trajectories represent well-learned, high-reward behavioral patterns that are deeply embedded in the policy.
Unlearning incomplete trajectories improves task success rate from 0% to 40%, since suppressing a failed action sequence implicitly steers the model away from suboptimal behaviors and toward successful alternatives.