Measuring LLM Unlearning Depth with Activation Patching and UDS
The Unlearning Depth Score (UDS) quantifies LLM unlearning depth using activation patching.
The Unlearning Depth Score (UDS) is a mechanistic metric that quantifies how much target knowledge can be recovered through two-stage activation patching. This method provides engineers with a valuable tool for evaluating the depth of knowledge erasure in machine learning models, benchmarked across 150 unlearned models.
This synthesis was produced from its source by AI; there is no human editor or manual review step. How we work