Stealing Reasoning Traces from Proprietary LLM APIs
We found a way to extract reasoning traces from frontier AI APIs, highlighting significant security risks and potential data leaks for engineers.
We have discovered a method to extract hidden reasoning from frontier models using vulnerabilities in the APIs of leading AI companies. This indicates that distilling reasoning traces may have been feasible for a long time, raising significant security concerns. Additionally, some models exhibit reasoning processes that could lead to sensitive data leaks.