研究发现,Anthropic、OpenAI 和 Google 等专有 LLM 的加密推理轨迹块可跨会话、用户和模型互换,攻击者将其注入同提供商防护较弱的模型,即可强制其以明文输出推理内容。
论文
·HuggingFace Daily Papers(社区热门论文)
窃取专有 LLM API 的推理轨迹:加密块可跨会话互换引发解密越狱
— Stealing Reasoning Traces from Proprietary LLM APIs
— Stealing Reasoning Traces from Proprietary LLM APIs