What Do Your Logits Know? (The Answer May Surprise You!)

Recent work has shown that probing model internals can reveal a wealth of information not apparent from the model generations. This poses the risk of unintentional or malicious information leakage, where model users are able to learn information that the model owner assumed was inaccessible. Using vision-language models as a testbed, we present the first systematic comparison of information retained at different “representational levels” as it is compressed from the rich information encoded in the residual stream through two natural bottlenecks: low-dimensional projections of the residual...

What Do Your Logits Know? (The Answer May Surprise You!)

Affected Roles

Time Horizon

What Changes

Recommended Action

Ready to dive deeper?

Discussion

Related Stories

Can Large Language Models Understand Context?

What is Codex?

AI agents (Grok vs. GPT-4o mini) compete in live crypto paper trading