FirsthandTech
arXiv — cs.AI preprintsInternational9 October 2026

Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception

This is an official announcement record

Firsthand records what arXiv — cs.AI preprints announced and links to the original. The wording below is theirs, not ours.

arXiv:2610.12445v1 Announce Type: cross Abstract: Recent incidents have highlighted the challenge of monitoring LLM agents and the danger of models deceiving people. We show that white-box deception detection via probes can be scaled up to frontier monitoring settings by collecting the largest deception dataset to date for training probes and introducing a novel probe architecture which can aggregate information across many layers and tokens. Our probes achieve 98.8% AUC in SHADE-Arena, surpassing an Opus 5.5 text-monitoring baseline, and show improved efficacy as the underlying model is scale
— arXiv — cs.AI preprints

More from arXiv — cs.AI preprints

This content is for informational purposes only and is not professional advice. Specifications, prices, plan tiers, and features change frequently and may differ from what is shown here; verify current details on the manufacturer's or company's official page before purchasing. Ratings are based on analysis of published documentation, not independent lab testing.