FirsthandTech
arXiv — cs.AI preprintsInternational9 October 2026

Deconstructing Off-Policy Ratios: Entropy-Normalized Trust Regions for Asynchronous Reinforcement Learning

This is an official announcement record

Firsthand records what arXiv — cs.AI preprints announced and links to the original. The wording below is theirs, not ours.

arXiv:2607.22186v5 Announce Type: replace Abstract: Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, off-policy data destabilizes optimization and can cause policy collapse. Existing methods gate tokens by ratio magnitude alone, applying one threshold at every position. We show that the ratio's natural scale is set by token entropy, so deviations from mid-trajectory weight updates stay within this scale and carry genuine exploration. We further identify an overlooke
— arXiv — cs.AI preprints

More from arXiv — cs.AI preprints

This content is for informational purposes only and is not professional advice. Specifications, prices, plan tiers, and features change frequently and may differ from what is shown here; verify current details on the manufacturer's or company's official page before purchasing. Ratings are based on analysis of published documentation, not independent lab testing.