arXiv — cs.AI preprintsInternational5 October 2026
Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript
This is an official announcement record
Firsthand records what arXiv — cs.AI preprints announced and links to the original. The wording below is theirs, not ours.
arXiv:2610.02638v1 Announce Type: new Abstract: Full-duplex voice agents make many small, closed decisions, which current systems answer by slow autoregressive decoding. We propose DuplexJev, which feeds ASR-encoder hidden states through a small connector into a frozen LLM and reads each question as a single-token distribution over its options. Nothing is decoded, and an 8-GPU node answers 80 decisions about eight utterances in about 0.1 s. With a last-layer connector, spoken QA stays close to reading the transcript (90% vs. 91%). DuplexJev also hears the speaker: gender and emotion accuracy b
Read the official announcement
Opens arxiv.org
More from arXiv — cs.AI preprints
- MintFlow: Minimal Trajectory Intervention for Constrained Flow Matching5 October 2026
- Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses5 October 2026
- The AI Risk Observatory: What Can We Learn from AI Disclosures in Annual Reports About Societal Resilience?5 October 2026
- Keep It CALM: Analyzing the Limits of Global Unsafety in Text-to-Image Generation5 October 2026
- Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents5 October 2026
This content is for informational purposes only and is not professional advice. Specifications, prices, plan tiers, and features change frequently and may differ from what is shown here; verify current details on the manufacturer's or company's official page before purchasing. Ratings are based on analysis of published documentation, not independent lab testing.