Announcements tracker
What tech authorities announced
627 announcements from official sources, each one linking straight to the document that issued it. We record what was announced and who announced it. We do not rewrite it.
- arXiv — cs.AI preprintsInternational2 Oct 2026
Heavy-Tailed Memory Traces in Long-Horizon Language Agents
arXiv:2610.00010v1 Announce Type: new Abstract: Long-horizon language agents increasingly rely on external memory as a frozen world model, yet current memory systems are usually judged only by task success or token…
- arXiv — cs.AI preprintsInternational2 Oct 2026
When Do Causal World Models Help Modular LLM Agents
arXiv:2610.00012v1 Announce Type: new Abstract: LLM agents increasingly act through modular systems, such as order, payment, inventory, and shipment services, where actions in one module change which transitions are…
- arXiv — cs.AI preprintsInternational2 Oct 2026
From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent Execution
arXiv:2610.00015v1 Announce Type: new Abstract: Large-language-model agents can propose and execute actions, but proposal, authority, dispatch, verified external effect, and serving promotion are different claims. We…
- arXiv — cs.AI preprintsInternational2 Oct 2026
What Do Rationales Communicate? A Message-Intervention Study in Role-Specialized QA
arXiv:2610.00018v1 Announce Type: new Abstract: Role-specialized QA pipelines increasingly pass rationales from a reasoner to a verifier, but it is unclear what this message actually buys: better answers, stronger…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Measuring the Microtask Eligibility Gap: When Is an Off-the-Shelf SLM Enough for an Agent Harness?
arXiv:2610.00025v1 Announce Type: new Abstract: Agent harnesses increasingly want to run small language models (SLMs) on the microtasks around a frontier large language model (LLM) planner: auto-approving shell…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Characterizing a Configuration Where Inference-Time PRM-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMs
arXiv:2610.00047v1 Announce Type: new Abstract: Diversity collapse in parallel chain-of-thought has motivated inference-time interventions built on a natural design: when a process reward model (PRM) prunes a chain, its…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Gradient-Aligned Pair Selection for Personalized Preference Optimization
arXiv:2610.00061v1 Announce Type: new Abstract: Personalizing large language models (LLMs) requires aligning generation behavior with user-specific preferences rather than aggregate quality. While Direct Preference…
- arXiv — cs.AI preprintsInternational2 Oct 2026
K-Dense BYOK: An Open-Source AI Research Assistant That Runs Locally and Keeps a Hash-Chained Lab Notebook
arXiv:2610.00074v1 Announce Type: new Abstract: K-Dense BYOK (bring your own keys) is a free, open-source AI research assistant for scientists in any field that runs on the researcher's own computer. The researcher…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks
arXiv:2610.00084v1 Announce Type: new Abstract: Detailed profession-specific system prompts raise token use and estimated cost per response without a consistent accuracy gain. We evaluate Scientific Agents, an…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Comedic Fool's Gold: Reward Exploits and Countermeasures in Conversational Humor
arXiv:2610.00197v1 Announce Type: new Abstract: We investigate automated rewards for training language models in conversational humor, focusing on reward exploits and countermeasures. Two approaches aim to capture…
- arXiv — cs.AI preprintsInternational2 Oct 2026
EviGraph: Proof-Carrying Selective Recommendation over Temporal Public-Service Knowledge Graphs
arXiv:2610.00212v1 Announce Type: new Abstract: Public-service recommendations require evidence that matches the requested service, scope, and date. Yet treating every missing detail as decisive can withhold useful…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Build2SPARQL: A Large-Scale Text-to-SPARQL Benchmark Dataset for Building Knowledge Graph Querying
arXiv:2610.00224v1 Announce Type: new Abstract: Building automation systems are increasingly represented as semantic knowledge graphs (KGs) using ontologies such as Brick and ASHRAE 223P, creating a machine-readable…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Robust Is Salient: An Informed Adversary Moves the Optimal Signal onto the Salience Pole
arXiv:2610.00233v1 Announce Type: new Abstract: When an informed adversary shares the audience of a constrained signalling channel, the signal that best protects the truth is the signal that best describes it. On 108…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Conflicting Supervision Moves Commitment, Not Capability: A 12.29{\sigma} arrangement effect that is exactly zero under a convention-agnostic score
arXiv:2610.00234v1 Announce Type: new Abstract: "Train a model on the same problems written under two incompatible conventions, both correct, and ask what the ordering of that data writes into the parameters. The…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Knowing When to Yield: Grounded Arbitration of User Corrections in Text-Based Embodied Agents
arXiv:2610.00282v1 Announce Type: new Abstract: How should an embodied agent respond when a person's correction may be wrong? We formulate grounded correction arbitration as a choice among accepting, rejecting,…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Rules to Tools: Executable Checks for LLM Agents in Scientific Computing
arXiv:2610.00313v1 Announce Type: new Abstract: Scientific coding agents receive equations, boundary conditions, and output requirements in writing, then must assess the programs they revise. Rules to Tools (R2T)…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts
arXiv:2610.00314v1 Announce Type: new Abstract: Research agents explain planned experiments. We measure predictive credit with paired forecasts sharing an intervention, forecaster, and outcome while varying description,…
- arXiv — cs.AI preprintsInternational2 Oct 2026
ContractRL: Shielded Group-Relative Policy Optimization for Auditable Tool-Call Repair
arXiv:2610.00328v1 Announce Type: new Abstract: Structured tool calls often fail after only a small number of fields violate a schema or an execution contract. Regenerating the complete object enlarges the action…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Mathematical Transfer in LLMs Follows Reasoning Approach More Than Topic
arXiv:2610.00331v1 Announce Type: new Abstract: When selecting mathematical training data for LLMs, a natural organizing principle is topic: probability examples for probability targets. An alternative is reasoning…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Fault-Tolerant Budget Conservation in Distributed Multi-Agent Delegation
arXiv:2610.00349v1 Announce Type: new Abstract: Resource limits are becoming an authorization boundary for AI agents that delegate work across concurrent and failure-prone workers. Parent-child allocation constraints,…
- arXiv — cs.AI preprintsInternational2 Oct 2026
JusticeAxis: Benchmarking Legal Judgment between Rigid Rule Application and Ungrounded Discretion
arXiv:2610.00353v1 Announce Type: new Abstract: A sound judgment applies the law to established facts and weighs the circumstances in which they arose. However, existing methods swing between rigid statute matching and…
- arXiv — cs.AI preprintsInternational2 Oct 2026
What Should an Agent Remember? Disentangling Retention from Retrieval in Bounded-Memory Evaluation
arXiv:2610.00366v1 Announce Type: new Abstract: A persistent agent must decide both what to retain as information arrives and what to surface once a query appears, yet memory evaluations can confound these decisions by…
- arXiv — cs.AI preprintsInternational2 Oct 2026
When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents
arXiv:2610.00372v1 Announce Type: new Abstract: Large language model agents rely on external harnesses to pass information between the model and its environment and to recover from execution errors. Yet recovery is…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Benchmarking Prompt Optimization of Large Language Models With Chess
arXiv:2610.00416v1 Announce Type: new Abstract: Evaluating large language models becomes increasingly challenging as their capabilities advance: benchmarks can saturate, public test sets risk contamination, and…
- arXiv — cs.AI preprintsInternational2 Oct 2026
JevSpawn: Adaptive Agentic Inference through Compositional Action Spaces
arXiv:2610.00437v1 Announce Type: new Abstract: LLM agents generate intermediate reasoning and actions token by token, making extended interactions slow and computationally expensive. Jev-style models offer fast…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Frozen Scenes, Shifting Winners: Configuration Fragility in Text-to-3D Evaluation
arXiv:2610.00447v1 Announce Type: new Abstract: Can a text-to-3D leaderboard change when every generated scene stays fixed? We audit this question for rendered-image evaluation, where camera settings and caption wording…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Before Agents Decide: Epistemic Action in LLM-Based Systems
arXiv:2610.00511v1 Announce Type: new Abstract: Before a difficult decision, people often act simply to understand the situation better. We turn an object to see another side, place alternatives next to each other, or…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Ontology-Based Contextual AI Evaluations (OB-CAIE) Methodology
arXiv:2610.00529v1 Announce Type: new Abstract: The ontology-based contextual AI evaluation (OB-CAIE) methodology was developed to address a lack of scientific rigor that arises from unclear testing coverage, to balance…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers
arXiv:2610.00531v1 Announce Type: new Abstract: AI-generated content, often called AI slop, is increasingly common everywhere, particularly in academia. Slop in AI-generated scientific papers, however, has more complex…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams
arXiv:2610.00583v1 Announce Type: new Abstract: People are increasingly delegating tasks to AI agents, and those agents are increasingly encountering other people's agents over shared resources such as a codebase, a…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents
arXiv:2610.00609v1 Announce Type: new Abstract: Legal research is a core and time-consuming legal workflow. Lawyers must identify controlling authority, verify that it remains valid, reconcile statutes and cases, and…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Spatial Strategies, Not Actions: Vector-Quantized Geodesics as Tools for LLM-Driven Agents
arXiv:2610.00613v1 Announce Type: new Abstract: Large language model (LLM) based agents are often criticized for lacking spatial understanding and mainly exploiting statistical text patterns. We investigate their…
- arXiv — cs.AI preprintsInternational2 Oct 2026
CompMat-Bench: Benchmarking AI Agents for Computational Materials Science
arXiv:2610.00636v1 Announce Type: new Abstract: Evaluating AI agents on scientific research tasks is constrained by the time and resources required for the underlying experiments or calculations. In computational…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Incident-Arena: Getting agents to the last nine of reliability
arXiv:2610.00648v1 Announce Type: new Abstract: AI coding agents are ubiquitous in engineering workflows amongst industry and academia. Yet, despite their use in app coding, relatively less attention has been paid to…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard
arXiv:2610.00651v1 Announce Type: new Abstract: Agent evaluations are increasingly used to compare LLMs and inform deployment decisions, yet ranks can reflect not only the model but also the effects of the evaluation…
- arXiv — cs.AI preprintsInternational2 Oct 2026
When More Data Is Not Enough: The Context-Sufficiency Frontier in Generative AI Personalization
arXiv:2610.00654v1 Announce Type: new Abstract: Personalization has long relied on customer data to infer what an individual is likely to value. We call this customer evidence: the customer's historical behavior and…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Backdoor Containment via Expert Quarantine and Shutdown in LLMs
arXiv:2610.00663v1 Announce Type: new Abstract: Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four…
- arXiv — cs.AI preprintsInternational2 Oct 2026
A Simple Doxastic Deontic Logic for Norm-Guided Decision Making
arXiv:2610.00668v1 Announce Type: new Abstract: Making decisions despite conflicting norms and incomplete or unreliable information is a fundamental challenge for autonomous systems. We introduce a simple doxastic…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Ontology-Grounded, Reasoner-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI
arXiv:2610.00682v1 Announce Type: new Abstract: Large language models (LLMs) increasingly underpin scientific AI applications that reason over structured knowledge, from biomedical question answering to materials…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection
arXiv:2610.00685v1 Announce Type: new Abstract: With the rapid adoption of large language models (LLMs) and parameter-efficient fine-tuning (PEFT) methods, the risk of backdoor attacks has become more severe. Existing…
- arXiv — cs.AI preprintsInternational2 Oct 2026
R-GroundBench: A Diagnostic Benchmark for R-Group Groundingin Markush Molecular Editing
arXiv:2610.00700v1 Announce Type: new Abstract: Recent advances in AI for scientific discovery enable molecular understandingand design, yet reasoning over incomplete chemical representations remainsunclear.Markush…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Meta-Multi-Agent Reinforcement Learning for Fast Adaptation of Interactive Policies with Applications to Autonomous Driving
arXiv:2610.00705v1 Announce Type: new Abstract: This paper develops a meta-multi-agent reinforcement learning (meta-MARL) framework to enable fast adaptation of interactive policies in a multi-agent system (MAS).…
- arXiv — cs.AI preprintsInternational2 Oct 2026
ReLiveGym: Evaluating Long-Lived Agents over Weeks of Replayed Reality
arXiv:2610.00710v1 Announce Type: new Abstract: As large language model (LLM) agents become widely adopted, they are increasingly deployed for tasks that require persistent monitoring or recurring actions (e.g., market…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Robust Nash Alignment under Preference Uncertainty
arXiv:2610.00715v1 Announce Type: new Abstract: Preference-based alignment methods typically optimize against a single preference model, and can therefore be brittle when pairwise preferences are uncertain: noisy,…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Enterprise Representation Simplification (ERS): Reducing Representational Complexity for Enterprise AI
arXiv:2610.00791v1 Announce Type: new Abstract: Enterprise information is represented through artifacts shaped by applications, projects, technologies, organizational boundaries, and local requirements. These structures…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Sapien: A Stateful Policy Engine for Autonomous AI Agents
arXiv:2610.00797v1 Announce Type: new Abstract: Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Kepler: Auditable World Models for ARC-AGI-3
arXiv:2610.00834v1 Announce Type: new Abstract: ARC-AGI-3 evaluates agents in interactive environments whose rules and objectives must be inferred from observation. We present Kepler, an open-source harness that…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning
arXiv:2610.00849v1 Announce Type: new Abstract: Existing approaches to offline goal-conditioned reinforcement learning (GCRL) struggle with long-horizon tasks. Discounting shrinks value differences between distant…
- arXiv — cs.AI preprintsInternational2 Oct 2026
An Educator-Guided LLM Pedagogical Agent for Scaffolded Feedback in Conceptual Database Design
arXiv:2610.00870v1 Announce Type: new Abstract: We present an educator-guided LLM pedagogical agent for scaffolded feedback in conceptual database design. Integrated into an entity--relationship diagram (ERD) editor,…
- arXiv — cs.AI preprintsInternational2 Oct 2026
MemFit: Efficient Long-Term Agentic Memory
arXiv:2610.00872v1 Announce Type: new Abstract: Long-term memory systems for large language models (LLMs) have gained popularity for extending reasoning capabilities across applications. Current memory systems rely on…
- arXiv — cs.AI preprintsInternational2 Oct 2026
ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization
arXiv:2610.00906v1 Announce Type: new Abstract: Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback.…
- arXiv — cs.AI preprintsInternational2 Oct 2026
OR for AI That Does OR: Routing LLMs up the Escalator inside the OSCAR Framework
arXiv:2610.00912v1 Announce Type: new Abstract: Large language models can translate business descriptions into optimization models, but executable code may misrepresent constraints or objectives. A solver can then…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Finding the Right Fit: Model-Harness Interactions across Agent Tasks
arXiv:2610.00917v1 Announce Type: new Abstract: Choosing an agent system means choosing both a language model and the harness through which it acts. We ask whether a strong model, harness, or pairing stays strong when…
- arXiv — cs.AI preprintsInternational2 Oct 2026
ABDA-NL: A Natural-Language Scenario Explorer for Argument-Based Reasoning
arXiv:2610.00947v1 Announce Type: new Abstract: ABDA-NL adds a natural-language interface to ABDA, a system for argument-based discussion using ASPIC- knowledge bases under grounded semantics. Users see which…
- arXiv — cs.AI preprintsInternational2 Oct 2026
PG-SFT: Balancing Capability Acquisition and Retention in Offline Agent Fine-Tuning
arXiv:2610.00949v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) on offline agent trajectories is the standard approach for training specialized tool-using agents, but forcing models to imitate reasoning and…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Cybernetic and Epistemic: A Missing Vocabulary for Trustworthy Agentic Delegation
arXiv:2610.00961v1 Announce Type: new Abstract: As code generation is increasingly delegated to AI systems, the bottleneck is shifting from writing code to supervising the systems that write it --- a shift CS-education…
- arXiv — cs.AI preprintsInternational2 Oct 2026
VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks
arXiv:2610.00972v1 Announce Type: new Abstract: As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be…
- arXiv — cs.AI preprintsInternational2 Oct 2026
RISED: RubrIcs for agentic multi-environment Selection and sElf-Distillation
arXiv:2610.00979v1 Announce Type: new Abstract: Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Evaluating LLM-Generated Preference Distributions
arXiv:2610.01000v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used as probabilistic generators for simulation, synthetic data generation, and decision support in settings where real-world…
- arXiv — cs.AI preprintsInternational2 Oct 2026
Calibration-risk routing for controlled world-model adaptation
arXiv:2610.01001v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL) can exploit simulated experience, but a simulator-to-target shift creates a model-selection problem: correcting the simulator and…