AI Agent Research Papers and Surveys
Clear reviews of the research shaping AI agents—from reasoning and tool use to self-evolution, evaluation, and safety.
18 agent research digests.
No agent research matches your search.

Self-Evolving Agents: Survey, Taxonomy, and How Self-Improving AI Agents Work
A self-evolving agents survey and taxonomy: how self-improving AI agents update memory, tools, workflows, and weights, with pap...

When AI Agents Get Internal Access: The Security Boundary Has Changed
How autonomous agents turn familiar security weaknesses into multi-stage intrusions—and why identity, least privilege, isolatio...

Coding Agents Enter the FinOps Era: Cost per Successful Task
Why coding-agent economics depend on cost per successful task, model routing, deterministic verification, two-tier budgets, and...

Stateless MCP and the Rise of the Agent Gateway
What MCP 2026-07-28 changes: stateless requests, gateway-centered authorization, per-request versioning, observability, and saf...

Measuring Reward-Seeking in RL-Trained Models
How Contrastive SDF tests whether RL-trained models follow user intent or inferred grader preferences, with evidence from o3 tr...

Stateful Long-Horizon Agents: 10 Key Papers
A review of 10 recent papers on stateful long-horizon agents, covering memory operations, execution continuity, verified recove...

HarnessDev Benchmark: Can LLMs Improve Their Harnesses?
How HarnessDev tests LLM-built agent harnesses: creation results, executor transfer, held-out evolution gains, and the cost of ...

Pi Harness v2: Crash Recovery and Tool Replay
How Pi Agent Harness v2 handles crash recovery, durable operations, safe tool replay, staged results, and the limits of exactly...

How DeepSeek Harness PTC Mode Actually Works
DeepSeek Harness PTC Mode explained from source: how the generated Code Mode SDK and run_code execute multi-tool programs, plus...

DeepSeek Harness Modes Explained
Compare DeepSeek Harness Standard, PTC, Minimal, and Creator modes by tools, execution model, security boundary, and best use c...

DeepSeek Harness Creator Mode: How It Works and Its Risks
DeepSeek Harness Creator Mode explained from source: runtime inspection, dynamic plugins, custom agent presets, process lifetim...

DeepSeek Harness vs Pi Agent: Plugin Runtime or Minimal Coding Harness?
DeepSeek Harness and Pi Agent compared across architecture, extensions, MCP, sandboxing, sessions, deployment, and best-fit eng...

DeepSeek Harness Architecture: How It Works and Whether to Adopt It
How DeepSeek Harness uses Cordis plugins for models, tools, sessions, approvals, and sandboxes—and what its developer-preview s...

Cordis Spatiotemporal Composability: Revertible Effects, Reactive Coeffects, and Fibers
A practical explanation of Cordis temporal and spatial composability, including revertible effects, reactive coeffects, Fiber l...

Real-Time Voice Agents Are Distributed Systems
Why production voice agents are distributed systems: media transport, full-duplex inference, state handoffs, capacity, and audi...

Agents Enter the Physical World: From VLA Policies to World Action Models
Why physical agents need world action models alongside vision-language-action policies—and what current robotics results do and...

The Model Gateway Is Becoming the Control Plane for Enterprise AI
Why model gateways are becoming the enterprise control plane for AI routing, budgets, quality, policy, and audit.

The Agent Framework Is Not the Runtime: Why Harnesses Are Taking Over Production
Why production agent systems separate framework APIs from a harness that owns state, tool boundaries, control loops, and observ...

Self-Evolving Agents: Survey, Taxonomy, and How Self-Improving AI Agents Work
A self-evolving agents survey and taxonomy: how self-improving AI agents update memory, tools, workflows, and weights, with papers and a build guide.

When AI Agents Get Internal Access: The Security Boundary Has Changed
How autonomous agents turn familiar security weaknesses into multi-stage intrusions—and why identity, least privilege, isolation, and behavioral monitoring n...

Coding Agents Enter the FinOps Era: Cost per Successful Task
Why coding-agent economics depend on cost per successful task, model routing, deterministic verification, two-tier budgets, and unified traces.

Stateless MCP and the Rise of the Agent Gateway
What MCP 2026-07-28 changes: stateless requests, gateway-centered authorization, per-request versioning, observability, and safer enterprise deployment.

Measuring Reward-Seeking in RL-Trained Models
How Contrastive SDF tests whether RL-trained models follow user intent or inferred grader preferences, with evidence from o3 training checkpoints.

Stateful Long-Horizon Agents: 10 Key Papers
A review of 10 recent papers on stateful long-horizon agents, covering memory operations, execution continuity, verified recovery, and delayed actions.

HarnessDev Benchmark: Can LLMs Improve Their Harnesses?
How HarnessDev tests LLM-built agent harnesses: creation results, executor transfer, held-out evolution gains, and the cost of running the generated systems.

Pi Harness v2: Crash Recovery and Tool Replay
How Pi Agent Harness v2 handles crash recovery, durable operations, safe tool replay, staged results, and the limits of exactly-once execution.

How DeepSeek Harness PTC Mode Actually Works
DeepSeek Harness PTC Mode explained from source: how the generated Code Mode SDK and run_code execute multi-tool programs, plus token and safety limits.

DeepSeek Harness Modes Explained
Compare DeepSeek Harness Standard, PTC, Minimal, and Creator modes by tools, execution model, security boundary, and best use case.

DeepSeek Harness Creator Mode: How It Works and Its Risks
DeepSeek Harness Creator Mode explained from source: runtime inspection, dynamic plugins, custom agent presets, process lifetime, and shell-level risks.

DeepSeek Harness vs Pi Agent: Plugin Runtime or Minimal Coding Harness?
DeepSeek Harness and Pi Agent compared across architecture, extensions, MCP, sandboxing, sessions, deployment, and best-fit engineering workflows.

DeepSeek Harness Architecture: How It Works and Whether to Adopt It
How DeepSeek Harness uses Cordis plugins for models, tools, sessions, approvals, and sandboxes—and what its developer-preview status means for adoption.

Cordis Spatiotemporal Composability: Revertible Effects, Reactive Coeffects, and Fibers
A practical explanation of Cordis temporal and spatial composability, including revertible effects, reactive coeffects, Fiber lifecycles, and system limits.

Real-Time Voice Agents Are Distributed Systems
Why production voice agents are distributed systems: media transport, full-duplex inference, state handoffs, capacity, and audio-aware evaluation.

Agents Enter the Physical World: From VLA Policies to World Action Models
Why physical agents need world action models alongside vision-language-action policies—and what current robotics results do and do not show.

The Model Gateway Is Becoming the Control Plane for Enterprise AI
Why model gateways are becoming the enterprise control plane for AI routing, budgets, quality, policy, and audit.

The Agent Framework Is Not the Runtime: Why Harnesses Are Taking Over Production
Why production agent systems separate framework APIs from a harness that owns state, tool boundaries, control loops, and observability.