An AI agent is a system where a large language model dynamically directs its own processes and tool usage — as opposed to a workflow, where LLMs and tools are orchestrated through predefined code paths. The distinction comes from Anthropic's engineering guide on building effective agents, published December 19, 2024, and it structures everything else.
What is the difference between an agent and a workflow?
A workflow is a pipeline the developer draws in advance: retrieve, summarize, classify, each step an LLM call wired into fixed code. An agent is given a goal and decides its own path — which tools to call, in what order, when to stop. Anthropic's guide draws the line precisely: workflows are systems where LLMs and tools are orchestrated through predefined code paths, while agents are systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks.
The practical consequence is a trade between predictability and flexibility. Workflows are debuggable, cheap, and auditable, because behavior is bounded by the code. Agents handle open-ended tasks but consume more tokens, fail in stranger ways, and need guardrails. The guide's most quoted finding is that the most successful implementations were not using complex frameworks or specialized libraries — they relied on simple, composable patterns.
What are agents built from?
The anatomy is smaller than the marketing suggests. Per Anthropic's guide, the basic building block of agentic systems is an LLM enhanced with augmentations such as retrieval, tools, and memory. Everything else — the agent loop, the planner, the multi-agent swarm — is arrangement of those pieces.
| Component | Role | Documented example |
|---|---|---|
| Model | Reasoning and decision core | Frontier LLM with tool-calling |
| Tools | Actions: search, code execution, APIs | Function calling to external systems |
| Retrieval | Grounding in private knowledge | Document search before answering |
| Memory | State across steps and sessions | Conversation and scratchpad state |
| Orchestration | The loop that decides next steps | Predefined workflow or dynamic agent loop |
The engineering discipline is in the boundaries: which tools exist, what each returns, and when the loop must stop. An agent with a vague stopping condition will happily run until the budget does, which is why cost ceilings and step limits are architectural features, not afterthoughts.
What is the agent loop, concretely?
Strip the marketing and a single control cycle remains. The model receives a goal and a tool list; it emits either a tool call or a final answer; the environment executes the call and returns the result; the model looks at the result and chooses again. That is the whole engine — one loop, iterated until the model produces an answer, hits a step limit, or trips a cost ceiling. Everything called an agent framework is scaffolding around this cycle: formatting tool results, persisting the running transcript, retrying failed calls, and enforcing the guardrails.
The transcript, usually called the context or scratchpad, deserves particular attention because it is both the agent's memory and its bottleneck. Everything the agent knows about its task must fit there, which is why long tasks degrade: the loop accumulates tool outputs, the context fills, and earlier instructions lose influence. Production architectures manage this deliberately — summarizing completed steps, offloading state to external stores, and splitting tasks before the context forces the issue. An engineer who can draw the loop on a whiteboard and say where its context overflows understands agent architectures better than most product pages explain them.
How do multi-agent systems divide labor?
When one loop is not enough, systems split into specialists: a coordinating agent decomposes the task, subagents own bounded slices with their own tool access, and results flow back through a synthesis step. Anthropic's guide describes this as a pattern worth reaching for when subtasks are genuinely parallel and separable — different tools, different contexts, different responsibilities — and cautions against it otherwise, because every additional agent multiplies token spend and adds a failure surface.
The division of labor follows organizational logic more than computational necessity. A research agent with web tools, a code agent with a sandbox, and a writing agent with document access mirror the desks of a small team, and the coordinator plays project manager. The design questions are the same ones any manager faces: who owns what, who reports to whom, and what happens when one contributor fails. The A2A protocol extends that org chart across company boundaries, letting the research agent come from one vendor and the code agent from another while both speak the same task language.
How do agents talk to each other?
Single-agent systems hit a ceiling quickly, and 2025-2026 produced a standard answer: the Agent2Agent protocol, A2A. Google introduced it in April 2025 and donated it to the Linux Foundation, and the Foundation announced on April 9, 2026, that the protocol surpassed 150 supporting organizations in its first year, with deep integration across Google, Microsoft, and AWS platforms and active production deployments across multiple industries.
A2A addresses the horizontal problem: agents built on different frameworks, by different vendors, discovering each other, exchanging task descriptions, and negotiating capabilities. It is deliberately complementary to the tool-facing protocols — one standard connects agents to tools and data, another connects agents to agents. Together they let an enterprise compose a travel-booking agent from one vendor with an expense agent from another without bespoke integration.
How do you evaluate an agent framework's claims?
Because the space is crowded and the demos are polished, a repeatable checklist beats impressions:
- Ask whether the product is a workflow or an agent; the label is often swapped for the more exciting one.
- Identify the tools the agent can call and who controls their permissions.
- Check what bounds the loop: step limits, cost limits, and stopping conditions.
- Look for named evaluations of the underlying model on the task's domain, not aggregate benchmark scores.
- Confirm how state and memory persist across sessions, and who can inspect them.
What can't these systems do yet?
Documentation is candid if read closely. Anthropic's guide recommends starting with the simplest solution and adding complexity only when measured performance demands it — an admission that agentic behavior still costs reliability. The Linux Foundation's first-year A2A report is a milestone of adoption, not of capability: production deployments exist, and cross-vendor agent collaboration at scale is still early. Long-horizon autonomy, verified correctness of multi-step plans, and predictable cost remain open engineering problems, and any architecture that claims to have solved them is claiming more than its documentation supports. The reliable pattern so far: small, composable, instrumented — and simple until proven insufficient. That is also the best summary of the architecture as a whole: one loop, a handful of tools, a guardrail, and a standard way to talk to the next agent over.

