Design Patterns for Agentic Workloads: A 2026 Field Guide

•21 min read
Design Patterns for Agentic Workloads: A 2026 Field Guide

Introduction

Every team building with LLMs eventually hits the same wall. The demo agent that researched, coded, and filed tickets on Friday afternoon burns through your token budget on Monday, spins on a malformed tool response, or returns a confident answer built on a tool call that silently failed. The model didn’t get worse. The system around it was never designed for a component that is probabilistic, expensive, and occasionally wrong.

The good news is that the field has converged on a small vocabulary of composable patterns. SitePoint’s 2026 guide makes the case that fluency in a handful of patterns matters more than mastery of any single framework, and the sources reviewed for this article, from Anthropic’s engineering posts to independent production write-ups, describe largely the same six shapes. This guide walks through them with framework-free Python (exercised against a scripted stand-in model), then covers the cross-cutting layer that makes them production-safe: budgets, typed tool errors, context management, and the 2026 interop stack of MCP and A2A.

We focus on patterns and operational design, not on any one SDK. If you can call an LLM API and parse JSON, you can build everything below.

🎯 Key Takeaways
  • Default to workflows (predefined code paths). Reach for autonomous agents only when you genuinely cannot predict the steps.
  • Six patterns cover most production systems: chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer, and the agent loop.
  • Budgets, stop conditions, and typed tool errors are what turn a pattern from a demo into a service.
  • Multi-agent designs cost roughly 15× the tokens of a chat, so reserve them for parallelizable, high-value work.
  • MCP (agent-to-tools) and A2A (agent-to-agent) are the 2026 interop layer. MCP's latest spec went stateless, so make your application state explicit.
ℹ️ Prerequisites

Python 3.11+, comfort with async/await, and a basic grasp of LLM tool calling. The tests run against a stub model, so no API key is needed. To run the live adapter you need pip install anthropic and an ANTHROPIC_API_KEY.

Workflows, Agents, and the Augmented LLM

Anthropic draws the most useful line in this space. Workflows are systems where LLMs and tools are orchestrated through predefined code paths. Agents are systems where the LLM dynamically directs its own process and tool usage. Both are built on the same primitive, the augmented LLM: a model connected to retrieval, tools, and memory. Anthropic’s post carries a note that much of the surrounding tooling has changed since its December 2024 publication, but the pattern vocabulary it introduced is still the baseline that most of the sources I reviewed build on.

In practice the line blurs. One production write-up observes that most real systems sit in between: a router picks one of a few prebuilt workflows, or an orchestrator hands a narrow task to a tightly constrained agent. Fully free-roaming agents remain rare for good reasons: cost, latency, and compounding errors.

"We recommend finding the simplest solution possible, and only increasing complexity when needed."
— Anthropic Engineering, "Building effective agents"

That advice becomes a decision procedure. Read the chart top-down and stop at the first exit that works. Many applications never get past the first box, because a single well-prompted call with retrieval and good examples is enough.

Yes

No

Yes

Yes

No

Yes

No

Yes

No

Yes

No

New task

Does one well-prompted

call work?

Ship the single call

plus retrieval

Can you list the

steps in advance?

Distinct input

categories?

Routing

Independent

subtasks?

Parallelization

Prompt chaining

with gates

Clear pass/fail

criteria?

Add evaluator-optimizer

Decomposable but

unpredictable?

Orchestrator-workers

Agent loop

with budgets

The Six Patterns at a Glance

Each pattern buys something and charges something. Knowing the bill up front is most of the design work.

PatternUse it whenYou pay withCharacteristic failure
Prompt chainingSteps are fixed and cleanly decomposableLatencyErrors cascade past weak gates
RoutingInputs fall into distinct categoriesOne extra (cheap) callMisclassification
ParallelizationSubtasks are independent, or you want multiple votesTokens from fan-outNaive aggregation
Orchestrator-workersSubtasks can’t be predicted in advanceTokens, plus orchestrator qualityOrchestrator is a single point of failure
Evaluator-optimizerCriteria are clear and feedback measurably helpsExtra iterationsEvaluator and generator converge on the same wrong answer
Agent loopThe path is open-endedCost, compounding errorsRunaway loops, drift

The last row deserves a picture, because every other pattern is a constrained version of it. An agent is an LLM using tools in a loop, grounded by what the environment returns at each step.

LLM decides Tool call Observe result Repeat until done, or until the budget says stop

The loop terminates in one of two ways: the model decides it’s finished, or your code decides it’s finished. Anthropic’s guidance is to include stopping conditions such as a maximum iteration count. That second exit is the one this article keeps returning to.

Practical Implementation

We’ll build the patterns on a deliberately tiny foundation. Frameworks can make these patterns faster to assemble, but Anthropic warns that their abstractions can obscure the underlying prompts and responses and make debugging harder, and suggests starting with direct API calls. So here an “LLM” is just an async function.

1
Treat the model as a function
One signature, (system, user) -> text, makes every pattern testable with a scripted stub and portable across providers.
2
Share one budget across every call
A single wrapper enforces call and time limits no matter which pattern is running.
3
Compose patterns, with gates between steps
Validate in code wherever you can. Use another model call only where you can't.
4
Make tool failures loud, typed, and bounded
The model should see an explicit error, never an empty success.
5
Trace every step and evaluate on real inputs
You can't tune what you can't replay.

Foundation: the LLM type, a shared budget, and a live adapter

# agentic_patterns.py — Python 3.11+, standard library only (adapter needs `anthropic`)
from __future__ import annotations

import asyncio
import json
import time
from dataclasses import dataclass, field
from typing import Awaitable, Callable

# An "LLM" is just an async function: (system_prompt, user_prompt) -> text.
LLM = Callable[[str, str], Awaitable[str]]


class BudgetExceeded(RuntimeError):
    pass


@dataclass
class Budget:
    """One budget shared by every pattern in a run."""
    max_calls: int = 25
    max_seconds: float = 120.0
    calls: int = 0
    started: float = field(default_factory=time.monotonic)

    def charge(self) -> None:
        self.calls += 1
        if self.calls > self.max_calls:
            raise BudgetExceeded(f"call limit {self.max_calls} exceeded")
        if time.monotonic() - self.started > self.max_seconds:
            raise BudgetExceeded(f"time limit {self.max_seconds}s exceeded")


def budgeted(llm: LLM, budget: Budget) -> LLM:
    async def call(system: str, user: str) -> str:
        budget.charge()
        return await llm(system, user)
    return call


# Live adapter (not used by the tests). Pass whichever model ID you run in production.
def anthropic_llm(model: str, max_tokens: int = 1024) -> LLM:
    from anthropic import AsyncAnthropic
    client = AsyncAnthropic()  # reads ANTHROPIC_API_KEY from the environment

    async def call(system: str, user: str) -> str:
        resp = await client.messages.create(
            model=model, max_tokens=max_tokens, system=system,
            messages=[{"role": "user", "content": user}],
        )
        return "".join(b.text for b in resp.content if b.type == "text")
    return call

Chaining and routing

Chaining trades latency for accuracy by making each call an easier task, and the gate between calls is plain code. Routing keeps one input type’s prompt from degrading another’s, and it should always have a safe default for labels the classifier invents.

async def chain(llm: LLM, topic: str) -> str:
    outline = await llm("Write a 5-bullet outline. One bullet per line.", topic)
    if len([l for l in outline.splitlines() if l.strip()]) != 5:  # the gate
        raise ValueError("outline gate failed: expected exactly 5 bullets")
    return await llm("Write a short article that follows this outline.", outline)


Handler = Callable[[str], Awaitable[str]]


async def route(llm: LLM, ticket: str, handlers: dict[str, Handler]) -> str:
    label = (await llm(
        f"Classify the ticket as one of: {', '.join(handlers)}. Reply with one word.",
        ticket,
    )).strip().lower()
    handler = handlers.get(label, handlers["other"])  # unknown label -> safe default
    return await handler(ticket)

Parallelization: sectioning and voting

Sectioning runs independent subtasks side by side. Voting runs the same check several times and thresholds the result. Anthropic notes that splitting guardrail screening into its own call tends to beat asking one call to both answer and police itself.

async def vote(llm: LLM, code: str, n: int = 3, threshold: int = 2) -> bool:
    system = "Review the code. Reply VULNERABLE or SAFE as the first word."
    verdicts = await asyncio.gather(*(llm(system, code) for _ in range(n)))
    flagged = sum(v.strip().upper().startswith("VULNERABLE") for v in verdicts)
    return flagged >= threshold

Orchestrator-workers

Structurally this resembles parallelization, but the subtasks aren’t predefined: the orchestrator decides them per input. That flexibility is also its weakness, so the code below validates the plan, caps the fan-out, and tolerates worker failure. One practitioner write-up describes orchestrators that produce a 14-step plan when three steps would do, and Anthropic’s research team reported that vague delegation led to duplicated work, which they fixed with explicit objectives and explicit boundaries per subtask. The prompt below bakes both lessons in.

Worker BWorker AOrchestratorCallerWorker BWorker AOrchestratorCallerpar[to Worker A][to Worker B]taskplan as JSON, max 5 subtasksobjective + boundariesobjective + boundariesresult or errorresult or errorsynthesized answer
Worker = Callable[[dict], Awaitable[str]]
MAX_SUBTASKS = 5


async def orchestrate(llm: LLM, task: str, workers: dict[str, Worker]) -> str:
    plan_system = (
        "Break the task into at most 5 independent subtasks. Return ONLY a JSON list of "
        '{"worker": <one of: ' + ", ".join(workers) + '>, "objective": str, "boundaries": str}. '
        "Boundaries must say what this subtask must NOT cover."
    )
    plan = json.loads(await llm(plan_system, task))
    plan = [s for s in plan if s.get("worker") in workers][:MAX_SUBTASKS]  # validate + cap
    if not plan:
        raise ValueError("orchestrator produced no valid subtasks")
    results = await asyncio.gather(
        *(workers[s["worker"]](s) for s in plan), return_exceptions=True
    )
    ok = [r for r in results if not isinstance(r, Exception)]
    failed = len(results) - len(ok)
    note = f"\n(NOTE: {failed} subtask(s) failed; say so in the answer.)" if failed else ""
    return await llm("Synthesize these worker results into one answer.", json.dumps(ok) + note)

Evaluator-optimizer

This pattern fits when you have clear criteria and feedback demonstrably improves the output. Two practices from production write-ups matter: give the evaluator its own prompt with binary, checkable criteria rather than a vague rubric, and have it return structured pass/fail results instead of free text.

async def refine(llm: LLM, task: str, criteria: list[str], max_rounds: int = 3) -> str:
    eval_system = (
        "You are a strict reviewer. For each criterion answer true/false. Return ONLY JSON: "
        '{"checks": {<criterion>: bool}, "feedback": "specific fixes for each false check"}'
    )
    draft = await llm("Write the best possible answer.", task)
    for _ in range(max_rounds):
        verdict = json.loads(await llm(
            eval_system, "Criteria:\n- " + "\n- ".join(criteria) + f"\n\nDraft:\n{draft}"))
        if all(verdict["checks"].values()):
            return draft
        draft = await llm("Revise the draft using the feedback.",
                          f"{draft}\n\nFeedback:\n{verdict['feedback']}")
    return draft  # best effort once the round budget is spent

The agent loop, with a leash

Three details separate this from a tutorial loop. Tools return a typed result, so failures can’t masquerade as empty successes. Tool output is truncated before it re-enters the context. And a circuit breaker catches the model repeating an identical call.

@dataclass
class ToolResult:
    ok: bool
    content: str


Tool = Callable[..., Awaitable[str]]


async def run_tool(tools: dict[str, Tool], name: str, args: dict, timeout: float = 10.0) -> ToolResult:
    if name not in tools:
        return ToolResult(False, f"unknown tool '{name}'. Available: {', '.join(tools)}")
    try:
        out = await asyncio.wait_for(tools[name](**args), timeout)
        return ToolResult(True, out[:2000])  # cap what re-enters the context
    except asyncio.TimeoutError:
        return ToolResult(False, f"{name} timed out after {timeout}s")
    except Exception as exc:  # explicit error instead of a silent empty result
        return ToolResult(False, f"{name} failed: {type(exc).__name__}: {exc}")


AGENT_SYSTEM = (
    'Reply ONLY with JSON. To use a tool: {"tool": str, "args": object}. '
    'To finish: {"final": str}. Never repeat an identical call.'
)


async def agent(llm: LLM, task: str, tools: dict[str, Tool], max_steps: int = 8) -> str:
    transcript = [f"TASK: {task}"]
    last_call = None
    for _ in range(max_steps):
        reply = json.loads(await llm(AGENT_SYSTEM, "\n".join(transcript)))
        if "final" in reply:
            return reply["final"]
        call = (reply["tool"], json.dumps(reply.get("args", {}), sort_keys=True))
        if call == last_call:  # circuit breaker on identical consecutive calls
            transcript.append("ERROR: identical call repeated. Change approach or finish.")
            continue
        last_call = call
        result = await run_tool(tools, reply["tool"], reply.get("args", {}))
        transcript.append(f"TOOL {reply['tool']} -> {'OK' if result.ok else 'ERROR'}: {result.content}")
    raise BudgetExceeded(f"agent used all {max_steps} steps without finishing")

Here is what the loop looks like against a scripted model that first asks for a log file that doesn’t exist, repeats itself, then recovers. This is the context the model sees on its final turn:

TASK: Why are we seeing 500s?
TOOL read_log -> ERROR: read_log failed: FileNotFoundError: api.log
ERROR: identical call repeated. Change approach or finish.
TOOL read_log -> OK: 12:01 ERROR db pool exhausted (max=10)

ANSWER: 500s come from DB pool exhaustion (max=10) at 12:01. | calls used: 4
⚠️ This is a teaching harness, not a drop-in

The JSON-in-text tool protocol keeps the code provider-agnostic and testable. In production, prefer native structured tool calling with schema validation, which SitePoint describes as table stakes across the major providers in 2026. The patterns and guardrails carry over unchanged.

What Autonomy Costs

Agentic designs spend tokens to buy capability, and Anthropic’s published numbers make the trade explicit. In its data, agents use about four times the tokens of a chat interaction and multi-agent systems about fifteen times. Its analysis suggests that a large share of multi-agent gains comes simply from spending enough tokens on the problem, which is why the payoff concentrates in breadth-first work with independent threads. Anthropic also notes that most coding tasks have fewer parallelizable pieces than research does.

~4×
Agent vs. Chat Tokens
Typical agent token usage relative to a chat (Anthropic)
~15×
Multi-Agent vs. Chat Tokens
Typical multi-agent token usage relative to a chat (Anthropic)
+90.2%
Research Eval Gain
Opus 4 lead + Sonnet 4 subagents vs. single Opus 4, on Anthropic's internal eval

Treat the 90.2% as a data point about one system on one internal benchmark, not a promise. The durable lesson is the shape of the trade: multi-agent pays off only when the task decomposes into independent parallel threads and the outcome is worth the spend.

The Cross-Cutting Layer

The six patterns are the skeleton. These four disciplines decide whether it stands up under load.

Context is a budget too. Every turn appends to the context window, and accumulated tool output can push the model past the point where it follows its own instructions. A 2025 context-rot report covering eighteen models found performance generally degrades as input grows. Four moves help: compact old turns or stale tool results once a threshold, tuned from real traces, is crossed; isolate work in subagents with their own windows that return distilled summaries; cap tool output (our run_tool truncates at 2,000 characters); and keep the tool manifest small. Every registered tool’s schema is injected on every call, so a sprawling toolset inflates cost and invites wrong-tool selection. Namespaced routing keeps each call’s manifest short.

Tools are an interface you design. Anthropic urges investing in the agent-computer interface as seriously as you would a human one: write tool descriptions like docstrings for a junior engineer, include edge cases and boundaries from similar tools, and poka-yoke the arguments so mistakes are hard to make. Its SWE-bench agent reportedly took more optimization effort on tools than on the prompt, and one fix was requiring absolute file paths after the agent kept stumbling on relative ones.

Assume misfires and shrink the blast radius. A narrower tool signature helps, but a model determined to act destructively will find the narrowest destructive tool it has. Default workers to read-only, scope credentials per worker, and require human approval for irreversible actions. Anthropic’s agent guidance also calls for pausing at checkpoints or blockers.

Observe at the step level. Model-layer metrics miss agent failures. Track tool calls per task (average and p95), steps per run, context tokens per turn, and cost per run, and alert on spikes. Anthropic reported that standard logging wasn’t enough for its multi-agent system and built custom monitoring of agent decisions.

Interop in 2026: MCP and A2A

Two protocols now cover the seams between systems. MCP moved to the Linux Foundation’s Agentic AI Foundation in December 2025, and A2A reached version 1.0 in 2026. They answer different questions.

🔌 MCP: agent ↔ tools and data
  • Host-client-server contract exposing tools, resources, and prompts
  • 2026-07-28 spec: stateless core, so any server instance can handle a request
  • Tasks extension returns a handle for long-running work; clients drive it with tasks/get, tasks/update, tasks/cancel
  • Best for: giving one agent well-described, governed capabilities
🤝 A2A: agent ↔ agent
  • Messages and tasks between independent agents, discovered through Agent Cards
  • v1.0 adds signed Agent Cards for verifiable identity
  • Built to cross organizational and trust boundaries
  • Best for: delegating to a partner's or vendor's agent you don't control
ℹ️ Stateless protocol, stateful application

The MCP maintainers' own example is a server that returns a basket_id which the model passes back as an argument on later calls. Copy that habit: return explicit handles from tools instead of leaning on hidden session state. It also makes workers resumable. Note the 2026-07-28 release contains breaking changes, and Tasks code written against the experimental 2025-11-25 API needs migration.

Common Pitfalls and Troubleshooting

⚠️ The over-eager orchestrator

The orchestrator is a single point of failure for decomposition quality. Cap subtasks, validate the plan before dispatch, and give every worker an objective plus explicit boundaries so two workers don't research the same thing.

⚠️ Silent tool failures

An empty payload, a timeout, or a changed schema can all return HTTP 200. If the model never sees the failure, it improvises and sounds confident. Validate tool responses against a schema, return typed errors the agent must handle, and add a circuit breaker for repeated identical calls, as run_tool and agent do above.

⚠️ Evaluator collapse

Generator and evaluator can converge on the same wrong answer within a couple of iterations, especially with a vague rubric. Use binary criteria, a separate evaluator prompt that never sees the generator's reasoning, and a fixed round cap. Pin the evaluator prompt and sample-grade its verdicts on a regular schedule to catch drift.

🚨 Tool results are untrusted input

Prompt injection arrives through tool results: a web page, an email body, an issue comment, even a filename. Never let content fetched by a tool grant new permissions or trigger destructive tools without an approval step outside the model's control.

⚠️ Unbounded loops and cost spikes

A few runs that consume most of your spend usually mean a loop with an objective but no exit condition. Enforce a shared Budget, a per-agent step limit, and per-tool rate limits, and alert on step-count outliers rather than waiting for the invoice.

Conclusion

Agentic workloads reward restraint. Start with one good call, add a workflow pattern only where your evals show a gap, and let a model steer only the slice of the problem you can’t script. Whatever you build, wrap it in the same operational discipline: one shared budget, typed tool errors, bounded context, step-level traces, and approval gates for anything irreversible. Then use MCP and A2A for the seams rather than bespoke glue.

A rollout order that follows this logic:

Phase 1: Week 1
Baseline and evals
Ship a single call with retrieval. Collect 20 to 50 real inputs as your eval set.
Phase 2: Weeks 2–3
Workflow patterns where evals show gaps
Add chaining, routing, or parallelization with programmatic gates.
Phase 3: Weeks 4–6
Agent loop for the open-ended slice
Budgets, typed errors, tracing, and human checkpoints come first, not last.
Phase 4: Week 7+
Multi-agent and protocols, only if earned
Orchestrator-workers for parallelizable, high-value work. Expose tools over MCP, and use A2A across trust boundaries.
💡 Next Steps

Pick one workflow you run today and put the Budget wrapper and step-level tracing on it before changing anything else. Then read Anthropic's two engineering posts below, and if you expose tools to agents, review the MCP 2026-07-28 release notes before building on session-based assumptions. Anthropic's original agents post also points to its newer Managed Agents write-up for its current approach.


References:

  1. Building effective agents (Anthropic Engineering) - https://www.anthropic.com/engineering/building-effective-agents - Workflow vs. agent distinction, the five workflow patterns, agent loop guidance, tool (ACI) design advice.
  2. How we built our multi-agent research system (Anthropic Engineering) - https://www.anthropic.com/engineering/multi-agent-research-system - Orchestrator-worker architecture, 4× / 15× token figures, 90.2% internal-eval result.
  3. The 2026-07-28 MCP Specification Release Candidate (Model Context Protocol blog) - https://blog.modelcontextprotocol.io/posts/2026-07-28-release-candidate/ - Stateless core, extensions, Tasks lifecycle, basket_id example, breaking-change notice.
  4. MCP 2026-07-28: What’s Changing and How to Migrate (Agentic AI Foundation) - https://aaif.io/blog/mcp-2026-07-28-whats-changing-and-how-to-migrate - Migration context for the stateless spec.
  5. The State of Agentic AI Standards in 2026 (DEV Community) - https://dev.to/alexmercedcoder/the-state-of-agentic-ai-standards-in-2026-mcp-a2a-webmcp-osi-and-the-protocol-stack-taking-3o2l - A2A 1.0 and the layered protocol stack.
  6. Agent Registry release notes (Google Cloud) - https://docs.cloud.google.com/agent-registry/release-notes - A2A v1.0 support at Agent Registry GA, June 2026.
  7. Agentic AI Orchestration: The Architecture Patterns That Actually Work at Scale (DEV Community) - https://dev.to/amasen/agentic-ai-orchestration-the-architecture-patterns-that-actually-work-at-scale-2ip2 - Orchestrator single point of failure; evaluator design with binary criteria.
  8. Agent workflow patterns and anti-patterns (Respan) - https://www.respan.ai/articles/agent-workflow - Hybrid production shapes, over-planning orchestrators, evaluator collapse.
  9. The Definitive Guide to Agentic Design Patterns in 2026 (SitePoint) - https://www.sitepoint.com/the-definitive-guide-to-agentic-design-patterns-in-2026/ - Patterns-over-frameworks thesis; structured tool calling with schema validation.
  10. The 7 New Failure Modes You Inherit the Moment You Ship an Agent (DEV Community) - https://dev.to/gabrielanhaia/the-7-new-failure-modes-you-inherit-the-moment-you-ship-an-agent-1kjm - Runaway loops, context rot reference, blast-radius and injection guidance.
  11. Why Your Agent Keeps Failing: 7 Production Mistakes to Avoid (MagmaLabs) - https://blog.magmalabs.io/?p=11123 - Silent tool errors, schema validation, circuit breakers.
  12. Agentic RAG Failure Modes (Towards Data Science) - https://towardsdatascience.com/agentic-rag-failure-modes-retrieval-thrash-tool-storms-and-context-bloat-and-how-to-spot-them-early - Tool storms, context bloat, monitoring signals.