Module 6 of 14
Tools, APIs and the Real World
How does an agent take an action outside of a conversation?
Learning objectives. Explain tools and APIs at three levels of depth; design a tool interface an agent can use reliably; understand MCP and A2A as interoperability layers and what they do and do not solve; and apply the reversibility principle to tool permissions.
Core lesson
An AI model generates decisions. Tools are the only way those decisions touch reality.
Without tools, an agent is a very well-read colleague locked in a room with no phone. Everything it produces is a suggestion.
START HERE — what a tool actually is
A tool is a capability the agent can invoke: look up a customer, send a draft, read a file, search the web, create a calendar event, run a calculation.
Mechanically, the model does not “use” the tool. It emits a structured request — call lookup_customer with `{“email”: “…”}“ — and your system executes that request and hands the result back into context. This distinction matters enormously for security: the model proposes; your code disposes. Every authorisation decision, every validation, every rate limit lives in your code, not in the model’s judgement.
An API — application programming interface — is a controlled doorway that lets one software system ask another to do something. Your CRM has one. Your accounting system has one. When people say “we integrated the AI with our systems”, they mean the agent was given tools that call those APIs.
PRACTITIONER — designing tools agents can use well
| Design rule | Why |
|---|---|
| One clear purpose per tool | Overlapping tools produce inconsistent selection |
| Explicit, strict input schemas | Prevents malformed calls and makes validation testable |
| Descriptions that say when not to use it | Negative guidance prevents the most common misuse |
| Meaningful, actionable errors | “Error 500” teaches the agent nothing; “customer_id not found — search by email first” lets it recover |
| Small, curated toolsets | Large toolsets degrade selection accuracy and consume context |
| Read/write separation | Read tools can be liberal; write tools must be deliberate |
| Idempotency where possible | Retries must not create three invoices |
| Confirmation for irreversible actions | The blast-radius control that matters most |
The tool inventory exercise. For each tool, record: name, purpose, inputs, outputs, cost per call, latency, reversibility, required permission, failure behaviour and owner. This table is one of the most valuable documents in any agent system and takes an afternoon to produce.
Categories of tool in business agents: search and retrieval; systems of record (CRM, ERP, ticketing, accounting); communication (email, chat, calendar); documents and files; data and analytics; computation and code execution; browser and computer use; and internal domain tools you build yourself.
Browser and computer-use tools let an agent operate software with no API by seeing the screen and controlling mouse and keyboard. Current status: usable and improving, materially slower, more expensive and less reliable than an API, and with a much larger security surface. Treat it as the integration of last resort, appropriate for legacy systems and one-off tasks, and always sandboxed.
ARCHITECT — interoperability: MCP and A2A
Two open standards address two different problems. Confusing them is common.
MCP — the Model Context Protocol. An open standard for connecting AI applications to tools and data sources: agent-to-tool. It defines how a client discovers and calls tools, reads resources and uses prompt templates, with an OAuth-based authorisation model. The value proposition is combinatorial: build a connector once and any MCP-compatible client can use it, replacing N×M bespoke integrations.
Current status (as of this edition): the specification is actively evolving, with the 2026-07-28 revision moving the protocol to a stateless request/response model, adding header-based routing for gateways, cacheable list results, and an extensions framework; earlier features including sampling and the HTTP+SSE transport are deprecated with a migration window. Expect further change.
The durable principle underneath: tool interfaces should be discoverable, described, authenticated and versioned, so that capability can be composed rather than hard-wired. That principle will outlive any particular revision of the specification.
A2A — Agent2Agent. An open standard for agent-to-agent communication: how one agent discovers another’s capabilities (via an “agent card”), delegates a task, and tracks it to completion across organisational boundaries. Originally contributed by Google, it is now hosted by the Linux Foundation; by April 2026 the project reported version 1.0, 150+ supporting organisations, five official SDKs, signed agent cards for cryptographic identity, and integration into major cloud agent platforms.
The durable principle: agents that work across trust boundaries need identity, capability description, task semantics and version negotiation — the same things any distributed system has always needed.
A caution. Both standards expand what an agent can reach, and therefore expand the attack surface. Tool poisoning — malicious instructions hidden inside a tool’s description or its returned data — and “rug pulls”, where a previously benign remote tool changes its behaviour after approval, are documented risks in MCP ecosystems. Treat every third-party connector as untrusted code: pin versions, review descriptions, isolate credentials, and log every call. Module 11 develops this.
Business example
A proposal agent had five tools: retrieve the service catalogue, retrieve similar past proposals, look up the prospect record, compute pricing from the rate card, and generate a draft document. It had no email tool at all. That single omission — an act of design, not oversight — meant no proposal could ever reach a client without a human sending it. The agent saved roughly two hours per proposal and could not, structurally, embarrass the firm. Sometimes the most important tool decision is the tool you do not provide.
Common mistakes
- Giving an agent an admin API key because scoping was inconvenient.
- Thirty tools where six would do.
- Swallowing errors so the agent never learns that its action failed.
- Allowing irreversible actions without confirmation, particularly send, delete, pay and publish.
- Trusting tool output as fact. Tool results are untrusted input (Module 11).
- Adopting a protocol because it is fashionable rather than because it solves an integration problem you actually have.
Expert insight
Tool design is agent design. Teams spend weeks on prompts and minutes on tool interfaces, then wonder why behaviour is erratic. If you improve one thing this month, rewrite your tool descriptions and tighten your schemas; the return is immediate and permanent.
Knowledge check
- Explain an API at beginner, practitioner and architect level.
- Why does the model never execute the tool itself, and why does that matter for security?
- Give four properties of a well-designed tool interface.
- What problem does MCP solve, and what problem does A2A solve?
- Why is idempotency important for agent tools?
- When is computer use justified, and what must accompany it?
- Why are tool results treated as untrusted input?
Challenge
Build a complete tool inventory for one agent you use or plan to build. Classify each tool by reversibility. For every irreversible action, define the control that prevents an unintended one. Present the result as a one-page table.
Related terms