Skip to content

Tool Integration

Draft

Models are more useful when they can act, and more dangerous.

  • Models can read files, run code, and query APIs
  • Tools extend capability beyond text generation
  • Tools enable agentic workflows
  • Tool results ground output in real data
  • Define functions the model can invoke
  • The model decides when and how to call
  • Results feed back into generation
  • A standard interface such as the Model Context Protocol (MCP)
  • A server provides tools, resources, and prompts
  • The harness connects to servers and orchestrates calls
  • The model generates API calls
  • The system executes them and returns results
  • Verify API responses before acting on them
  • Descriptions written for the model: purpose, inputs, edge cases, and how it differs from similar tools (Anthropic)
  • Minimal required parameters: fewer ways to call it wrong
  • Structured errors: error code, cause, and next step the model can act on
  • Idempotency: a retried call does not double the effect; accept a request key on writes
  • Pagination and filters: return a page and a cursor, not the whole dataset, to protect context
  • Concise results: return the fields the task needs, not the full record
  • Every loaded tool definition takes context on every request
  • Large catalogs slow agents and raise cost, and so do large intermediate results (Anthropic)
  • Load tools on demand: a small core set, plus search or per-task loading for the rest
  • Similar tools with vague descriptions raise wrong-tool calls
  • Tool output is untrusted input: fetched pages, files, tickets, and API results can carry instructions
  • Indirect prompt injection plants instructions in data the model will retrieve (Greshake et al., OWASP LLM01)
  • RAG and fine-tuning do not fully mitigate prompt injection (OWASP LLM01)
  • Least privilege: scope each credential to the tools and resources the task needs
  • Break the trifecta: avoid private data, untrusted content, and an exfiltration path in one session (Willison)
  • Approval gates: outward-facing or destructive calls wait for a human (Checkpoint Gates, Human in the Loop)
  • Sandbox execution: run generated code with no ambient credentials
  • Audit log: every tool call, arguments, caller, and result status (Observability)
  • Tool: a ticket-tracker integration for a triage agent
  • Read tools: search_tickets, get_ticket; paginated, return summary fields only
  • Write tools: add_comment, update_status; separate server, separate credential
  • Scope: read token limited to one project; write token issued only after approval
  • Injection handling: ticket bodies are treated as data; the session that reads them has no write tools loaded
  • Checkpoint: the agent outputs a proposed comment and status change; a human approves before a write step runs it
  • Errors: update_status returns invalid_transition with the allowed states, so the agent can correct itself
  • Too many tools loaded at once (model gets confused, context fills)
  • Vague tool descriptions
  • No error handling, or errors the model cannot act on
  • Trusting tool output without validation
  • One broad credential shared by read and write tools
  • Outward-facing actions with no approval step