Postman Agent Mode on Amazon Bedrock: Production Lessons for AI Agents
· Engineer's Notes · Cem Koyluoglu
Analysis of Postman's Agent Mode architecture, tool management, schema reads, and context handling for scalable AI agents on Bedrock.
What happened
Postman introduced Agent Mode, an AI‑native interface that lets developers interact with the product for testing, documentation, discovery, and implementation. The service runs on Amazon Bedrock, leveraging managed foundation models and Guardrails to redact personal data. Early design assumed tool availability would be the main obstacle, but practical experience revealed three dominant challenges: tool sprawl, context scarcity, and the need for schema‑based data access. The team reduced an initial catalog of more than 170 tools to roughly 15 relevant ones per request by querying a vector database of tool embeddings. They also shifted from atomic UI‑driven tools to schema‑aware query tools that operate on ClickHouse tables, allowing a single read tool to answer many analytical questions. Context handling required dedicated handlers that distill UI entities into purpose‑shaped information, because the rendering data model proved noisy and exceeded the model’s context window. Human oversight remains integral; actions that modify application state require explicit user approval, and Guardrails can be toggled by enterprise admins. Model flexibility is achieved through Bedrock’s support for multiple Claude models, enabling routing of latency‑sensitive traffic to faster models and complex reasoning to larger, higher‑quality models.
Why it matters in production
The patterns described illustrate how large‑scale AI agents must balance latency, cost, and reliability. Dynamically scoping tools reduces round‑trip latency by limiting the number of model calls required for multi‑step workflows, which is critical when traffic exhibits sharp bursts from a global developer base of 40 million users. Schema‑based reads replace a proliferation of single‑purpose tools, lowering operational overhead and improving cost efficiency because fewer inference calls are needed to generate complex queries. Context budgeting directly impacts failure rates; insufficient or noisy context leads to incorrect tool usage even when the model is capable. By treating context as a scarce resource and using purpose‑shaped handlers, Postman mitigates hallucinations and improves deterministic behavior. Guardrails and mandatory user approval add safety layers that protect against unintended state changes and data exposure, aligning the service with responsible AI standards. Model flexibility via Bedrock’s Claude family allows Postman to adapt to workload characteristics without redeploying custom model servers, simplifying operations and reducing infrastructure maintenance.
Engineering takeaways
- Scope tools per request: Query a vector store of tool embeddings to expose only the most relevant subset (e.g., ~15 of >170) to the model, reducing latency and tool‑selection errors.
- Prefer schema‑aware read APIs: Provide agents with access to structured query engines (e.g., ClickHouse) instead of building numerous single‑purpose read tools; this trades tool count for data modeling effort and scales better.
- Design purpose‑shaped context handlers: Extract only the information the model needs from UI entities, treat the context window as a limited budget, and implement truncation strategies to avoid crowding the prompt.
- Enforce human approval for state changes: Require explicit user consent before any action that mutates application state, and use Guardrails to redact sensitive data before model ingestion.
- Leverage model flexibility: Use a managed inference platform that supports multiple model families, routing latency‑critical calls to faster models and complex reasoning to larger models without code changes.
Related work on this site
- Agentic & LLM Automation — Multi-step LLM pipelines that run unattended: model cascades with fallbacks, validation gates, scheduled automation and alerting when something breaks.
- RAG & Grounded LLM Systems — LLM features that answer from the right source instead of from memory: document-grounded assistants, transcript-grounded chat and quality gates around generated output.
- YouTube AI Summarizer — Open-source Chrome extension on the Chrome Web Store: AI summaries, key points, deep analysis, a two-host AI podcast (Gemini TTS) and transcript-grounded chat for any YouTube video, using Groq or Ollama Cloud with the user’s own key.
- Automated AI News Pipeline (this site) — The pipeline behind this site’s Tech News: scheduled GitHub Actions scrape sources, an LLM cascade on Groq rewrites and enhances articles, and layered quality gates (date integrity, language checks, instruction-leak detection, duplicate detection) decide what gets published, with Telegram alerts.
This note was drafted with AI assistance from the primary source credited on this page, and automatically checked against that source before publishing.