Implementing Amazon S3 Vectors as Persistent Memory for NVIDIA NeMo Agent Toolkit

· Engineer's Notes · Cem Koyluoglu

A guide to using Amazon S3 Vectors as a custom memory provider in NVIDIA NeMo Agent Toolkit, covering setup, plugin implementation, and deployment on EKS.

What happened

The post details how to integrate Amazon S3 Vectors as the persistent memory layer for the NVIDIA NeMo Agent Toolkit (NAT). NAT is an open‑source framework that supports agent orchestration, profiling, evaluation, and optimization across multiple LLM‑based agent frameworks. Its memory subsystem is extensible through a plugin interface that requires a MemoryEditor implementation with add_items(), search(), and remove_items() methods. The article walks through three concrete steps: (1) creating an S3 Vectors bucket and index with a 1024‑dimensional float32 space that matches Amazon Titan Text Embeddings V2, (2) implementing a custom MemoryEditor plugin that generates embeddings via Bedrock, stores vectors in S3 Vectors, and translates search filters into metadata queries, and (3) configuring NAT’s YAML workflow to use the new provider and the auto_memory_agent wrapper, which automatically captures and injects memory for each agent call. The example uses an investment‑research use case and demonstrates how user messages and agent responses are stored and retrieved without explicit LLM calls.

Why it matters in production

Using S3 Vectors as a memory backend satisfies several production requirements. First, semantic retrieval is enabled by the vector index, allowing agents to find relevant past conversations or facts based on cosine similarity. Second, the index schema supports rich metadata—including user_id, agent_id, team_id, and custom fields such as ticker—while marking the content field as non‑filterable to preserve embedding quality. Third, S3 Vectors offers strong write consistency and elastic scaling to billions of vectors, which is critical for large‑scale multi‑agent deployments. Fourth, the plugin design keeps NAT’s core logic unchanged; only the memory provider is swapped, simplifying maintenance. Finally, the article highlights responsible AI considerations: retention policies, PII redaction, and least‑privilege IAM scopes, which mitigate data‑handling risks in persistent storage.

Engineering takeaways

  • Create a dedicated S3 Vectors bucket and index with 1024‑dimensional float32 vectors and a metadata configuration that marks the content field as non‑filterable. This matches the output of Amazon Titan Text Embeddings V2.
  • Implement a MemoryEditor plugin that: (1) generates embeddings via Bedrock, (2) stores each MemoryItem as a vector with a UUID‑based key, and (3) builds a metadata dictionary that includes user_id, memory_type, agent_id, team_id, task_id, confidence, created_at_epoch, is_shared, source, and optional domain fields such as ticker.
  • Configure NAT’s YAML to reference the new provider (_type: s3vectors_memory) and to enable the auto_memory_agent wrapper with options such as save_user_messages_to_memory, retrieve_memory_for_every_response, and save_ai_messages_to_memory.
  • Apply retention and security controls by setting a retention policy on the S3 Vectors bucket, removing obsolete vectors through the remove_items() method, and scoping IAM permissions to the specific bucket and index.
  • Validate performance by profiling token usage, latency, and throughput with NAT’s built‑in profiling tools, ensuring that vector queries and embedding generation do not become bottlenecks in the agent workflow.

Related work on this site

  • Agentic & LLM Automation — Multi-step LLM pipelines that run unattended: model cascades with fallbacks, validation gates, scheduled automation and alerting when something breaks.
  • RAG & Grounded LLM Systems — LLM features that answer from the right source instead of from memory: document-grounded assistants, transcript-grounded chat and quality gates around generated output.
  • YouTube AI Summarizer — Open-source Chrome extension on the Chrome Web Store: AI summaries, key points, deep analysis, a two-host AI podcast (Gemini TTS) and transcript-grounded chat for any YouTube video, using Groq or Ollama Cloud with the user’s own key.
  • Automated AI News Pipeline (this site) — The pipeline behind this site’s Tech News: scheduled GitHub Actions scrape sources, an LLM cascade on Groq rewrites and enhances articles, and layered quality gates (date integrity, language checks, instruction-leak detection, duplicate detection) decide what gets published, with Telegram alerts.

This note was drafted with AI assistance from the primary source credited on this page, and automatically checked against that source before publishing.

Source: https://aws.amazon.com/blogs/machine-learning/build-agent-memory-with-nvidia-nemo-agent-toolkit-and-amazon-s3-vectors/

← All tech news