
AI Agent Memory Poisoning Testing Across Sessions
A CTO is preparing to enable persistent memory in a customer operations agent. The feature remembers account preferences, previous decisions and unfinished work, reducing the information users must repeat. Security has already reviewed authentication and tried prompt injection during a single conversation. Engineering now wants approval to keep those memories between sessions. The unanswered question is whether information from a support attachment can become a lasting instruction that influences a different task next week.
AI agent memory poisoning testing should answer that question before the feature receives production authority. The assessment needs to follow source content into stored memory, observe its use in a genuinely later session, and verify that repair removes harmful influence without losing necessary customer context. For leadership, the result should support a specific decision: enable the feature, narrow its scope, change the design, or hold rollout until material gaps are resolved.
Why persistent memory changes the approval decision
Persistent agent memory is application-managed information retained for future interactions. Depending on the architecture, it may include saved preferences, summaries of previous work, retrieved records or reusable task notes. Storing information does not necessarily change model weights. Its security importance comes from how later workflows select, interpret and act on it.
A source can be removed while a derived memory remains. A new conversation can begin while the application silently loads earlier notes. An agent can remain within its permitted tools and still use a remembered false instruction to make a poor decision. Each possibility changes the evidence a feature owner should require before approval.
Two dated research sources make the lifecycle worth examining. PMPA, published as an arXiv preprint on 12 September 2026, studies malicious external instructions being stored and affecting later sessions in evaluated harness-based agents. MemSecBench, submitted on 29 July 2026, evaluates persistence, downstream consequences and selective repair across controlled agent configurations.
These studies support asking better assessment questions. They do not establish how often businesses encounter memory poisoning, prove that every memory backend is vulnerable, or predict a success rate for your deployment. The practical trigger is a feature or architecture change, rather than a claim that all agents are already compromised.
Include the memory lifecycle explicitly in an AI penetration testing assessment. A scope that mentions chat safety but omits persistent writes, later recall and repair may leave the launch decision unsupported. Existing tests can still contribute evidence, provided their versions, permissions and tested boundaries remain applicable.
The business exposure depends on what memory can influence. Incorrect personalization may create support rework. A remembered routing preference could misdirect confidential content if destination controls permit it. A false account exception could influence a proposed transaction, although completion still depends on independent authorization and approval. Keep those outcomes separate when assigning severity and deciding treatment.
Define ownership and provenance before testing
Begin with a memory inventory that describes who owns each record, who may write it, which future users may recall it, and which business decisions it may influence. The inventory should distinguish private user memories, tenant-shared context, agent-specific notes and organization-wide knowledge. Calling everything “memory” hides differences that determine the potential blast radius.
Ownership must survive every transformation. A tenant-scoped source should not become an unrestricted summary because the summarization process uses a shared service identity. A revoked user should not retain access through previously saved context. Ask engineering to explain how the current user's permissions are applied when a memory is selected and before its content reaches the model.
Keep source trust separate from write authority
A legitimate user may authorize the agent to read an attachment without authorizing that attachment to change future operating rules. Similarly, an internal tool can return customer-controlled text. The connector's identity does not make every returned statement trustworthy. The assessment should examine whether content keeps its original trust classification when it becomes a summary or remembered preference.
The upstream boundary is explained in how indirect prompt injection enters RAG pipelines. This assessment adds a distinct question: after the immediate retrieval ends, what information was saved, under whose authority, and for which future purpose?
Useful provenance identifies the source, originating user or connector, affected owner or tenant, creation time, writer component and relevant policy version. Derived records also need a relationship to their source or predecessor. Record identifiers and access-controlled evidence references can support that relationship without duplicating full confidential documents into every log.
Provenance supports investigation and policy decisions; it does not prove a statement is true or safe. A correctly attributed source can contain misleading information. A signed record establishes integrity and origin under its signing assumptions, rather than permission to override business rules. Treat provenance as evidence used by controls, with those controls tested independently.
Agree on prohibited memory classes before active testing. Examples include credentials, instructions that waive required approval, another tenant's information and unsupported claims of administrative authority. Also name useful memories that must remain available after repair. Without both sets, a team can appear to fix the problem simply by removing the feature's value.
Scope AI agent memory poisoning testing across sessions
Use a representative environment with synthetic accounts, bounded actions and observable memory records. The assessment should compare normal behavior with controlled adverse conditions while retaining the same relevant permissions and business workflow. This is assessment design, not a requirement to expose production users to harmful content.
| Boundary | Question to resolve | Evidence to request | Business decision |
|---|---|---|---|
| Source ingestion | Can lower-trust content influence a proposed memory? | Source identity, ownership and observed processing path | Approve or restrict eligible sources |
| Persistent write | Does the saved record preserve authorized purpose and ownership? | Before-and-after record evidence and write decision | Require review or narrower write policy |
| Later recall | Does a fresh session retrieve the record under current permissions? | Session separation, retrieved record IDs and authorization result | Accept or redesign recall boundaries |
| Business use | Does recalled content change an answer, proposed action or completed action? | Output, tool request and downstream system state | Set severity and restrict authority |
| Selective repair | Is harmful influence removed while needed context remains? | Repair trace, fresh-session checks and benign-workflow results | Re-enable, extend repair or hold rollout |

Require evidence of a genuinely later session
A repeated question in the same conversation can show immediate influence without proving persistence. The report should explain what was reset, what was retained, and why the later interaction represents the deployed session boundary. Depending on the product, that may involve a new conversation, a renewed authenticated context or a restarted worker loading the persistent store.
The assessor should distinguish saved memory from ordinary conversation history, live source retrieval and caches. If the suspect source is still available during the later task, the evidence may not establish which component caused the effect. Controlled comparisons can help isolate the contribution of the stored record without claiming certainty where visibility is missing.
Agree on realistic later tasks and relevant intervening activity. Some records may influence only a matching workflow; others may be loaded more broadly. The assessment should reflect the product's intended use rather than assuming that one failed recall attempt proves a record harmless. Record the tested observation period and conditions.
Separate influence from completed impact
Report whether the content was stored, retrieved, interpreted as guidance, used to propose an action, or used in an action that actually completed. These are different results. A model saying that it changed an account is insufficient evidence that the account changed. A blocked tool request can demonstrate a functioning authorization boundary alongside a memory-integrity weakness.
When memory uses vector search, include retrieval-layer security testing for the relevant collections and ownership filters. A file-based memory design needs its own read and write controls instead. Do not apply vector-specific conclusions to every store merely because both retain information.
Repeat meaningful observations under documented configurations because model-mediated behavior can vary. Report the number of assessed cases, observed outcomes and untested conditions. Avoid turning a small assessment sample into a population-level percentage or a promise that no future poisoning attempt can succeed.
Hypothetical scenario: a remembered exception changes a later decision
This is a hypothetical illustration, not a client case study. A B2B support agent reads customer attachments and remembers account preferences for later tickets. It can draft communications and propose account updates. A separate service controls whether those updates are permitted, and a human must approve sensitive changes.
During an authorized assessment, a synthetic attachment contains an unsupported claim about how a test account should be handled. The system saves a derived note without retaining the source's lower-trust status. In a later session, the agent recalls the note and proposes an exception to the normal review process. No real customer record or production transaction is involved.
If the downstream service rejects the proposal, the evidence supports a narrower finding: persisted influence affected a recommendation, but the tested action boundary held. The remaining consequences could include reviewer confusion, repeated support work or inaccurate customer communications. If a bounded synthetic update completes without the required approval, the assessment has stronger evidence of business impact.
Now remove the original attachment. If the saved note still affects the later workflow, source removal alone did not repair this tested path. The repair must address the derived record and any relevant copies, while confirming that legitimate account preferences remain usable. This is the difference between cleaning an input and restoring reliable operation.
Independent runtime containment evidence remains useful for limiting reachable resources. It does not settle whether an agent makes an incorrect decision through a permitted application path. Leadership should require both forms of evidence when both controls contribute to the approved operating model.
Validate repair before re-enabling persistent memory
Repair has two objectives: stop harmful influence from returning and preserve the useful information the feature is supposed to retain. Deleting the visible record may be necessary, but the assessment should first establish where the same content or its derived meaning can remain active.
If suspicious behavior has occurred in production, coordinate evidence preservation before modifying the store. Retain access-controlled copies of relevant records, identifiers, timestamps, versions and event history under the organization's incident process. Avoid making broad deletions that obscure the source, affected users or historical actions. A validation test can establish possible behavior; it does not by itself establish what happened previously.

Follow the record's descendants and return paths
The repair scope may include summaries, replicas, indexes, caches or pending memory writes, depending on the architecture. A corrected source may coexist with an older derived summary. A delayed job may recreate a removed record. A backup restoration may reintroduce suspect content. These are paths to assess where present, not assumptions about every product.
Require evidence that the repair reached active serving paths. Then examine whether the original ingestion or write weakness remains open. A clean store can become contaminated again if the same policy still promotes lower-trust content into authoritative memory. Store cleanup and prevention of recurrence should have separate closure criteria.
The validation should include fresh sessions and unaffected tasks. Confirm that approved account preferences remain available, ownership boundaries still hold and legitimate workflows have not silently lost necessary context. Assess meaning and observed behavior, rather than relying solely on a search for the exact original wording.
Define what disabling memory actually changes
A product control labelled “memory off” can have several implementation meanings. It might stop new writes, stop recall, or affect only a selected workflow. Verify the behavior of the deployed control, including already running tasks and separate retrieval paths. Do not infer that disabling new writes makes previously saved content inaccessible.
A complete reset may be appropriate when selective repair cannot be demonstrated, but it has a business cost. Users may need to rebuild context, workflows may need revalidation and the team must prevent re-ingestion of the same suspect source. Choose that option explicitly, with a recovery owner and an acceptance test.
Close repair only for the tested configuration and coverage. If a supplier prevents inspection of relevant copies, mark that limitation clearly. Procurement can seek stronger evidence, or the product team can keep memory restricted until it can demonstrate a defensible recovery path.
Evidence, effort and framework mapping
A useful report connects each material observation to a product decision. Request the relevant source and memory identifiers, owner and tenant, session boundaries, model and backend versions, write and recall decisions, downstream results, repair actions and retest outcome. Preserve the distinction between observed failure, configuration concern and unavailable evidence.
Logs should support correlation without becoming another sensitive memory store. Access restrictions, redaction and retention need to cover prompts, summaries and tool parameters as well as conventional application records. Use synthetic content for testing where possible, and agree on the handling of any customer information encountered.
Effort follows the lifecycle and its visibility
Scoping needs an architecture diagram, eligible sources, memory ownership rules, permitted actions and recovery procedures. Engineering, application security and the feature sponsor should agree on the acceptance criteria. Resolve supplier testing restrictions before scheduling work, and document any differences between the test environment and production.
A single-user feature with one store and observable writes is simpler to assess than shared memories across several agents, connectors and background jobs. Permission changes, multiple model configurations, migration paths and derived summaries increase the combinations that need evidence. The broader AI penetration test scope helps identify surrounding controls that should be included or explicitly excluded.
Budget separately for environment preparation, assessment, engineering remediation, data repair and retest. Fixing an overly broad source policy may be limited work. Adding lineage to existing records or separating tenant ownership can require design changes and migration validation. Confirm price and schedule after those dependencies are understood; this guide does not prescribe a universal memory-testing fee or duration.
Use frameworks to organize evidence
| Reference | Application to memory testing | Limit |
|---|---|---|
| OWASP LLM01:2025 Prompt Injection | Untrusted instructions influence later model behavior through saved memory | Persistence must be demonstrated separately from immediate influence |
| OWASP LLM02:2025 Sensitive Information Disclosure | Stored or recalled information becomes available to an unauthorized recipient | Apply when disclosure exposure is relevant; not every poisoned preference leaks data |
| OWASP LLM08:2025 Vector and Embedding Weaknesses | Retrieval-based memory has weaknesses in storage, selection or access partitioning | Conditional on vector or embedding architecture, not a universal memory category |
| NIST AI RMF Core | GOVERN assigns ownership; MAP describes exposure; MEASURE records test evidence; MANAGE supports treatment and acceptance | Voluntary risk-management guidance, not certification or a standalone pass criterion |
These are practical mappings for assessment evidence, not claims that every framework explicitly prescribes this memory test. The NIST AI RMF page identifies AI RMF 1.0 and states that revision is in progress. Record the edition used. Framework references can support an audit or vendor review without establishing that a particular law applies or guaranteeing certification.
What leadership should decide after AI agent memory poisoning testing
Approve an operating envelope that names eligible sources, memory owners, allowed content, recall permissions and the actions remembered information may influence. A useful approval is configuration-specific. It should identify the responsible owner, residual risks and changes that require further validation.
| Decision | Evidence supporting it | Required follow-through |
|---|---|---|
| Enable within the assessed scope | Material write, ownership, recall, action and repair criteria are demonstrated | Monitor relevant changes and retain the approved configuration |
| Run a restricted pilot | Demonstrated controls bound remaining exposure to an accepted level | Name an owner, expiry and criteria for expansion |
| Hold or narrow memory use | Unacceptable cross-owner influence, sensitive exposure, action bypass or unresolved material evidence | Remediate the affected boundary and validate before expansion |
| Keep affected memory disabled after repair | Cleanup is incomplete, harmful influence returns or safe recovery remains unverified | Complete repair and benign-workflow validation before reactivation |

A restricted pilot should have enforceable restrictions. Limiting eligible sources, removing sensitive tools or confining memory to one owner can reduce exposure when the restrictions themselves are verified. A warning in the user interface is insufficient evidence that a shared store respects those limits.
Plan focused reassessment when the backend, write policy, summarizer, retrieval logic, connector set, ownership model or action authority changes. Compare the new design with the previous evidence rather than treating an earlier report as approval for all future configurations.
Before enabling persistent memory or changing its backend, request a memory-lifecycle assessment. Bring the workflow, ownership rules and recovery design. The outcome should show what persists, what later sessions can do with it, and whether the organization can repair harmful influence while retaining useful operation.
Frequently asked questions
Can we test a memory backend before purchasing it?
Yes, if the supplier permits testing. Compare candidates using the same representative workflows, ownership rules and recovery criteria. A backend demonstration alone cannot establish the security of your complete agent integration.
Does encrypted memory prevent poisoning?
Encryption protects confidentiality under its configured threat model. It does not determine whether an authorized writer saves misleading content or whether an agent later treats that content as an instruction.
Must we share production customer conversations with the assessor?
Usually, representative synthetic records can establish the intended boundaries. Agree on any additional evidence needed to reproduce production-specific behavior, with explicit access, confidentiality and retention limits.
Is a memory poisoning assessment the same as investigating an incident?
No. An assessment validates behavior under agreed conditions. An investigation examines historical records to establish affected users, actions and timing. Combine them when needed, while preserving evidence before repair.
Can we assess a supplier's memory feature without backend access?
Some observable behavior can be tested through supported interfaces. Conclusions about internal writes, lineage or deletion may remain limited. Record those gaps and seek supplier evidence for material approval claims.
Should every unexplained agent mistake be classified as memory poisoning?
No. Incorrect output can result from stale information, ordinary model error, application defects or permissions issues. Classification needs evidence linking an adverse memory change to later behavior.
What should a backend migration acceptance test include?
Include ownership and provenance preservation, equivalent recall permissions, exclusion of quarantined content, and recovery validation. Confirm that migration has not revived removed records or broadened who can retrieve retained information.

