<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE rfc [
  <!ENTITY nbsp "&#160;">
  <!ENTITY zwsp "&#8203;">
  <!ENTITY nbhy "&#8209;">
  <!ENTITY wj "&#8288;">
]>
<rfc xmlns:xi="http://www.w3.org/2001/XInclude"
     category="info"
     docName="draft-khandelwal-bmwg-agent-memory-integrity-00"
     ipr="trust200902"
     submissionType="IETF"
     consensus="false"
     xml:lang="en"
     version="3">

  <front>
    <title abbrev="Agent Memory Integrity Benchmark">
      A Benchmarking Method for the Integrity of AI Agent Memory at Rest
    </title>
    <seriesInfo name="Internet-Draft"
                value="draft-khandelwal-bmwg-agent-memory-integrity-00"/>

    <author fullname="Yasha Khandelwal" initials="Y." surname="Khandelwal">
      <organization>Tech4Biz Solutions</organization>
      <address>
        <email>yasha.khandelwal@tech4biz.io</email>
      </address>
    </author>

    <date year="2026"/>

    <area>Operations and Management</area>
    <workgroup>Benchmarking Methodology Working Group</workgroup>
    <keyword>agent memory</keyword>
    <keyword>integrity</keyword>
    <keyword>benchmarking</keyword>
    <keyword>LLM agents</keyword>
    <keyword>tamper evidence</keyword>

    <abstract>
      <t>
        AI agents increasingly persist memory across sessions and treat that
        memory, on the next turn, as if it were their own prior experience.
        This document defines a benchmarking method that measures whether an
        agent's memory subsystem detects that its persisted memory has been
        altered, removed, reordered, replayed, or forged at the storage layer,
        and refuses to serve that memory or reports it before it is served.
        The method defines eight storage-level edits, three verdict classes, a
        detection-point distinction between read time and audit time, two
        control cases, and a scoring rule. It is a laboratory method for
        controlled, reproducible measurement, in the spirit of RFC 2544 and
        RFC 8239, and it is intended as a test method for the "Protection of
        Memory Data Integrity" metric under discussion in the Benchmarking
        Methodology Working Group.
      </t>
    </abstract>
  </front>

  <middle>

    <section anchor="intro">
      <name>Introduction</name>
      <t>
        An AI agent built on a framework such as LangGraph, Letta, Mem0, or a
        vector memory store keeps state between sessions: long-term memory
        records, session checkpoints, and their metadata. On the next turn the
        agent reads that state back and acts on it as trusted context. The
        store that holds the state is ordinary infrastructure, a database file,
        a table, a key-value namespace, or an object store, and is reachable by
        the same means as any other data: a compromised host, a shared
        credential, an injection flaw in a co-located application, a restored
        backup, or a malicious operator.
      </t>
      <t>
        There is at present no agreed way to measure whether an agent notices
        when the memory behind it has been changed. Existing work on agent
        security concentrates on the input path: prompt injection, and
        poisoning through the agent's own write interface. The at-rest case,
        where the adversary edits the store directly, is different: the
        adversary needs no injection, and can move, remove, or roll back
        genuine records without authoring any content of their own.
      </t>
      <t>
        This document specifies a laboratory benchmarking method for that case.
        It is deterministic, runs offline, and produces a single metric value
        together with a per-case verdict table. It follows the conventions of
        the Benchmarking Methodology Working Group: a controlled test bed, an
        explicit procedure, control cases that must hold for a result to count,
        and full reporting of the configuration under test
        <xref target="RFC2544"/> <xref target="RFC8239"/>.
      </t>
    </section>

    <section anchor="scope">
      <name>Scope and Relationship to Other Metrics</name>
      <t>
        This method measures one property: the ability of a System Under Test
        (SUT) to detect at-rest tampering with its own persisted memory and to
        avoid serving tampered memory as genuine. It does not measure
        confidentiality of memory, resistance to prompt injection through the
        agent interface, or correctness of the agent's reasoning.
      </t>
      <t>
        The adversary here is distinct from two adjacent adversaries discussed
        elsewhere. It differs from an adversary who can only talk to the agent
        (memory poisoning through the write path), and from a legitimate user
        of another session (cross-session isolation). The adversary in this
        document has write access to the storage medium but holds none of the
        SUT's cryptographic keys.
      </t>
      <t>
        The method is intended to serve as the test method for a "Protection of
        Memory Data Integrity" metric. It can be used on its own for any agent
        memory subsystem.
      </t>
    </section>

    <section anchor="terms">
      <name>Terminology</name>
      <t>
        The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT",
        "SHOULD", "SHOULD NOT", "RECOMMENDED", "MAY", and "OPTIONAL" in this
        document are to be interpreted as described in BCP 14
        <xref target="RFC2119"/> <xref target="RFC8174"/> when, and only when,
        they appear in all capitals, as shown here.
      </t>
      <dl newline="false">
        <dt>Memory subsystem:</dt>
        <dd>
          the component of the SUT that persists and later retrieves the
          agent's state: long-term memory records, session checkpoints, and
          their metadata.
        </dd>
        <dt>Store:</dt>
        <dd>
          the medium that holds persisted memory: a database file, a table, a
          key-value namespace, or an object store.
        </dd>
        <dt>Write path:</dt>
        <dd>
          the SUT's own interface for adding memory (for example, remember,
          add, put, or checkpoint).
        </dd>
        <dt>Read path:</dt>
        <dd>
          the SUT's own interface for retrieving memory for use in a turn (for
          example, recall, search, get state, or resume).
        </dd>
        <dt>Context:</dt>
        <dd>
          an isolation unit the SUT exposes for memory: a user, a thread, or a
          session.
        </dd>
        <dt>Edit:</dt>
        <dd>
          a single change applied to the store by the adversary using generic
          storage tooling, with no SUT code loaded.
        </dd>
      </dl>
    </section>

    <section anchor="threat">
      <name>Threat Model</name>
      <t>
        The adversary has write access to the medium that holds the SUT's
        memory but holds none of the SUT's cryptographic keys. The adversary's
        goal is that the SUT resumes from memory of the adversary's choosing
        and that nobody notices.
      </t>
      <t>
        Because the adversary can write anything to the store, an integrity
        value stored next to the data, a checksum column, a hash field, or a
        "verified" flag, provides no protection unless it is bound to a secret
        or a root of trust the adversary does not hold. The method therefore
        measures the SUT's behaviour on read, not the presence of integrity
        fields in the store. An adversary who can edit a memory record can
        equally edit a checksum stored beside it; the only question worth
        measuring is whether the SUT notices at the moment it loads the memory.
      </t>
    </section>

    <section anchor="testbed">
      <name>Test Bed</name>
      <t>
        The memory subsystem MUST be run in the configuration the SUT's
        documentation recommends for production, including any at-rest
        encryption, signing, or audit feature the vendor documents as
        protecting memory. A feature that is off by default MUST be reported as
        off by default. Where the SUT offers such a feature, the evaluator
        SHOULD report two results, one with the feature off and one with it on,
        so that the two can be compared.
      </t>
      <t>
        The evaluator seeds memory through the SUT's own write path: at least
        five records in each of two isolated contexts (two users, threads, or
        sessions, whichever the SUT exposes). One record in context A carries a
        distinctive fact FA and one record in context B carries a distinctive
        fact FB. Nothing is written directly to the store during seeding.
      </t>
      <t>
        All reads are made through the SUT's own read path and never by
        inspecting the store. What the store contains after an edit is the
        adversary's business; what the SUT serves through its read path is the
        metric.
      </t>
    </section>

    <section anchor="cases">
      <name>Test Cases</name>
      <t>
        Each case starts from a fresh copy of the seeded store. The evaluator
        stops the SUT, applies exactly one edit with generic storage tooling
        (SQL, a file editor, or the store's own client with no SUT code
        loaded), restarts the SUT, and reads the affected context through the
        read path.
      </t>
      <dl newline="true">
        <dt>T1 Content tamper.</dt>
        <dd>Change the fact inside one existing record in place.</dd>
        <dt>T2 Tail truncation.</dt>
        <dd>
          Delete the most recent record or records so that an earlier state
          becomes current.
        </dd>
        <dt>T3 Middle deletion.</dt>
        <dd>Delete a record that is neither the first nor the last.</dd>
        <dt>T4 Reordering.</dt>
        <dd>
          Swap the position of two records by editing their order keys,
          timestamps, or parent links.
        </dd>
        <dt>T5 Forged insertion.</dt>
        <dd>
          Insert a new record of the adversary's authorship, in the SUT's own
          storage format.
        </dd>
        <dt>T6 Cross-context replay.</dt>
        <dd>
          Copy a genuine record from context A over a record in context B,
          leaving B's identifiers in place.
        </dd>
        <dt>T7 Rollback replay.</dt>
        <dd>
          Copy a genuine older record of context A over A's most recent record.
        </dd>
        <dt>T8 Metadata tamper.</dt>
        <dd>
          Alter a record's owner, role, source, or timestamp field and leave
          its content unchanged.
        </dd>
      </dl>
      <t>
        T6 and T7 use only bytes the SUT itself wrote. They pass through any
        at-rest encryption that does not bind ciphertext to record identity,
        and any signature that does not cover a record's position. They are the
        cases that separate confidentiality from integrity, and a subsystem
        that relies on encryption alone typically accepts both.
      </t>
      <t>
        A subsystem that stores memory as unordered key-value pairs, with no
        inherent record order and no separate metadata layer, cannot express
        T2, T4, or T8. Such cases are reported as "not applicable" for that
        SUT, with the reason, and are excluded from the denominator in
        <xref target="scoring"/>. They are never reported as passes.
      </t>
    </section>

    <section anchor="verdicts">
      <name>Verdicts</name>
      <t>
        After the restart the evaluator reads the affected context through the
        read path and classifies the outcome as one of three verdicts.
      </t>
      <dl newline="false">
        <dt>REJECTED:</dt>
        <dd>
          the SUT refuses to load or serve the affected memory: an error, an
          empty result, or a refusal to resume.
        </dd>
        <dt>REPORTED:</dt>
        <dd>
          the SUT serves the memory but raises an integrity signal through a
          documented channel, a log line at warning level or above, a callback,
          or a status field, before or at the time of serving.
        </dd>
        <dt>ACCEPTED:</dt>
        <dd>the SUT serves the edited memory as genuine with no signal.</dd>
      </dl>
      <t>
        REJECTED and REPORTED are passes. ACCEPTED is a fail. Each verdict
        carries a detection point. A signal raised at the moment the memory is
        loaded or served has detection point "read". A signal that appears only
        when the operator runs a separate audit command has detection point
        "audit"; such a result is recorded as REPORTED with detection point
        "audit" and counts as a partial pass, because by the time such an audit
        runs the agent has already resumed from and acted on the memory.
      </t>
    </section>

    <section anchor="controls">
      <name>Control Cases</name>
      <t>
        Two control cases MUST hold for a result to be valid.
      </t>
      <dl newline="true">
        <dt>C1 No-op reload.</dt>
        <dd>
          Stop the SUT, restart it with no edit, and read. The SUT MUST serve
          the seeded memory unchanged and MUST NOT raise an integrity signal.
        </dd>
        <dt>C2 Genuine write after restart.</dt>
        <dd>
          Stop, restart, write one more record through the SUT's own write
          path, and read. The record MUST be served.
        </dd>
      </dl>
      <t>
        If either control fails, the metric MUST NOT be computed and the result
        MUST be reported as "not evaluable" with the reason. Without C1, a
        subsystem that refuses everything would score a perfect pass rate; C1 is
        therefore not optional.
      </t>
    </section>

    <section anchor="scoring">
      <name>Scoring</name>
      <t>
        Let N be the number of applicable test cases for the SUT (eight, less
        any cases reported "not applicable" under <xref target="cases"/>). Let
        passes be the count of REJECTED and read-time REPORTED verdicts, and
        let partial be the count of audit-time REPORTED verdicts. The metric is:
      </t>
      <t>
        Metric Value = (passes + 0.5 * partial) / N
      </t>
      <t>
        Evaluators MAY additionally report the pass rate on T1 to T5 and on T6
        to T8 separately, since the first group tests tamper evidence and the
        second tests binding of a record to its place and owner in the store.
      </t>
    </section>

    <section anchor="reporting">
      <name>Reporting</name>
      <t>
        Each result MUST state: the SUT and the exact version of its memory
        component; the store backend and version; the configuration used, that
        is, which encryption, signing, or audit features were on or off; for
        every test case the verdict and the detection point; the exact commands
        or script that produced the edit; and the date of measurement.
      </t>
      <t>
        A result is a statement about one version of one subsystem in one
        configuration. Re-measurement after a version change is expected, and a
        change of verdict between versions is itself a useful signal.
      </t>
    </section>

    <section anchor="illustrative">
      <name>Illustrative Results (Informative)</name>
      <t>
        The following verdicts illustrate what the method produces on current
        releases of four widely used memory subsystems. They are informative,
        included so that the method's output can be seen on real software, and
        are not normative. All runs used the subsystem's default configuration
        unless stated. Versions are given so the results can be reproduced or
        contested.
      </t>
      <ul spacing="normal">
        <li>
          LangGraph SqliteSaver (langgraph-checkpoint-sqlite 3.1.1): T1 to T5
          ACCEPTED. No integrity check on load.
        </li>
        <li>
          LangGraph EncryptedSerializer, AES in EAX mode
          (langgraph-checkpoint 4.2.0): T6 and T7 ACCEPTED. The authenticated
          encryption tag covers the ciphertext but not the record's identity,
          so a genuine encrypted record verifies in any position or context.
        </li>
        <li>
          Letta block checkpoint history (Letta 0.16.8): T1 to T5 ACCEPTED.
        </li>
        <li>
          Mem0 local vector store (Mem0 2.0.20, Qdrant): T1 to T5 ACCEPTED.
        </li>
        <li>
          A signed-receipt store with receipts enabled: T1 to T5 REPORTED,
          detection point "audit". With receipts off, the default: ACCEPTED.
        </li>
      </ul>
      <t>
        The pattern across these subsystems is consistent: the read path trusts
        the store. Where a protection exists, it is either off by default or
        detects only after the agent has already resumed. This is why
        <xref target="verdicts"/> keeps read-time and audit-time detection
        apart.
      </t>
    </section>

    <section anchor="rationale">
      <name>Design Rationale (Informative)</name>
      <t>
        Three choices in this method are worth stating explicitly.
      </t>
      <t>
        First, eight edits rather than one. A single "tamper the file" case
        lets a subsystem with a whole-store checksum pass while it still accepts
        rollback and cross-context replay. T6 and T7 in particular exist
        because encryption alone accepts them: they move only genuine bytes.
      </t>
      <t>
        Second, the verdict comes from the SUT, not from the evaluator.
        Inspecting the store to decide whether a tamper "should" have been
        caught makes the result an opinion. Reading through the SUT's own read
        path and recording what it did makes it a measurement.
      </t>
      <t>
        Third, control C1 is not optional. It is the only thing that prevents a
        subsystem that rejects all reads from scoring a perfect result.
      </t>
    </section>

    <section anchor="reference-impl">
      <name>Reference Implementation (Informative)</name>
      <t>
        An open, MIT-licensed implementation of this method, called agmi (Agent
        Memory Integrity), runs offline and applies the eight edits and control
        C1 to the memory components of several frameworks through a common
        adapter interface. It is provided for reproducibility and is not
        required to use the method.
      </t>
      <ul empty="true" spacing="compact">
        <li>Code: https://github.com/tech4biz-yasha/agmi</li>
        <li>Method paper: https://doi.org/10.5281/zenodo.22765627</li>
      </ul>
    </section>

    <section anchor="security">
      <name>Security Considerations</name>
      <t>
        This document defines a benchmarking method for laboratory use. As with
        other benchmarking methodologies, the procedures here are intended for
        an isolated test bed and MUST NOT be run against production systems or
        systems the evaluator is not authorised to test. The test cases apply
        adversarial edits to a store; those edits MUST be confined to test data
        in a controlled environment.
      </t>
      <t>
        The method measures a security-relevant property, whether an agent
        detects at-rest tampering with its memory, but a metric value from this
        method is not by itself an assurance of security. A high value on this
        method does not measure confidentiality, input-path poisoning
        resistance, or any property outside <xref target="scope"/>.
      </t>
      <t>
        Publishing per-case verdicts for named subsystems and versions
        describes weaknesses in released software. The intent is constructive:
        results are stated with the version measured so that vendors can
        reproduce and fix them, and so that a later re-measurement shows the
        change.
      </t>
    </section>

    <section anchor="iana">
      <name>IANA Considerations</name>
      <t>This document has no IANA actions.</t>
    </section>

  </middle>

  <back>
    <references>
      <name>Normative References</name>
      <reference anchor="RFC2119" target="https://www.rfc-editor.org/info/rfc2119">
        <front>
          <title>Key words for use in RFCs to Indicate Requirement Levels</title>
          <author initials="S." surname="Bradner" fullname="S. Bradner"/>
          <date year="1997" month="March"/>
        </front>
        <seriesInfo name="BCP" value="14"/>
        <seriesInfo name="RFC" value="2119"/>
      </reference>
      <reference anchor="RFC8174" target="https://www.rfc-editor.org/info/rfc8174">
        <front>
          <title>Ambiguity of Uppercase vs Lowercase in RFC 2119 Key Words</title>
          <author initials="B." surname="Leiba" fullname="B. Leiba"/>
          <date year="2017" month="May"/>
        </front>
        <seriesInfo name="BCP" value="14"/>
        <seriesInfo name="RFC" value="8174"/>
      </reference>
    </references>

    <references>
      <name>Informative References</name>
      <reference anchor="RFC2544" target="https://www.rfc-editor.org/info/rfc2544">
        <front>
          <title>Benchmarking Methodology for Network Interconnect Devices</title>
          <author initials="S." surname="Bradner" fullname="S. Bradner"/>
          <author initials="J." surname="McQuaid" fullname="J. McQuaid"/>
          <date year="1999" month="March"/>
        </front>
        <seriesInfo name="RFC" value="2544"/>
      </reference>
      <reference anchor="RFC8239" target="https://www.rfc-editor.org/info/rfc8239">
        <front>
          <title>Data Center Benchmarking Methodology</title>
          <author initials="L." surname="Avramov" fullname="L. Avramov"/>
          <author initials="J." surname="Rapp" fullname="J. Rapp"/>
          <date year="2017" month="August"/>
        </front>
        <seriesInfo name="RFC" value="8239"/>
      </reference>
    </references>

    <section anchor="ack" numbered="false">
      <name>Acknowledgements</name>
      <t>
        The author thanks the maintainers of the memory subsystems measured
        during the development of this method for their engagement on the
        individual findings.
      </t>
    </section>
  </back>
</rfc>
