CATS Y. Mo Internet-Draft D. Yang Intended status: Informational C. Zhou Expires: 6 April 2027 Huazhong University of Science and Technology 06 October 2026 State, Storage, and Compute Affinity for AI Agent Service Selection in Computing-Aware Traffic Steering draft-mo-cats-agent-state-affinity-00 Abstract AI agent services are stateful and long-running: the speed with which a step can be served depends on whether the session's context, retrieved data, tool results, and model-side state such as a key-value (KV) cache are already available at, or near, the selected service contact instance, and on whether the computation that the step needs is ready there. Computing-Aware Traffic Steering (CATS) exposes computing and network metrics and defines service contact instance affinity, but it does not expose the availability of reusable state, does not distinguish a hard locality constraint on state from a preference for reusing it, does not provide a way to compare the cost of moving state with the cost of recomputing it, and does not represent the readiness of a specific computation. This document describes affinity in three coupled dimensions -- state, storage, and compute -- for agent service selection. It states the motivation and the goals of introducing affinity, the requirements and constraints that affinity places on the mapping of an agent workload onto storage and compute resources, the metrics and measurement methods by which affinity is observed, the mechanisms and the procedure by which affinity is assured, and two cases in detail: a long-horizon session that passes through several stages, and a group of similar agents or of tenants served by shared reusable state. This document defines no wire protocol, no encoding, and no data model, and it does not define the transfer mechanisms themselves. Status of This Memo This Internet-Draft is submitted in full conformance with the provisions of BCP 78 and BCP 79. Internet-Drafts are working documents of the Internet Engineering Task Force (IETF). Note that other groups may also distribute working documents as Internet-Drafts. The list of current Internet- Drafts is at https://datatracker.ietf.org/drafts/current/. Mo, et al. Expires 3 April 2027 [Page 1] Internet-Draft Agent Affinity September 2026 Internet-Drafts are draft documents valid for a maximum of six months and may be updated, replaced, or obsoleted by other documents at any time. It is inappropriate to use Internet-Drafts as reference material or to cite them other than as "work in progress." This Internet-Draft will expire on 3 April 2027. Copyright Notice Copyright (c) 2026 IETF Trust and the persons identified as the document authors. All rights reserved. This document is subject to BCP 78 and the IETF Trust's Legal Provisions Relating to IETF Documents (https://trustee.ietf.org/ license-info) in effect on the date of publication of this document. Please review these documents carefully, as they describe your rights and restrictions with respect to this document. Code Components extracted from this document must include Revised BSD License text as described in Section 4.e of the Trust Legal Provisions and are provided without warranty as described in the Revised BSD License. Discussion Venues Discussion of this document takes place on the CATS Working Group mailing list (cats@ietf.org), which is archived at https://mailarchive.ietf.org/arch/browse/cats/. Table of Contents Mo, et al. Expires 3 April 2027 [Page 2] Internet-Draft Agent Affinity September 2026 1. Introduction . . . . . . . . . . . . . . . . . . . . . . . . . 4 2. Conventions and Definitions . . . . . . . . . . . . . . . . . .6 3. Terminology . . . . . . . . . . . . . . . . . . . . . . . . . .6 4. Motivation and Goals . . . . . . . . . . . . . . . . . . . . . 8 4.1. Why Agent State, Storage, and Compute Are Coupled . . . . 8 4.2. Failure Modes without an Affinity View . . . . . . . . . .9 4.3. Goals . . . . . . . . . . . . . . . . . . . . . . . . . .10 4.4. Non-Goals . . . . . . . . . . . . . . . . . . . . . . . .11 5. Problem Statement . . . . . . . . . . . . . . . . . . . . . . 11 5.1. Storage Is Represented Only as Capacity . . . . . . . . .11 5.2. Affinity Is Instance-Level . . . . . . . . . . . . . . . 11 5.3. Transfer and Re-computation Are Not Comparable . . . . . .11 5.4. Compute Affinity Is Not Represented . . . . . . . . . . .11 5.5. Reuse across Similar Agents and Tenants Is Not Represented . . . . . . . . . . . . . . . . . . . . . . . . 12 5.6. The Long Horizon Amplifies Each of These Gaps . . . . . .12 6. Working Set Elements Relevant to Agent Services . . . . . . . 13 6.1. Element Classes . . . . . . . . . . . . . . . . . . . . .13 6.2. Compute-Side Readiness . . . . . . . . . . . . . . . . . 14 6.3. Why the Properties Are Pairwise . . . . . . . . . . . . .14 7. State, Storage, and Compute Affinity . . . . . . . . . . . . .14 7.1. Affinity, Preference, and Locality . . . . . . . . . . . 14 7.2. How the Three Affinities Interact . . . . . . . . . . . .15 7.3. The Granularity of an Affinity Decision . . . . . . . . .16 8. Requirements and Constraints on the Mapping . . . . . . . . . 17 8.1. The Mapping Relation . . . . . . . . . . . . . . . . . . 17 8.2. Constraint Families . . . . . . . . . . . . . . . . . . .17 8.3. Mapping Requirements . . . . . . . . . . . . . . . . . . 18 9. Affinity Information Requirements . . . . . . . . . . . . . . 20 9.1. State Availability . . . . . . . . . . . . . . . . . . . 20 9.2. Residency Tiers and Retrieval Cost . . . . . . . . . . . 20 9.3. Constraints, Consistency, and Sharing . . . . . . . . . .20 9.4. Compute Affinity and Reuse Scope . . . . . . . . . . . . 21 10. State Handles and Their Relationship to CATS Identifiers . . 21 11. Measuring Affinity: Metrics and Methods . . . . . . . . . . .22 11.1. What Has to Be Measured . . . . . . . . . . . . . . . . 23 11.2. Affinity Metric Catalogue . . . . . . . . . . . . . . . 23 11.3. Metric Semantics and Reporting Standards . . . . . . . .24 11.4. Measurement Methods . . . . . . . . . . . . . . . . . . 25 11.5. Measurement Requirements . . . . . . . . . . . . . . . .26 12. Affinity Assurance: Mechanisms and Procedures . . . . . . . .27 12.1. Mechanism Catalogue . . . . . . . . . . . . . . . . . . 27 12.2. Assurance Procedure . . . . . . . . . . . . . . . . . . 28 12.3. Triggers for Re-evaluation . . . . . . . . . . . . . . .29 12.4. Abort, Failure, and Fallback . . . . . . . . . . . . . .30 12.5. Assurance Requirements . . . . . . . . . . . . . . . . .31 13. Affinity Assurance for Long-Horizon and Multi-Stage Sessions . . . . . . . . . . . . . . . . . . . . . . . . . . . 32 13.1. Stage Model . . . . . . . . . . . . . . . . . . . . . . 32 13.2. Phase-Differentiated Assurance . . . . . . . . . . . . .32 13.3. Pinning, Retention, and Tier Budget . . . . . . . . . . 33 Mo, et al. Expires 3 April 2027 [Page 3] Internet-Draft Agent Affinity September 2026 13.4. Checkpoint and Resume after a Pause . . . . . . . . . . 33 13.5. Stage Transitions . . . . . . . . . . . . . . . . . . . 34 13.6. Drift, Decay, and Re-evaluation . . . . . . . . . . . . 34 13.7. Budget Pacing across Stages . . . . . . . . . . . . . . 35 13.8. Failure and Recovery . . . . . . . . . . . . . . . . . .35 13.9. Multi-Agent Stages . . . . . . . . . . . . . . . . . . .36 13.10. Long-Horizon Requirements . . . . . . . . . . . . . . .36 14. Affinity for Similar Agents and for Tenants . . . . . . . . .37 14.1. Similarity and Reuse Groups . . . . . . . . . . . . . . 37 14.2. Elements That Can Be Shared . . . . . . . . . . . . . . 37 14.3. Conditions for Safe Reuse . . . . . . . . . . . . . . . 38 14.4. Tenant Affinity Domains and Isolation . . . . . . . . . 39 14.5. Fairness, Herding, and Quota . . . . . . . . . . . . . .39 14.6. Measurement and Observability across Tenants . . . . . .40 14.7. Requirements for Similar Agents and Tenants . . . . . . 41 15. Interaction with Existing Work . . . . . . . . . . . . . . . 42 15.1. Traceability to the Agent Service Requirements . . . . .43 16. Operational Considerations . . . . . . . . . . . . . . . . . 43 17. Security Considerations . . . . . . . . . . . . . . . . . . .45 18. Privacy Considerations . . . . . . . . . . . . . . . . . . . 45 19. IANA Considerations . . . . . . . . . . . . . . . . . . . . .46 20. Normative References . . . . . . . . . . . . . . . . . . . . 46 21. Informative References . . . . . . . . . . . . . . . . . . . 46 Acknowledgments . . . . . . . . . . . . . . . . . . . . . . . . . 47 Authors' Addresses . . . . . . . . . . . . . . . . . . . . . . . .48 1. Introduction The CATS framework selects a service contact instance using computing and network metrics [I-D.ietf-cats-framework]. Agent services add a consideration that is not represented in that metric set: how much of what a step needs is already in place at the candidate. A step of an agent session frequently can be served in two ways at a candidate instance -- by using state that the instance already holds and by using the computation that is already ready there, or by obtaining that state and that readiness again, either by transferring them from where they are held or by reconstructing them. The difference between the two paths is often larger than the difference between two candidate instances on any currently defined metric. Three dimensions of that consideration are distinguished in this document. State affinity is the preference for a candidate because it already holds the working set elements that a step needs. Storage affinity is the preference for a candidate because the tier at which those elements are held is usable for the step and because the durable store that can serve them is close to it. Compute affinity is the preference for a candidate because the model revision, precision, adapter, runtime, and accelerator that the step needs are resident and ready there, so that no artifact load, quantization change, or capability substitution is required. Mo, et al. Expires 3 April 2027 [Page 4] Internet-Draft Agent Affinity September 2026 The three are coupled, and that coupling is the reason this document treats them together. State can be reused only by a computation that can consume it, so a state that is present at an instance whose runtime cannot consume it has no reuse value at that instance. Changing the instance to gain compute affinity pays for the change in state: the accumulated state of a session is either moved, which costs transfer, or rebuilt, which costs accelerator time. Storage affinity is the bridge between the two, because it determines at which tier an element is held and how cheaply it can be made usable by the computation that needs it. This document states what a CATS system needs to know in order to take the three affinities into account, and what it has to be able to do about them. The document is organized around six questions. Section 4 states the motivation and the goals of introducing affinity. Section 8 states the requirements and the constraints that affinity places on the mapping of an agent workload onto storage and compute resources. Section 11 states how affinity is measured, and which metric properties make two measurements comparable. Section 12 states the mechanisms by which affinity is assured and the procedure that applies them. Section 13 applies that procedure to a long-horizon session that passes through several stages. Section 14 applies it to a group of similar agents and to tenants that share reusable state. The gap that motivates the document has been described for one element of the working set: the KV cache of a large language model, for which it has been observed that the existing CATS metrics do not expose cache state and that the distribution framework does not describe how cached content is distributed or synchronized across instances [I-D.li-cats-kv-cache-distribution]. The same reasoning applies to the other elements of an agent working set, and to the compute side of the same question. This document assumes the single-domain deployment model of the CATS framework. The intended standing of the document is informational groundwork in the sense of the CATS charter [CATS-CHARTER]: it states what a CATS system has to be able to express, so that the work can be taken up by the metric definition [I-D.ietf-cats-metric-definition] and by the data model [I-D.ietf-cats-data-model] rather than by a protocol extension. Three distinctions are introduced here that the current documents do not make. * A distinction between state availability, which is a quantity, and state locality, which is a constraint. Mo, et al. Expires 3 April 2027 [Page 5] Internet-Draft Agent Affinity September 2026 * A distinction between instance affinity, which is about which service contact instance serves a session, and state affinity, which is about where the session's state and its computation are held. * A distinction between the cost of obtaining state by transfer and the cost of obtaining it by re-computation or re-retrieval, and, on the compute side, between load, which is a property of an instance, and readiness, which is a property of the pair of a step requirement and an instance. 2. Conventions and Definitions The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in BCP 14 [RFC2119] [RFC8174] when, and only when, they appear in all capitals, as shown here. In this document these key words are used to state requirements on the design of a CATS system that supports agent services. They do not describe protocol behaviour, and this document defines no protocol, message format, or data model. 3. Terminology This document uses the terms of [I-D.ietf-cats-framework] and [I-D.mo-cats-agent-service-characteristics]. The following additional terms are used. State Handle: An identifier under which a working set element can be addressed. A state handle is not a network address and does not by itself authorize access to the state that it names. State Availability: The presence of a working set element at, or within a reachable tier of, a candidate instance. Mo, et al. Expires 3 April 2027 [Page 6] Internet-Draft Agent Affinity September 2026 State Affinity: The preference for a candidate instance because it already holds, or is able to reach cheaply, the working set elements that a step needs. Storage Affinity: The preference for a candidate instance because the tier at which a working set element is held there is usable for the step, and because the durable store that can serve the element is cheap to reach from it. Compute Affinity: The preference for a candidate instance because the computation that a step requires is ready there: the model revision, the precision, the adapter, the runtime, and the accelerator class that the step requires are resident, so that the step can start without a load, a conversion, or a substitution. State Locality Constraint: A rule that forbids a working set element from being placed, transferred, or reused in a given location or scope. Affinity Domain: The set of instances among which a working set element may be reused, or among which the sessions of a tenant may be placed, under a stated scope. Affinity Gain: The reduction in the cost of serving a step that is attributable to reuse and to readiness, measured against the cost that the same step would have had if the state and the computation had been re-established. Affinity Decay: The reduction of affinity gain over time, caused by eviction, expiry, replacement of a model or adapter, drift of the session's requirement, or growth of the working set beyond the capacity that the holder is willing to commit. Affinity Budget: The share of the cost, latency, or capacity budget of a session that the session may spend on establishing, maintaining, or changing its affinity. Mo, et al. Expires 3 April 2027 [Page 7] Internet-Draft Agent Affinity September 2026 Retention Commitment: A statement by a holder that it will keep a working set element for a stated period or until a stated event, within a stated capacity limit and subject to stated preemption rules. Reuse Group: The set of sessions, steps, agents, or tenants for which a given working set element may legitimately be reused. Similar Agent: An agent whose steps require the same computation or the same reusable elements as another agent, as determined by the conditions of reuse and not by the identity of the agent. Stage Transition: The boundary at which the requirement profile of the next stage becomes known, and at which the placement and the affinity of the session may be re-established. 4. Motivation and Goals 4.1. Why Agent State, Storage, and Compute Are Coupled An agent session consumes three kinds of readiness, and each of them is produced at a cost that the selection decision can either pay or avoid. First, the state of the session. A step that can use the accumulated context of a session avoids re-reading the transcript, re-retrieving the documents, re-invoking the tools, and re-running the pre-fill. The alternative to reuse is not a constant: it depends on the length of the context that has to be re-established, on the throughput of the candidate at that operation, and on whether the material that has to be re-retrieved is available at all. Second, the tier at which that state is held. The same element has different costs depending on whether it is in accelerator memory, in host memory, in a node-local store, or only in a shared store that is reached over the network. The tier determines whether the element is usable as it stands or whether it has to be copied, paged, or deserialized before a step can use it, and it determines whether holding the element competes for the same accelerator capacity that the step itself needs. Third, the computation. A step of an agent session is not served by an arbitrary processor: it needs a model revision, a precision, sometimes an adapter, a runtime that can consume the state that is available, and an accelerator class that can execute it. When those are absent, the step can be served only after an artifact is loaded, a model is converted, or the step is redirected to a different capability tier, and each of those is a cost that is paid before the step begins. Mo, et al. Expires 3 April 2027 [Page 8] Internet-Draft Agent Affinity September 2026 The three are coupled because a state is reusable only by a computation that can consume it, and because a computation that is ready is worth little without the state to run on. A candidate that holds the KV state of a session but runs a different model revision cannot reuse it. A candidate that has the accelerator free but not the artifact must load it first, and the load costs more than the difference between many pairs of candidates. This is why an affinity decision that considers only one of the three dimensions produces a placement that is worse than one made without considering affinity at all: it moves the session to gain a dimension and pays for the gain in the other two. 4.2. Failure Modes without an Affinity View Six failure modes are observed when affinity is not represented in the selection input. * Repeated reconstruction. The same context is prefilled, the same documents are retrieved, and the same tools are invoked again because the selection did not know where the results were already held. * Migration thrash. Because the cost of moving state is not comparable with the benefit of the move, sessions are moved for gains that do not survive the move, and the traffic that the moves generate degrades the conditions that motivated them. * Concentration and its correction. If affinity is applied without a rule that overrides it, the instances that hold popular state absorb the demand, and if affinity is not applied at all, the state is re-established everywhere. Both extremes waste capacity in opposite directions. * Isolation that is assumed rather than enforced. A reuse key that is treated as a hint rather than as a scoped identifier can cause state of one tenant to be reused by another, or a locality rule that was expressed as a cost to be traded to be violated silently. * Loss of accumulated work. A long session that pauses and resumes on an instance that no longer holds its state, or that holds it at a tier the step cannot use, restarts from a checkpoint that may be much older than the session itself. Mo, et al. Expires 3 April 2027 [Page 9] Internet-Draft Agent Affinity September 2026 * Decisions that are not reviewable. When the reason for a placement is not recorded, an operator cannot distinguish a session that was correctly pinned from one that was pinned by the absence of a competitive alternative. 4.3. Goals The goals of introducing state, storage, and compute affinity into CATS are the following. They are stated as goals rather than as requirements; the requirements appear in Sections 8, 9, 11, 12, 13, and 14. * G1. Make readiness visible. A selection function is able to learn, for a candidate and for the elements that a step needs, whether the element is present, at which tier, for how long, and whether the computation that would consume it is ready. * G2. Make the alternatives comparable. The cost of reusing what is present is comparable with the cost of transferring it and with the cost of reconstructing it, on a common basis and for the same step. * G3. Keep constraints out of the ranking. Locality, tenancy, capability, and consistency conditions are evaluated as constraints before any affinity preference is valued, so that a preference is never satisfied by paying for it in a dimension in which the step has a constraint. * G4. Assure affinity rather than assume it. Affinity is established, maintained, verified, and released by stated mechanisms, with a procedure that reserves, commits, and can abort. * G5. Sustain affinity over a horizon. Affinity is maintained across the steps of a long session, across pauses and stage transitions, and is re-evaluated when the assumption behind it changes. * G6. Share without leaking. Reuse across similar agents and across the sessions of one tenant is enabled where it is permitted, and is bounded by scope, isolation, quota, and fairness rules where it is not. Mo, et al. Expires 3 April 2027 [Page 10] Internet-Draft Agent Affinity September 2026 5. Problem Statement 5.1. Storage Is Represented Only as Capacity The CATS metric definition includes storage among the raw metrics that may be collected, in the form of available capacity, read throughput, and write throughput [I-D.ietf-cats-metric-definition]. Those raw metrics describe an instance's storage as a resource that a workload consumes, in the same way that processor utilization or bandwidth describe other resources. They do not describe whether a particular piece of state is available, and they cannot be used to answer the question that matters for agent service selection. 5.2. Affinity Is Instance-Level The framework defines service contact instance affinity, which keeps the traffic of a session on the same instance [I-D.ietf-cats-framework]. That concept is binary with respect to the instance: traffic either stays on the instance or does not. For agent services, affinity has a structure: a session may be able to continue on a new instance without penalty if one element of its working set is transferable and cheap, while it must stay where it is if the element is not transferable or its transfer is prohibited. An instance-level affinity flag cannot express that distinction. 5.3. Transfer and Re-computation Are Not Comparable A CATS system that can see storage capacity but not which state is present cannot compare the two ways of obtaining missing state. Transferring a KV cache of a given size over a path with a given capacity and latency has a cost; recomputing it from the request has a cost that depends on the capability of the instance and the length of the context. The cheaper option determines which candidate is preferable, and the current metric set expresses neither. 5.4. Compute Affinity Is Not Represented Mo, et al. Expires 3 April 2027 [Page 11] Internet-Draft Agent Affinity September 2026 The computing metrics of the framework describe an instance: its capability, its utilization, its queue. A step, however, does not need an instance; it needs a specific computation to be ready. Nothing in the current metric set states whether the artifact of the model revision that the step requires is resident at the candidate, whether the adapter that the session uses is loaded, whether the precision that the step needs is available on the accelerator of that candidate, or whether the runtime at the candidate can consume the state that is available there. The consequence is that a capability requirement is expressed as a filter on an instance rather than as a property of the pair of a requirement and a candidate: a candidate that is capable in general may still be unable to serve the step without a load, and a load that takes longer than the step itself is not visible to a selection that compares only utilization and path quality. The same applies to the state: a candidate whose runtime cannot consume the KV layout in which a prefix is stored holds a state that has no reuse value for that step. 5.5. Reuse across Similar Agents and Tenants Is Not Represented An agent service usually serves many sessions whose steps require the same computation and the same reusable material: the same system prompt prefix, the same tool schemas, the same retrieval corpus, the same model artifact, the same adapter. Because the current information is per instance capacity and load, the fact that a prefix is already resident is invisible to the other sessions that could use it, and the pre-fill of a shared prefix is paid once per session instead of once per reuse group. Prefix-level steering for agent traffic has been proposed on the forwarding side [I-D.zhang-cats-token-aware-ts]; the corresponding reuse side is not represented. The tenant dimension has the same shape. A tenant has a placement domain, a capacity expectation, and an isolation requirement, and its sessions share state that may not be shared with another tenant. Neither the placement domain of a tenant nor the scope of a reusable element is expressible in the current selection input, and neither is the effect of affinity on the fairness between tenants. 5.6. The Long Horizon Amplifies Each of These Gaps Mo, et al. Expires 3 April 2027 [Page 12] Internet-Draft Agent Affinity September 2026 An agent session can run for a long time, with pauses that exceed the lifetime of a cache entry, with stages that have different requirements, and with a working set that grows while it runs. Every one of the gaps above becomes more expensive over a longer horizon: state that is lost at a pause costs a resume from an older checkpoint, a placement that is re-evaluated at every step costs a transfer per step, and a stage transition that is not anticipated costs a cold start in the middle of a task. A selection input that is valid for a single request is therefore not sufficient, and the assurance procedure for a long session is not the same as the one for a single step. 6. Working Set Elements Relevant to Agent Services 6.1. Element Classes The elements below are the ones that agent service selection is expected to take into account. They differ in size, in volatility, in whether they may be shared, and in whether they may move. Element | Volatility | Reuse scope | Locality -------------------+-------------+---------------+----------- Session context | grows per | per session | jurisdiction | turn | | Retrieved data set | per query | per tenant | policy | | | dependent Tool results | short-lived | per session | endpoint | | | bound KV cache or prefix | grows per | per session | accelerator state | token | or per prefix | bound Plan and scratch | per step | per session | jurisdiction state | | | Long-term memory | appended | per tenant or | policy or index | | per user | dependent Model artifact | versioned | per site or | hardware | | per tier | bound Adapter or | versioned | per model or | license quantization | | per tenant | bound artifact | | | Training | per | per job | jurisdiction checkpoint | iteration | | Runtime and kernel | versioned | per site | hardware image | | | bound The last four elements arise in long-running jobs, such as the distributed training use case of [I-D.ietf-cats-usecases-requirements], and are listed here because a long-running session raises the same selection question as an agent session. Mo, et al. Expires 3 April 2027 [Page 13] Internet-Draft Agent Affinity September 2026 The list is not exhaustive, and it is expected that deployments will add elements. What matters for this document is that the properties in the columns are the ones that selection needs, and that they are not properties of the instance alone but of the pair of the element and the instance. The reuse scope column is the basis of Section 14: an element whose reuse scope is wider than a session is material that several sessions, several agents, or one tenant can share, and it is the element class for which the cost of re-establishing the element is paid most often for no reason. 6.2. Compute-Side Readiness The compute side has its own working set, and it is as expensive to re-establish as the state side. It consists of the artifacts and the runtime properties that a step requires: the weights of the model revision that the step needs, the adapter or the quantization artifact that the session uses, the runtime and kernel images that the accelerator requires, and the compiled or cached forms of the operations that the step executes. These elements differ from the state elements in one respect that matters to the selection: they are shared by many sessions, they are versioned rather than volatile, and their re-establishment cost is a load rather than a re-computation. A selection that compares two candidates on the basis of free accelerator capacity alone will prefer the candidate that has the memory free and will then pay the load, which is exactly the cost that a compute affinity view makes visible. 6.3. Why the Properties Are Pairwise Each property in the table above is a property of a pair: the element and the instance. The tuple of the model revision, the tokenizer, and the precision that produced a KV state is part of the identity of that state for the purpose of reuse; the runtime of a candidate determines whether the state that the candidate holds is usable by a step; and the free capacity of a candidate determines whether the element fits. A summary of an instance that is not expressed against the element that a step needs therefore cannot answer the question that selection asks. 7. State, Storage, and Compute Affinity 7.1. Affinity, Preference, and Locality Mo, et al. Expires 3 April 2027 [Page 14] Internet-Draft Agent Affinity September 2026 Affinity is a preference, not a constraint. It states that a candidate is better because the material that a step needs is already in place there, and it may be traded against path quality, load, cost, and budget. A locality constraint, by contrast, is a rule that makes a candidate ineligible when the element is not where the rule permits it to be. The two are frequently confused in designs that express a locality rule as a very large weight in a score, which produces a system that admits a violation of the rule when the score happens to favor it. The three affinities are defined over different objects and have different lifetimes. * State affinity is defined over a working set element and a candidate. It lasts as long as the element is present at that candidate, and it is invalidated by eviction, expiry, or a change of the content of the element. * Storage affinity is defined over a residency tier, a durable store, and a candidate. It lasts as long as the element is held at a tier that the step can use, or as long as the store that can serve the element is reachable at the cost that the decision、 assumed. * Compute affinity is defined over a step requirement and a candidate. It lasts as long as the artifacts and the runtime properties that the requirement names are resident, and it is invalidated by a change of the model revision, the adapter, the precision, or the runtime. 7.2. How the Three Affinities Interact The three dimensions are not additive, and the reason is that reuse requires a compatible consumer. The relations below are stated as the pairwise conditions that a selection has to test. Pairwise conditions at a candidate C for a step of session S: Mo, et al. Expires 3 April 2027 [Page 15] Internet-Draft Agent Affinity September 2026 state_reusable | holds the element AND C can consume it: | same model revision, tokenizer, and runtime state_usable | element is at a tier that the step can use, or | can be promoted into one within the step budget compute_ready | the artifacts, precision, adapter, and | accelerator of the step are resident at C compute_fit | the accelerator has capacity for the state and | for the working set of the step together affinity_gain | cost_without_reuse - cost_with_reuse, where the | two costs are measured for the same step A candidate that satisfies state affinity but not compute affinity holds a state it cannot use. A candidate that satisfies compute affinity but not state affinity must either receive the state or reconstruct it, and the cheaper of those two is the cost that the decision has to compare against the gain of the move. A candidate that satisfies storage affinity without either is close to the material but not ready to use it, and its value is the difference between the transfer time from the store and the transfer time from a distant holder. The interaction also produces an ordering that the rest of this document follows: constraints are tested first (Sections 8 and 9), the cost of not holding the state and of not being compute-ready is derived next (Section 11), and only then is affinity valued, assured, and maintained (Sections 12, 13, and 14). This is the same ordering, and the same division between constraints and quantities, as the selection mapping of [I-D.mo-cats-agent-selection-mapping]. 7.3. The Granularity of an Affinity Decision Affinity is decided per element and per step, not per instance and not per session. The practical consequences are as follows. * A decision may hold for one element and not for another, so the unit of a decision is the element that a step needs. * A decision may be revisited at a step boundary without moving the session, because one element may be re-established where it stands while the rest of the working set stays in place. * A decision that covers several steps has a longer lifetime than a decision that covers one, and the longer lifetime is what makes a retention commitment worth its capacity. * A decision that applies to a group of sessions (Section 14) has a Mo, et al. Expires 3 April 2027 [Page 16] Internet-Draft Agent Affinity September 2026 lifetime that is bounded by the version of the element, not by the lifetime of any one session. 8. Requirements and Constraints on the Mapping 8.1. The Mapping Relation The mapping that this section constrains takes a step of an agent session and produces, for each candidate instance, the constraints that the step imposes and the quantities by which the candidates differ. It is the storage-and-compute part of the mapping described in [I-D.mo-cats-agent-selection-mapping], stated here at the level of detail that affinity requires. The demand side of the mapping is derived from the step and from the session it belongs to. The supply side is observed at the candidate. The result is a pairwise valuation. Mapping element | Item | Content ---------------------+------------------+---------------------- Demand of a step | needs_state | keys, tier, locality, | | age | needs_storage | capacity, tier, store | | locality | needs_compute | revision, precision, | | adapter, accelerator, | | rate Supply at a | holds_state | key, tier, retention candidate | | | offers_storage | capacity, tier, | | locality | offers_compute | readiness, free | | capacity, queue Pairwise result | state_ready | holds the element and | | can consume it | state_cost | min(transfer time, | | recompute time) | affinity_gain | cost without reuse, | | minus the cost with | | reuse Three properties of this mapping are required for affinity to be usable. The mapping is per step, because the requirement of a step is what a candidate is evaluated against. The mapping distinguishes constraints from quantities, because a constraint cannot be traded. And the mapping is recomputable from information that the candidate can report without disclosing the content of the state, because otherwise affinity would be expressible only inside a single trust domain. 8.2. Constraint Families Mo, et al. Expires 3 April 2027 [Page 17] Internet-Draft Agent Affinity September 2026 The constraints that affinity imposes on the mapping belong to seven families. Each of them is evaluated before candidates are ranked, and each of them can make a candidate ineligible. * Capacity: the element and the working set of the step fit in the tier that the step requires at the candidate. * Capability: the model revision, precision, adapter, and accelerator class that the step requires are available at the candidate. * Compatibility: the candidate can consume the state that it holds or receives, which requires a matching model, tokenizer, runtime, and layout. * Locality: the element and the execution remain within the jurisdiction, tenant domain, or device boundary that the element carries. * Consistency: the element is current with respect to the version that the step assumes, and is not under an eviction or a replacement that would make it unusable during the step. * Isolation: reuse does not cross a session, user, or tenant boundary that policy protects, and the timing of a hit does not disclose the presence of another session's state. * Temporal: the element remains held, and the computation remains ready, for the duration that the decision assumes, which for a long-horizon session is a commitment that exceeds the step. A budget condition is not a constraint of this kind: exceeding a budget reduces the value of a candidate rather than making it ineligible, and it is therefore expressed as a quantity, except where a session has a hard budget and a step that cannot be served within it must be failed rather than degraded. 8.3. Mapping Requirements Mo, et al. Expires 3 April 2027 [Page 18] Internet-Draft Agent Affinity September 2026 The requirements below are stated using the conventions of BCP 14 [RFC2119] [RFC8174]. They are requirements on the information and on the mapping that a CATS system needs in order to consider affinity, and are not protocol requirements. WM1. A CATS system SHOULD be able to express the demand of a step as a requirement on each of the state, storage, and compute dimensions, rather than as a single scalar. WM2. A CATS system MUST distinguish, in that expression, a constraint that makes a candidate ineligible from a preference that may be traded against other quantities. WM3. A CATS system SHOULD be able to express a state requirement as a set of required elements, each named by a state handle, with an indication of whether the element is required to be present at a stated tier or may be made present at a stated cost. WM4. A CATS system SHOULD be able to express a storage requirement as a capacity and a tier requirement, together with a locality constraint on the durable store that holds the element. WM5. A CATS system SHOULD be able to express a compute requirement as a capability requirement, comprising the model revision, the precision, the accelerator class, the runtime, and any adapter, together with the rate at which the step needs to consume. WM6. A CATS system MUST evaluate capacity, capability, compatibility, locality, consistency, isolation, and temporal constraints before it values any affinity preference, and MUST NOT satisfy a constraint by paying for it in another dimension. WM7. A CATS system SHOULD be able to express the mapping at the granularity of a step, and SHOULD be able to carry the part of a session's mapping that is unchanged from one step to the next without re-deriving it. WM8. A CATS system SHOULD be able to express the residual requirement of a step that cannot be served in full, so that a partial reuse, a partial result, or a degraded execution is expressed as a reduction of the requirement rather than as a silent failure. WM9. A CATS system SHOULD be able to state, for each element of the mapping, the unit in which it is expressed, the component that observes it, and the bound within which it remains valid. WM10. A CATS system SHOULD be able to compare the mapping across candidates of different capability tiers without assuming that the tiers are interchangeable, and SHOULD preserve the reason when a candidate is excluded by a constraint. Mo, et al. Expires 3 April 2027 [Page 19] Internet-Draft Agent Affinity September 2026 9. Affinity Information Requirements The requirements below are stated using the conventions of BCP 14 [RFC2119] [RFC8174]. They are requirements on the information that a CATS system needs in order to consider state, storage, and compute affinity, and are not protocol requirements. 9.1. State Availability S1. A CATS system SHOULD be able to determine, for a candidate instance, whether a working set element is available, and under which state handle it can be addressed. S2. State availability information SHOULD be expressed in a form that can be aggregated, so that an instance is not required to advertise every element that it holds. S3. State availability information SHOULD carry an indication of its freshness, so that a selection function can distinguish a recent observation from a stale one. S4. State availability SHOULD be expressed at the granularity at which reuse is meaningful, which for a KV cache means at the level of a reusable prefix or block rather than at the level of an entire instance. 9.2. Residency Tiers and Retrieval Cost S5. A CATS system SHOULD be able to distinguish the tier at which a working set element resides, at least between accelerator memory, host memory, node-local storage, and a shared store. S6. A CATS system SHOULD be able to represent the cost of making an element usable at a candidate instance, including both the cost of transferring it and the cost of recomputing or re-retrieving it. S7. A CATS system SHOULD be able to identify the cheaper of transfer and re-computation for a given element and candidate, so that the selection function can prefer the candidate that yields the lower total cost. 9.3. Constraints, Consistency, and Sharing S8. A CATS system MUST be able to express state locality constraints as constraints that are evaluated before candidates are ranked. S9. A CATS system SHOULD be able to distinguish state that may be reused across sessions, across tenants, or not at all, and SHOULD NOT require a candidate to disclose state that may not be shared. Mo, et al. Expires 3 April 2027 [Page 20] Internet-Draft Agent Affinity September 2026 S10. A CATS system SHOULD be able to represent the consistency conditions under which a stored element may be reused, including at least whether it is current and whether an eviction is pending. S11. A CATS system SHOULD be able to represent the cost and the conditions of changing the instance that serves a session, so that affinity is applied only when the migration is actually cheaper than the alternative. S12. A CATS system SHOULD NOT require the network to learn the content of a working set element in order to select an instance for it. 9.4. Compute Affinity and Reuse Scope S13. A CATS system SHOULD be able to determine, for a candidate, whether the computation that a step requires is ready there without being re-established: the model revision, the precision, the adapter, the runtime, and the accelerator class that the step needs. S14. A CATS system SHOULD be able to distinguish compute readiness, which is a property of the pair of a step requirement and a candidate, from computing load, which is a property of the candidate alone. S15. A CATS system SHOULD be able to represent the affinity that a session, an agent, or a tenant has accumulated with an instance in a form that survives the end of a turn and the interval between turns. S16. A CATS system SHOULD be able to represent the reuse scope of an element, that is, the set of sessions, agents, or tenants for which the element may be reused, without requiring the content of the element to be disclosed. 10. State Handles and Their Relationship to CATS Identifiers Selection requires that a candidate can state which elements it holds, and that a selection function can compare that statement with what a step needs. This requires an addressing scheme for working set elements. The following properties are proposed for such a handle. * A state handle identifies a working set element, not a location. It does not replace the CATS Service Identifier or the CATS Service Contact Instance ID, and it is not a routable address. * A state handle is scoped. It is meaningful between the parties that use it, and it carries or implies the scope within which it may be used, such as a session, a tenant, a reuse group, or a site. Mo, et al. Expires 3 April 2027 [Page 21] Internet-Draft Agent Affinity September 2026 * A state handle is subject to authorization. Knowledge of a handle is not authorization to read, transfer, or reuse the state that it names. Authorization is expected to be provided by the identity and authorization mechanisms of the environment in which the agent service runs. * A state handle is revocable. Eviction, expiry, a change of the reuse scope, or a policy change can invalidate a handle, and the selection function is expected to tolerate a handle that has become invalid. * A state handle is comparable only within a defined equivalence. Two handles denote the same reusable state only if the deployment defines the equivalence, for example equality of a content hash over a defined model, tokenizer, precision, and prefix. The last property is the one that the compute and sharing dimensions make sharper. For the reuse of a model-side state, the equivalence has to cover the model revision, the tokenizer, the precision, and the runtime layout, because a KV state that was produced by a different combination is not reusable by the step even when its content is identical. This is the same condition that a cache applies when it decides whether a stored representation may be reused for a request [RFC9111]: the stored material is reusable only if the key and the validators of the current request match. The parallel is stated here because it shows that the condition is a property of the pair of the material and the consumer, and not a property of the material alone. The granularity of a state handle is the granularity at which reuse is meaningful, as required by S4, and it is the key under which the storage dimension of agent service selection reports availability [I-D.mo-cats-agent-selection-mapping]. The framework does not define the syntax of a state handle, and this document does not propose one. A syntax is needed only if a protocol is later defined to exchange state availability, at which point the syntax becomes the subject of the document that defines that exchange. 11. Measuring Affinity: Metrics and Methods Mo, et al. Expires 3 April 2027 [Page 22] Internet-Draft Agent Affinity September 2026 11.1. What Has to Be Measured Affinity is an expectation about the cost of the next step, and it can be verified only by measuring what the step actually cost and what it would have cost without reuse. Four questions therefore have to be answerable from measurement. * Is the material there? Presence and residency have to be observable per reuse key, and not only as an aggregate capacity, because the value of a candidate depends on the specific element that a step needs. * Is the computation ready? Readiness has to be observable as a property of the pair of a requirement and a candidate, which means that a measurement of utilization or of installed capability is not a substitute for it. * What did the reuse save? Affinity gain has to be measured as a difference between two costs of the same step, with the basis of the comparison stated, because a gain measured against a different baseline is not comparable with a gain measured elsewhere. * What did affinity cost? The transfer that a decision caused, the capacity that a retention commitment withheld, and the degradation that the concentration of sessions caused are costs of affinity and need to be attributed to the decision that incurred them. 11.2. Affinity Metric Catalogue The metrics below are the ones that the four questions require. They are stated at the level of the quantity and its observation point, not as an encoding, and the levels follow the raw and derived distinction of [I-D.ietf-cats-metric-definition]. Mo, et al. Expires 3 April 2027 [Page 23] Internet-Draft Agent Affinity September 2026 Affinity metric | Unit | Level | Observed at -----------------------+-----------+--------+---------------- Element present under | boolean | raw | C-SMA, per key a reuse key | | | Residency tier of an | tier | raw | C-SMA, per key element | | | Retention remaining | seconds | raw | C-SMA, per key Compute readiness of a | boolean | raw | C-SMA, per pair step requirement | | | Reuse hit ratio | fraction | derived | C-PS, per key | | | and window Reuse distance of a | steps, | derived | C-PS, per session | seconds | | session Transfer volume caused | bytes | raw | C-NMA, per by a decision | | | decision Transfer time caused | seconds | derived | C-NMA, per by a decision | | | decision Recompute volume | tokens, | raw | C-SMA, per step (re-prefill, reload) | bytes | | State cost of a step | seconds, | derived | C-PS, per pair at a candidate | cost | | Affinity gain of a | seconds, | derived | C-PS, per step step | cost | | Selections changed for | count, | derived | C-PS, per lack of state | reason | | window Affinity concentration | fraction | derived | C-PS, per site on an instance | | | Tenant affinity | bytes, | derived | C-SMA, per footprint | fraction | | tenant Reuse refused by scope | count, | derived | C-SMA, per or policy | reason | | window The metrics are grouped by the question they answer. Presence, residency, retention, and readiness answer the first two questions. Reuse hit ratio, reuse distance, transfer volume and time, recompute volume, and state cost answer the third. Affinity gain answers the third and the fourth together, because it states the saving rather than the cost. Concentration, tenant footprint, and refused reuse answer the fourth question and are also the inputs of the fairness rules of Section 14. 11.3. Metric Semantics and Reporting Standards A quantity that two components report under the same name is comparable only if the properties below are fixed. They are the reporting standards that this document asks of a metric definition, and they follow the treatment of freshness and of unknown values in [I-D.zhu-cats-metric-semantics]. * Unit and basis. Each metric names the unit in which it is expressed and the population over which it is computed. A hit ratio without its window and its population is not comparable with another hit ratio. Mo, et al. Expires 3 April 2027 [Page 24] Internet-Draft Agent Affinity September 2026 * Level. Each metric is stated as raw, observed at a component, or as derived, computed from raw values. Affinity gain, state cost, and reuse distance are derived; presence, residency, retention, readiness, and transfer volume are raw. * Observation point. Each metric names the component that observes it and the granularity at which it is observed, which for the affinity metrics is the reuse key, the pair of a requirement and a candidate, the session, or the site. * Freshness. Each reported value carries the time at which it was observed and the bound within which it may be used, and a value outside that bound is treated as unknown. * Percentile and tail. Quantities that describe a distribution are reported at a stated percentile over a stated window, and a quantity that is a mean is identified as a mean, so that a tail objective is not evaluated against an average. * Unknown. A value that is absent, stale, or of unknown provenance is reported as unknown rather than as a default, and an unknown value does not satisfy a constraint. * Aggregation. Affinity metrics are aggregated per class of element, per key, per site, and per tenant. Aggregation reduces cardinality but must not merge populations whose objectives differ, and it must not turn a per-tenant quantity into a value that discloses another tenant. 11.4. Measurement Methods The quantities above are obtainable by four methods, and the choice among them is a deployment decision. * Observation at the holder: the instance that holds an element reports presence, tier, retention, and the outcome of the reuse of the Mo, et al. Expires 3 April 2027 [Page 25] Internet-Draft Agent Affinity September 2026 element. This is the most accurate method for the state side and the only one that can see eviction. * Observation at the decision point: the component that selects records the cost of the step for which a decision was taken, the alternative cost that the same step faced at the candidates that were rejected, and the reason for the choice. This is the only method that can produce affinity gain, because the gain is a difference between two costs of the same step. * Probing: a periodic or on-demand test of the presence of a key, used where the holder does not report, and used after a pause to re-verify the state on which a session depends. Probing has to be rate limited, because a probe is an observable event and because a probe storm competes with the traffic it measures. * Accounting of transfers: the transfer volume and time that a decision caused, separated from background traffic, so that the cost of affinity maintenance is visible. This is required for the operational rule that bounds transfer as a fraction of decisions. Measurement itself is subject to the constraints of the framework: a measurement that can say which session produced a byte has to be protected as session information, and multi-tenant measurement has to be aggregated so that it does not become an information channel (Section 18). The OAM functions of [I-D.ietf-cats-oam-fw] are the natural carrier of these quantities. 11.5. Measurement Requirements MA1. A CATS system SHOULD define each affinity metric with its unit, its observation point, and its level, so that two components that report the same metric report a comparable quantity. MA2. A CATS system SHOULD measure the presence and the residency of a working set element per reuse key, and SHOULD NOT require an instance to enumerate its keys to the network. MA3. A CATS system SHOULD measure the hit ratio of state reuse per key and per session over a stated window, and SHOULD report it together with the window and the population over which it was computed. Mo, et al. Expires 3 April 2027 [Page 26] Internet-Draft Agent Affinity September 2026 MA4. A CATS system SHOULD measure the transfer that a decision causes and the recomputation that a decision causes separately, so that the two ways of obtaining missing state can be compared. MA5. A CATS system SHOULD measure affinity gain as the difference between the cost of a step that reused state or ready compute and the cost that the same step would have had without that reuse, and SHOULD state the basis of the comparison. MA6. A CATS system SHOULD report latency and cost quantities at a stated percentile over a stated window, and MUST NOT report a quantity measured over a population as if it were a property of a single session. MA7. A CATS system SHOULD carry the observation time and the validity bound of each measurement, and SHOULD NOT use a measurement outside that bound to satisfy a constraint. MA8. A CATS system SHOULD aggregate affinity measurements per class of element, per site, and per tenant, and MUST NOT expose a per-session affinity measurement to a party that is not authorized to observe that session. MA9. A CATS system SHOULD distinguish, in its measurements, a miss that is caused by eviction from a miss that is caused by a scope or policy prohibition, because the two have different remedies. MA10. A CATS system SHOULD measure the effect of affinity on the distribution of load, including the concentration of sessions on the instances that hold popular state and the effect of an affinity-driven selection on the tail of other sessions. 12. Affinity Assurance: Mechanisms and Procedures 12.1. Mechanism Catalogue Affinity is assured by a small set of mechanisms. Each mechanism has an effect and a cost, and the choice among them is what turns the information of the previous sections into a placement that is stable and fair. Mo, et al. Expires 3 April 2027 [Page 27] Internet-Draft Agent Affinity September 2026 Mechanism | Effect | Cost or risk -------------------------+---------------------+-------------- Placement pinning | keeps a session | load | with the holder of | concentration | its state | Retention lease | a holder commits to | capacity | keep an element | withheld Replication or warm copy | shortens the path | refresh cost | to the state | Prefetch before a stage | state is ready | transfer if | before the first | unused | step | Recompute in place | avoids a transfer | accelerator | entirely | time Handoff with reserve and | moves a session | two-instance commit | without a gap | overlap Eviction protection | keeps a hot or | capacity for | shared element | others | resident | Admission control | bounds transfers in | queued | flight | sessions Quota or reservation | bounds per-session | under-use | or per-tenant use | Load-aware override | prevents herding on | affinity gain | one holder | lost Degradation ladder | keeps a session | worse | alive under | objective | pressure | Release and eviction | reclaims state that | early loss of policy | has no horizon | reuse The mechanisms are complementary and are combined rather than chosen between. A retention lease without a release policy exhausts the capacity of the holder; a release policy without a lease makes the state of a long session disappear during a pause; an override without a measurement of concentration cannot be applied at the right moment. 12.2. Assurance Procedure The procedure below is the loop that applies the mechanisms. It is stated as a sequence of steps with the information that each step consumes and the decision that it produces, and it is applied at each decision point of a session. Mo, et al. Expires 3 April 2027 [Page 28] Internet-Draft Agent Affinity September 2026 Step | Input | Output -------+--------------------------+--------------------------- A1 | presence, tier, | per-candidate supply view | retention, readiness, | | load | A2 | supply view, constraints | eligible candidate set A3 | eligible set, size, | state cost per candidate | rate, locality | A4 | gain, path, load, | preferred candidate | budget, concentration | A5 | preferred candidate, | commitment or reservation | next steps | A6 | missing elements | transfer, prefetch, or | | reconstruction A7 | committed state and | verified placement and | compute | record A8 | observed gain, decay, | next decision point | triggers | Step A2 applies the constraints of Section 8.2 and excludes the candidates that break one of them, before any affinity is valued. Step A3 derives the cost of making the working set usable, which is the cheaper of the transfer and the reconstruction of each missing element. Step A4 values affinity as the reduction of that cost and trades it against path quality, load, budget, and the concentration that the choice would create. Steps A5 and A6 act, and step A7 verifies rather than assumes that the material is in place, because a commitment can expire between the decision and the step. Step A8 closes the loop and is the reason the procedure is not a one-shot placement: the observed gain and the decay of the affinity are the inputs of the next decision. 12.3. Triggers for Re-evaluation The procedure is re-entered when one of the following events is observed. * Session admission, at which the first decision is taken. * A step boundary, at which the requirement of the next step is known and may differ from that of the previous one. * A stage transition of a long-horizon session (Section 13). * A pause that exceeds the freshness bound of the presence information that the current placement relies on. Mo, et al. Expires 3 April 2027 [Page 29] Internet-Draft Agent Affinity September 2026 * An eviction notice, an expiry of a retention commitment, or a replacement of a model revision, an adapter, or a runtime that the session depends on. * A capacity pressure or a rejection of an admission request, which indicates that the capacity that the decision assumed is no longer available. * A measured breach of the objective of the session, or a measured concentration of sessions on an instance that degrades the objective of other sessions. * A change of the locality, tenancy, or reuse-scope policy that governs an element of the working set. 12.4. Abort, Failure, and Fallback A move that has started may fail at any of its steps, and the failure semantics matter more for affinity than they do for a stateless decision, because a partially moved working set is worse than either endpoint. The following properties are expected of the procedure. * The source of a move remains usable until the target has verified that it holds the material that the session needs, so that a failure leaves the session where it was rather than in neither location. * The verification that a step performs before it runs is the point at which a failed commitment is detected, and the detection leads to the degradation classes of the step rather than to an unbounded retry. * A move that cannot be completed within the affinity budget of the session is abandoned and the session continues where it is, with the element re-established locally if that is cheaper than the move. * When no candidate satisfies the constraints, the session degrades in the order of the fallback ladder of the selection mapping rather than failing silently, which is consistent with the notion of a fallback decision [I-D.pang-cats-fallback-decision-framework]. Mo, et al. Expires 3 April 2027 [Page 30] Internet-Draft Agent Affinity September 2026 12.5. Assurance Requirements AM1. A CATS system SHOULD apply affinity as a preference that is valued after the constraints and alongside load, with a stated weight or order, and SHOULD NOT allow it to override a constraint. AM2. A CATS system SHOULD be able to hold a retention commitment for a working set element for a stated period or until a stated event, and SHOULD be able to release that commitment before its end. AM3. A CATS system SHOULD be able to establish the state and the compute readiness that the next steps need before those steps start, rather than only at the moment at which a step is served. AM4. A CATS system SHOULD be able to choose reconstruction in place as an alternative to a transfer, when the transfer would be larger, slower, or prohibited. AM5. A CATS system SHOULD verify, before a step is served, that the state and the compute readiness that the decision assumed are still present and usable, and SHOULD fall back when they are not. AM6. A CATS system SHOULD bound the fraction of decisions that may trigger a transfer, and SHOULD attribute the transfer that a decision causes to the session that the decision serves. AM7. A CATS system SHOULD override affinity when the concentration of sessions on the instances that hold popular state would degrade the objective of other sessions, and SHOULD make that override visible in the selection context. AM8. A CATS system SHOULD release state that is no longer eligible for reuse, and SHOULD NOT hold state for a session whose horizon has ended. AM9. A CATS system SHOULD re-evaluate an affinity decision when an assumption behind it changes, including an eviction, an expiry, a change of model revision or adapter, a change of policy, or a sustained deviation of the observed gain from the expected gain. AM10. A CATS system SHOULD be able to abandon a move after it has started, and SHOULD leave both the source and the target in a usable state when it does so. AM11. A CATS system SHOULD apply the same assurance procedure to a change of residency tier within an instance as to a change of instance, because both change the cost of the next step. AM12. A CATS system SHOULD retain the decision and its outcome for the steps that follow, so that the affinity of a session is not re-derived from scratch at every step. Mo, et al. Expires 3 April 2027 [Page 31] Internet-Draft Agent Affinity September 2026 13. Affinity Assurance for Long-Horizon and Multi-Stage Sessions 13.1. Stage Model A long-horizon agent session does not have one requirement profile; it has a sequence of them. A session that plans, then retrieves, then analyzes, then calls a tool, and then reports has stages whose working sets overlap in part, whose compute requirements differ, and whose tolerable latency differs. The stage is the unit at which the assurance procedure of Section 12 can be applied without paying for a change at every step. Stage | Working set that grows | Compute profile ---------------+--------------------------+------------------- Plan | task state, plan | capable model Retrieve | retrieved set, index | embedding, search Analyze | context, intermediate | long context Act | tool results, arguments | low-latency calls Report | assembled result | capable model The model is a description, not a specification: a deployment defines its own stages. What matters for affinity is that the boundary between two stages is the point at which the requirement profile becomes known in advance, and therefore the point at which affinity can be re-established proactively rather than reactively. 13.2. Phase-Differentiated Assurance The assurance procedure has a different emphasis in each phase of a long-horizon session. * Admission: the session is placed on the basis of its first stage, and the elements of the later stages that are already known are prefetched where that is cheap. * Steady state within a stage: the placement is kept stable, the retention commitments are renewed as they approach their end, and the affinity gain is measured rather than assumed. * Pause: the retention commitment is what protects the session, and its length is chosen by comparing the cost of holding the state with the cost of re-establishing it at resume. Mo, et al. Expires 3 April 2027 [Page 32] Internet-Draft Agent Affinity September 2026 * Resume: the presence of the elements on which the session depends is re-verified, and an element that is no longer present is treated as absent, so that the resume decision is taken on observed rather than on remembered state. * Stage transition: the affinity of the next stage is established proactively, and the placement of the session is changed at that boundary if the next stage is better served elsewhere. * Completion: the state is released, and the elements whose reuse scope is wider than the session (Section 14) are kept where the reuse group benefits from them. 13.3. Pinning, Retention, and Tier Budget The capacity that a session may keep warm is finite, and the decision of what to keep is a comparison of the cost of holding against the cost of re-establishing. Four rules make that decision tractable. * Keep the elements whose reconstruction is expensive and whose reuse is certain, which for a long session is the accumulated context and the model-side prefix state. * Keep the elements whose reconstruction is impossible, such as the result of a tool invocation that cannot be repeated safely or a retrieval that is no longer reachable. * Do not keep an element whose reconstruction is cheaper than its retention, which for a small derived artifact is usually the case. * Place the elements that are kept at the cheapest tier from which the step can use them, and promote an element to a faster tier only for the stage that needs it, because promotion competes with the state of the steps that are running. 13.4. Checkpoint and Resume after a Pause Mo, et al. Expires 3 April 2027 [Page 33] Internet-Draft Agent Affinity September 2026 A pause longer than the retention commitment of the holder, a failure of the holder, or an administrative move makes the state of a session unavailable. Three properties make the session resumable. * A checkpoint of the durable part of the working set, taken at a stage boundary, bounds the work that a resume can lose. * The part of the working set that is sufficient to continue is stated, so that a resume does not attempt to re-establish the whole set before the session can make progress. * The resume is a decision point like any other: the presence of the elements is verified, the cost of the resume is compared across candidates, and the placement that was in force before the pause is not assumed to be valid. 13.5. Stage Transitions A stage transition is the natural point at which a placement may change, and a change elsewhere is what the stability rule of the selection mapping is intended to prevent. Three properties apply. * The requirement profile of the next stage is known at the boundary, so the constraint test of the next stage can be performed before the stage starts, and a candidate that will fail a constraint can be excluded without waiting for the failure. * The elements that the next stage needs can be prefetched or reconstituted during the last part of the current stage, in parallel with work that is still running, so that the transition does not begin with a serial load. * The cost of the transition is attributed to the stage that benefits from it, and a transition that serves several remaining stages is charged to the session rather than to one step. 13.6. Drift, Decay, and Re-evaluation Mo, et al. Expires 3 April 2027 [Page 34] Internet-Draft Agent Affinity September 2026 Affinity decays, and the decay is gradual rather than binary. The working set grows with the session, so the element that fits at admission may not fit later. The context that the session uses may become less similar to the prefix that is cached. The load of the holder may grow until the affinity gain no longer compensates for it. A session whose placement was correct at step ten can be wrong at step thirty without any single event having changed. Two quantities make the decay observable: the observed affinity gain, compared with the gain that the placement was expected to deliver, and the reuse hit ratio of the session over a recent window. A sustained reduction of either, beyond the tolerance of the session, is the trigger for re-evaluation. Re-evaluation is not the same as movement: it may conclude that re-establishing one element at the current instance is cheaper than moving the session, which for a large working set is usually the case. 13.7. Budget Pacing across Stages A long-horizon session has a budget, and affinity maintenance spends it. The cost of a stage transition, the capacity that a retention commitment withholds, and the traffic of a prefetch are all affordable in isolation and not affordable in a loop. The procedure therefore paces affinity maintenance over the remaining horizon: at each stage boundary, the budget that the remaining stages need is reserved before the affinity of the current stage is extended, and the transfer that a decision causes is charged to the session budget that the decision serves. 13.8. Failure and Recovery The failure modes that matter for affinity over a horizon are the loss of the holder, the loss of the material, and the loss of the placement's value. Each has a recovery path, and the path that was taken is part of the decision record. * Loss of the holder: the session is resumed from a checkpoint, on an instance chosen by the same procedure, with the elements that survived reused where they are reachable. * Loss of the material: the element is reconstructed if that is cheaper than re-establishing it elsewhere, and the reconstruction is measured as a recompute event so that the failure is visible in the accounting. * Loss of the value: the placement no longer delivers a gain, and the session is moved at the next stage boundary rather than immediately, Mo, et al. Expires 3 April 2027 [Page 35] Internet-Draft Agent Affinity September 2026 unless a constraint makes the current placement ineligible. 13.9. Multi-Agent Stages A stage may be served by several agents that share a working set: a planner and a set of workers, or a set of agents that read the same retrieved material. Affinity for such a stage is a property of the shared element as well as of the session, and the placement of the agents of one stage is decided together where the shared element is large, because duplicating it across instances costs as much as the transfer that the shared placement avoids. 13.10. Long-Horizon Requirements LH1. A CATS system SHOULD represent a long-horizon session as a sequence of stages, and SHOULD be able to state, for each stage, the elements and the compute capabilities that the stage needs. LH2. A CATS system SHOULD keep the placement of a session stable across the steps of a stage, and SHOULD change it at a stage boundary rather than within a stage, unless a failure or a violated constraint forces the change. LH3. A CATS system SHOULD maintain the affinity of a session across a pause, and SHOULD carry a retention commitment that covers the expected idle interval where the cost of losing the state exceeds the cost of holding it. LH4. A CATS system MUST re-verify the presence of the elements that a session depends on when the session resumes, and MUST treat an element that is no longer present as absent. LH5. A CATS system SHOULD be able to checkpoint the durable part of a working set, so that a session can be resumed on another instance after a failure, an eviction, or an administrative move. LH6. A CATS system SHOULD state the part of a working set that is sufficient to continue a session, so that a resume reconstructs only the material that the session needs. LH7. A CATS system SHOULD be able to establish the affinity of a stage before the stage starts, including the prefetch of the artifacts and the elements that the next stage needs. LH8. A CATS system SHOULD detect affinity decay, that is, a sustained reduction of the observed gain of the current placement, and SHOULD re-evaluate the placement when that reduction exceeds the tolerance of the session. Mo, et al. Expires 3 April 2027 [Page 36] Internet-Draft Agent Affinity September 2026 LH9. A CATS system SHOULD pace the cost of affinity maintenance over the remaining horizon of a session, and SHOULD NOT spend the budget of the remaining stages to preserve the state of the current one. LH10. A CATS system SHOULD attribute the cost of a stage transition to the stage that benefits from it, and SHOULD attribute a transfer that serves several remaining stages to the session rather than to a single step. LH11. A CATS system SHOULD be able to recover the affinity of a session after the failure of the instance that holds its state, by falling back to a checkpoint or to a reconstruction path, and SHOULD record which path was taken. LH12. A CATS system SHOULD decide the placement of a multi-agent stage that shares a working set as one decision, and SHOULD account for the duplication when the agents of a stage are placed on different instances. 14. Affinity for Similar Agents and for Tenants 14.1. Similarity and Reuse Groups Two agents are similar for the purpose of this document when their steps require the same reusable elements, and not when they are the same software or serve the same user. The reusable elements of Section 6 are frequently shared: a system prompt prefix is common to every session of an agent service, a tool schema is common to every agent that uses the tool, a retrieval corpus is common to a tenant, and a model artifact is common to every step of every session that runs on it. Affinity for a group has a different economics from affinity for a single session. The cost of establishing an element is paid once and is amortized over the reuse group, so an element whose re-establishment is expensive is worth keeping even when the individual session that caused it has ended. The risk is also different: the group is what makes a position on one instance popular, and it is what makes the isolation and fairness rules of the following subsections necessary. 14.2. Elements That Can Be Shared Mo, et al. Expires 3 April 2027 [Page 37] Internet-Draft Agent Affinity September 2026 Reusable element | Reuse group | Precondition -----------------------+---------------------+-------------- System prompt prefix | agents of one agent | same model state | service | and tokenizer Tool schema and | agents of one tool | same template templates | set | version Retrieved corpus and | tenant or user | same rights index | | and version Embeddings of shared | tenant | same documents | | embedding | | model Model weights and | site or capability | same adapter | tier | precision and | | revision Policy and evaluation | agent service | same policy state | | version Tenant memory and plan | tenant only | no state | | cross-tenant | | scope The conditions of reuse in the third column are what make a group a reuse group: an element may be reused by the members of the group only when the conditions hold at the candidate that holds it. A prefix state produced by one model revision is not reusable by a step that runs another, even when the prompt text is identical, because the state is not a copy of the prompt. 14.3. Conditions for Safe Reuse Shared reuse is safe when the following conditions are met, and each of them is a condition that the holder evaluates rather than one that the requester asserts. * Identity of the computation: the model revision, tokenizer, precision, runtime, and adapter that produced the element match those that the reusing step requires. * Identity of the material: the element is the same version of the same material, established by a defined equivalence (Section 10) rather than by a name that two producers may use for different content. * Authorization: the reuse group of the element includes the session that requests the reuse, and the holder enforces that membership. Mo, et al. Expires 3 April 2027 [Page 38] Internet-Draft Agent Affinity September 2026 * Isolation: the reuse does not disclose the content of one session to another, and does not disclose the presence of one session's state to another through the timing of a hit. * Currency: the element has not been superseded by a version that the reusing step assumes, and is not under a pending eviction or replacement that would make it unusable during the step. The fourth condition is the one that is most easily lost in design, because sharing is implemented for efficiency and its observability consequences are considered later. A shared hit that is faster than a miss is an observable signal, and in a multi-tenant system it is a signal about another tenant's activity if the sharing is not bounded. 14.4. Tenant Affinity Domains and Isolation A tenant is the unit at which placement policy, capacity expectation, and isolation are usually expressed. Three properties apply to the affinity of a tenant. * A tenant has an affinity domain, that is, the set of instances at which its sessions and its state may be placed. The domain is a constraint: a candidate outside the domain is ineligible for that tenant's sessions and for the state that belongs to the tenant, regardless of the affinity gain that it would produce. * The scope of a reusable element is part of the element, not a property of the requester. An element that a tenant holds is reusable by the sessions of that tenant within the domain, and it is not reusable outside it, even when the content would be identical for both tenants. * The affinity of one tenant is not visible to another. A tenant may observe its own affinity, and the operator may observe the aggregate, but neither may observe the presence or the reuse of another tenant's state. 14.5. Fairness, Herding, and Quota Affinity concentrates demand, and concentration has to be bounded by rules that are stated in advance rather than applied when the concentration has already degraded the service. Mo, et al. Expires 3 April 2027 [Page 39] Internet-Draft Agent Affinity September 2026 * Herding. Popular reusable elements attract sessions, and the attraction is self-reinforcing, because the sessions that are steered to the holder make the element more valuable there. A selection that applies affinity without a bound therefore produces a distribution that is worse than the one it started from in the tail, even when the mean improves. * Duplication. When the concentration of a reuse group on one instance degrades the objective of the sessions that share the element, the remedy is to admit a second copy of the element at another instance and to split the group, at the cost of establishing the element twice. The decision to duplicate is a comparison of the two costs, not a default. * Quota. A tenant with a large state footprint can occupy the capacity that other tenants need. The affinity of a tenant is therefore bounded by a quota on the capacity that its state may occupy and by a share of the reuse capacity of a shared instance, and the binding of the quota is observable rather than silent. * Fairness of measurement. The affinity gain of one tenant is not allowed to be achieved by degrading the tail of another, which requires that the objective of a session be evaluated per tenant and per communication mode rather than over the aggregate. 14.6. Measurement and Observability across Tenants The affinity metrics of Section 11 are reported per tenant as well as per key and per site, and the following restrictions apply to them. * A tenant may see its own footprint, its own hit ratio, and the binding of its own quota. * The operator may see the aggregate footprint, the concentration on an instance, and the refused reuse, because those are the quantities that the fairness rules bind to. Mo, et al. Expires 3 April 2027 [Page 40] Internet-Draft Agent Affinity September 2026 * Neither may see the per-session affinity of another tenant, and neither may infer it from a difference in latency, so the counters that are exposed are aggregated over the population of the tenant rather than over a key that another tenant could probe. 14.7. Requirements for Similar Agents and Tenants ST1. A CATS system SHOULD be able to identify the group of sessions, agents, or tenants that may reuse a given working set element, and SHOULD be able to express that group as part of the state information. ST2. A CATS system SHOULD be able to determine whether two agents may share an element by comparing the conditions of reuse, including the model revision, the tokenizer, the runtime, the precision, the template version, and the policy version, rather than by comparing their content or their identity. ST3. A CATS system SHOULD treat a shared element as available for a step only when the conditions of reuse are satisfied at the candidate that holds it. ST4. A CATS system SHOULD be able to express the affinity domain of a tenant, that is, the set of instances at which the sessions and the state of the tenant may be placed, and MUST enforce that domain as a constraint. ST5. A CATS system MUST NOT reuse a working set element across tenants, or across sessions of different users within a tenant, unless the policy of the holder of the element permits that reuse. ST6. A CATS system SHOULD account for the cost of a shared element once for the reuse group that establishes it, and SHOULD distinguish a shared reuse from a private one in its measurements. ST7. A CATS system SHOULD bound the capacity that the state of one tenant may occupy at a shared instance, and SHOULD make the binding of that bound observable. ST8. A CATS system SHOULD detect the concentration of a reuse group on a small number of instances, and SHOULD be able to admit an additional copy of an element when that concentration degrades the objective of the sessions that share the element. ST9. A CATS system SHOULD be able to state, per tenant, whether affinity may be traded against load, distance, or cost, and SHOULD apply the affinity policy of the tenant to the sessions of that tenant. Mo, et al. Expires 3 April 2027 [Page 41] Internet-Draft Agent Affinity September 2026 ST10. A CATS system SHOULD measure the affinity of a tenant without disclosing the affinity of an individual session to another tenant, and SHOULD NOT allow a tenant to observe the presence of another tenant's state. ST11. A CATS system SHOULD apply the isolation requirements of the tenant boundary to a shared element as well as to a private one, including to the timing of a hit and a miss. ST12. A CATS system SHOULD be able to revoke the reuse scope of an element when a tenant or a policy changes, without requiring the content of the element to be disclosed. 15. Interaction with Existing Work Metrics: The state, storage, and compute information described here is intended to appear as an additional category or categories in the metric framework of [I-D.ietf-cats-metric-definition], alongside the existing computing, communication, and service categories. In the terms used there, state availability, residency tier, compute readiness, and retention are naturally raw (Level 0) metrics, while the comparison of transfer against recomputation, affinity gain, and the concentration of affinity are derived (Level 1) quantities. Data model: If affinity information is to be configured or monitored, the data model [I-D.ietf-cats-data-model] needs to accommodate it. The minimum set of objects is a state handle or key space, a residency tier, a readiness descriptor for a computation, a reuse scope, a retention commitment, and a policy that states the permitted scope of reuse. OAM: The operational indicators for affinity-based selection are the hit ratio of state reuse, the volume and the latency of the state that a decision transfers, the affinity gain, the concentration of sessions on the instances that hold popular state, and the number of selections changed because state was not available where it was expected. These are candidates for the monitoring functions of [I-D.ietf-cats-oam-fw]. Selection mapping: The mapping described in [I-D.mo-cats-agent-selection-mapping] provides the framework in which the requirements of Sections 8 to 14 are applied: the descriptor of a step, the resource view of a candidate, the decision points, and the stability requirement are defined there, and this document adds the state, storage, and compute affinity content of each. Mo, et al. Expires 3 April 2027 [Page 42] Internet-Draft Agent Affinity September 2026 Distributed cache and storage protocols: Moving working set elements between instances is a data transfer problem that is expected to use existing transport building blocks. This document does not define a transfer protocol, and multiple such protocols may be used in one deployment. The reuse condition of Section 10 is the same condition that a cache applies when it decides whether a stored representation may be reused [RFC9111], and cache-control semantics are a useful model for the retention commitments of Section 12. 15.1. Traceability to the Agent Service Requirements The requirements of [I-D.mo-cats-agent-service-characteristics] are addressed as follows. The last two rows map the characteristics of that document that this document develops further. * R1, R2: WM1, WM2, WM6. * R8, R9, R10: S13, S14, WM5. * R11: S1, S2, S4. * R12: S6, S7, WM3. * R13: S8, S9, WM6. * R14: S15, S16, AM1, ST1, ST3. * R15, R16: AM6, LH9, LH10. * R17: MA6. * R19: AM3, AM4, LH5, LH6. * R20: AM4, WM8. * R21: MA4, MA5, MA9. * R22: MA8, MA10. * Long-horizon state and memory persistence (Section 5.8 of that document): LH1 to LH12, AM2, AM8. * Locality, governance, and tenancy (Section 5.9 of that document): S8, S9, ST4, ST5, ST11, ST12. 16. Operational Considerations Mo, et al. Expires 3 April 2027 [Page 43] Internet-Draft Agent Affinity September 2026 State-based selection changes the shape of the traffic that an operator sees. State transfer is additional traffic between service sites, it can be large, and it can be triggered by a steering decision, which means that a selection that ignores transfer cost can create the congestion that then degrades the next selection. Operators are therefore expected to bound the fraction of decisions that are allowed to trigger a transfer, and to treat the transfer network as a resource that the selection function must account for. Affinity can also concentrate load and can increase the impact of a failure: the instance that holds the state of many sessions is a higher-value target and a higher-impact failure than an instance that holds none. Deployments are expected to decide how much state is worth keeping warm, and to have a fallback path when the preferred instance is unavailable, which is consistent with the notion of a fallback decision [I-D.pang-cats-fallback-decision-framework]. Five operational properties follow from the mechanisms of Section 12. * The retention policy of an instance is a capacity decision, not only a performance decision, because a commitment that is not released withholds capacity from the steps that are running. * The override rule has to be testable in advance. An operator is expected to know at which measured concentration the override will be applied, rather than discovering it during an incident. * The failure of the holder of a large amount of state is a common-mode event for every session that it holds. A deployment is expected to bound the amount of state that one instance holds for one tenant, and to keep the recovery path (checkpoint or reconstruction) exercised. * Affinity metrics have to be attributable to a decision, because a hit ratio that cannot be attributed to a placement cannot be used to correct the placement. * The duplication decision of Section 14 has a cost that appears in the storage metrics and a benefit that appears in the tail of the objectives. An operator is expected to state which of the two is bounded, so that the decision can be taken consistently. Mo, et al. Expires 3 April 2027 [Page 44] Internet-Draft Agent Affinity September 2026 17. Security Considerations State is more sensitive than capacity. Exposing which elements an instance holds reveals patterns of use even when the content is not exposed, and the ability to request a transfer is the ability to move data between locations. The following considerations follow. Content protection: State that is transferred between instances should be protected in the same way as the session data from which it is derived, including at rest where the tier is persistent. Handle confidentiality and unforgeability: A state handle that can be guessed or forged is a handle that can be used to probe for the existence of state, and possession of a handle MUST NOT grant access to the state. Handles should be unguessable, scoped, and validated against authorization before use. A handle that names a shared element is scoped to the reuse group of that element, so that knowledge of it does not disclose the existence of state outside the group. Cross-tenant leakage: Reuse of state across sessions or tenants creates a direct path to information disclosure, including through the timing of a hit or a miss. Sharing scope should be enforced by the entity that holds the state, not only by the entity that requests it. Poisoning: A participant that can cause incorrect state to be associated with a valid handle can influence the output of the sessions that reuse it. Integrity of the association between a handle and its content should be established by the holder of the state. Denial through eviction: Because state availability is a resource that can be exhausted, a participant that can cause eviction can degrade other sessions. Implementations should not allow one session to evict the state of another without authorization. Affinity widens this surface, because an element that is shared by a reuse group is a single object whose eviction degrades the whole group, and because the retention commitments of one session compete with the state of the others. Compute substitution: A selection that treats capability tiers as interchangeable can place a step on a candidate that is cheaper but less capable, which changes the result of the step rather than only its cost. The constraint families of Section 8.2 keep capability out of the affinity valuation for this reason. 18. Privacy Considerations Mo, et al. Expires 3 April 2027 [Page 45] Internet-Draft Agent Affinity September 2026 Working set elements are derived from user inputs and therefore inherit their sensitivity. Even aggregated state availability information can reveal activity, and transfer events reveal which sessions are active between which sites. Where such information is exposed to the network it should be minimized, aggregated where possible, and retained only as long as it serves the selection function. Locality constraints are frequently imposed because of legal or regulatory requirements, and where a constraint is expressed in the information exchanged, the accuracy of that expression determines whether the requirement is met. Implementations SHOULD treat an unknown constraint as a prohibition rather than as permission. Shared reuse adds two considerations. First, a reuse group is an inference surface: a party that can observe the hit ratio of a shared element learns how many sessions use it and when, even when it learns nothing about their content. Second, the revocation of a reuse scope is a privacy control, and it is expected to be effective for the state that is already held and not only for the state that is subsequently created. The metrics of Section 11 are therefore reported per tenant and aggregated over a population, and not per key where a key could be probed by another tenant. 19. IANA Considerations This document has no IANA actions. 20. Normative References [I-D.ietf-cats-framework] Li, C., Du, Z., Boucadair, M., Contreras, L. M., et al., "A Framework for Computing-Aware Traffic Steering (CATS)", Work in Progress, Internet-Draft, draft-ietf-cats-framework-24, September 2026. [I-D.ietf-cats-metric-definition] Yao, K., et al., "CATS Metrics Definition", Work in Progress, Internet-Draft, draft-ietf-cats-metric-definition-12, September 2026. [I-D.mo-cats-agent-service-characteristics] Mo, Y., Yang, D., Zhou, C., "AI Agent Service Characteristics and Their Implications for Computing-Aware Traffic Steering", Work in Progress, Internet-Draft, draft-mo-cats-agent-service-characteristics-00, September 2026. 21. Informative References [CATS-CHARTER] IETF, "Computing-Aware Traffic Steering (CATS) Working Group Charter", . Mo, et al. Expires 3 April 2027 [Page 46] Internet-Draft Agent Affinity September 2026 [I-D.ietf-cats-usecases-requirements] Yao, K., et al., "Computing-Aware Traffic Steering (CATS) Problem Statement, Use Cases, and Requirements", Work in Progress, Internet-Draft, draft-ietf-cats-usecases-requirements-14, September 2026. [I-D.mo-cats-agent-selection-mapping] Mo, Y., Yang, D., Zhou, C., "A Selection Mapping Framework for AI Agent Services in Computing-Aware Traffic Steering", Work in Progress, Internet-Draft, draft-mo-cats-agent-selection-mapping-00, September 2026. [I-D.ietf-cats-data-model] Yao, H., Lin, C., et al., "Data Model for Computing-Aware Traffic Steering (CATS)", Work in Progress, draft-ietf-cats-data-model-00, September 2026. [I-D.ietf-cats-oam-fw] Fu, H., Xiong, Q., Du, Z., et al., "Computing- Aware Traffic Steering (CATS) Operations, Administration, and Maintenance (OAM) Framework", Work in Progress, Internet-Draft, draft-ietf-cats-oam-fw-01, July 2026. [I-D.li-cats-kv-cache-distribution] Li, Z., et al., "KV Cache Distribution for Distributed LLM Inference: Use Case and Requirements", Work in Progress, draft-li-cats-kv-cache-distribution-00, July 2026. [I-D.pang-cats-fallback-decision-framework] Pang, R., Ed., Han, M., Ed., Huang, T., Ed., "CATS Fallback Decision Framework", Work in Progress, Internet-Draft, draft-pang-cats-fallback-decision-framework-00, July 2026. [I-D.zhang-cats-token-aware-ts] Zhang, N., Ed., Han, M., Ed., Yi, X., Ed., "A token-aware traffic steering solution for agent service", Work in Progress, draft-zhang-cats-token-aware-ts-00, March 2026. [I-D.zhu-cats-metric-semantics] Zhu, M., "Operational Semantics for CATS Metric Consumption", Work in Progress, draft-zhu-cats-metric-semantics-01, August 2026. [RFC2119] Bradner, S., "Key words for use in RFCs to Indicate Requirement Levels", BCP 14, RFC 2119, DOI 10.17487/RFC2119, March 1997, . [RFC8174] Leiba, B., "Ambiguity of Uppercase vs Lowercase in RFC 2119 Key Words", BCP 14, RFC 8174, DOI 10.17487/RFC8174, May 2017, . [RFC9111] Fielding, R., Nottingham, M., Reschke, J., "HTTP Caching", STD 98, RFC 9111, DOI 10.17487/RFC9111, June 2022, . Acknowledgments Mo, et al. Expires 3 April 2027 [Page 47] Internet-Draft Agent Affinity September 2026 The authors would like to thank the participants of the CATS working group for the discussions that shaped this document. Authors' Addresses Y. Mo Huazhong University of Science and Technology Email: moyj@hust.edu.cn D. Yang Huazhong University of Science and Technology Email: d202581903@hust.edu.cn C. Zhou Huazhong University of Science and Technology Email: m202474228@hust.edu.cn Mo, et al. Expires 3 April 2027 [Page 48]