<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE rfc SYSTEM "rfc2629.dtd">
<?rfc toc="yes"?>
<?rfc tocdepth="3"?>
<?rfc symrefs="yes"?>
<?rfc sortrefs="yes"?>
<?rfc comments="yes"?>
<?rfc inline="yes"?>
<?rfc compact="yes"?>
<?rfc subcompact="no"?>

<rfc
    category="std"
    docName="draft-zhangb-cats-sci-implementation-01"
    ipr="trust200902"
    submissionType="IETF"
    consensus="true">

  <front>
    <title abbrev="CATS SCI Functional Implementation">
      CATS Service Contact Instance Functional Implementation
    </title>

    <author fullname="Bin Zhang" initials="B." surname="Zhang" role="editor">
      <organization>Pengcheng Laboratory</organization>
      <address>
        <postal>
          <street>Sibilong Street</street>
          <city>Shenzhen</city>
          <region>Guangdong</region>
          <code>518055</code>
          <country>CN</country>
        </postal>
        <email>zhangb@pcl.ac.cn</email>
      </address>
    </author>

    <author fullname="Bowen Shen" initials="B." surname="Shen" role="editor">
      <organization>Harbin Institute of Technology</organization>
      <address>
        <postal>
          <street>Taoyuan Street</street>
          <city>Shenzhen</city>
          <code>518055</code>
          <country>China</country>
        </postal>
        <email>shenbowen@stu.hit.edu.cn</email>
      </address>
    </author>

    <author fullname="Yixuan Yang" initials="Y." surname="Yang" role="editor">
      <organization>Sun Yat-sen University</organization>
      <address>
        <postal>
          <street>Gongchang Street</street>
          <city>Shenzhen</city>
          <region>Guangdong</region>
          <code>518055</code>
          <country>CN</country>
        </postal>
        <email>yangyx277@mail2.sysu.edu.cn</email>
      </address>
    </author>

    <date year="2026" month="September" day="15"/>

    <area>Routing</area>
    <workgroup>Computing-Aware Traffic Steering</workgroup>
    <keyword>CATS</keyword>
    <keyword>Service Contact Instance</keyword>
    <keyword>SCI</keyword>
    <keyword>metric aggregation</keyword>
    <keyword>metric reporting</keyword>
    <keyword>health monitoring</keyword>

    <abstract>
      <t>The Computing-Aware Traffic Steering (CATS) framework <xref target="I-D.ietf-cats-framework"/> introduces the concept of a Service Contact Instance (SCI) as the client-facing entity responsible for receiving and dispatching service requests. While the framework and the CATS metric documents define the components and the metrics, the concrete observable behavior of a Service Contact Instance - in particular, what it reports to the CATS Service Metric Agent (C-SMA), when it reports, and how service-instance health changes are reflected in the reported metrics - remains underspecified.</t>
      <t>This document fills that gap. It specifies the functional behavior of a CATS Service Contact Instance in terms of observable behavior and reporting semantics: how an SCI aggregates instance-level information into service-oriented metrics (e.g., Global Available Slots and Computing Time) as defined in <xref target="I-D.zhangb-cats-service-metrics-op"/>; how it monitors the health and status of underlying service instances and adjusts reported metrics accordingly; how it maintains affinity and handles failure scenarios; and how it reports metrics and status updates to the C-SMA, including update policies and threshold-based triggers. A decomposition of the SCI into internal functional components is provided as illustrative implementation guidance, not as a mandated software architecture.</t>
      <t>This document complements <xref target="I-D.ietf-cats-framework"/>, <xref target="I-D.ietf-cats-metric-definition-11"/>, and <xref target="I-D.zhangb-cats-service-metrics-op"/> by providing the operational execution layer for the SCI within the unified CATS architecture.</t>
    </abstract>
  </front>

  <middle>

    <section title="Introduction" anchor="intro">
      <t>The Computing-Aware Traffic Steering (CATS) framework <xref target="I-D.ietf-cats-framework"/> defines a Service Contact Instance (SCI) as a client-facing function responsible for receiving requests in the context of a given service. The framework states that an SCI may dispatch service requests to one or more service instances and that steering beyond an SCI is hidden to both clients and CATS components. The framework also notes that "the metrics of the service contact instance may be aggregate metrics from multiple service instances".</t>
      <t>However, the framework does not specify:</t>
      <t>
        <list style="symbols">
          <t>How an SCI derives and aggregates service-oriented metrics (e.g., Global Available Slots, Computing Time) from its underlying service instances.</t>
          <t>How an SCI monitors the health and availability of service instances, and how health changes are reflected in the reported metrics.</t>
          <t>How an SCI dispatches incoming requests to the appropriate service instance while preserving affinity.</t>
          <t>What an SCI reports to the CATS Service Metric Agent (C-SMA), when it reports, and how update frequency is controlled.</t>
        </list>
      </t>
      <t>This document specifies the functional behavior of a CATS Service Contact Instance to address these gaps. The normative requirements of this document are stated in terms of observable SCI behavior and of the reporting semantics towards the C-SMA; the decomposition of the SCI into internal functional components in <xref target="architecture"/> is provided as illustrative implementation guidance.</t>
      <t>The SCI acts as the boundary between the CATS-aware network and the service site. It is responsible for:</t>
      <t>
        <list style="symbols">
          <t>Collecting per-service-instance metrics (CPU, memory, GPU, throughput, queue depth, etc.).</t>
          <t>Aggregating these metrics into service-oriented abstractions (e.g., Global Available Slots, Computing Time) as defined in <xref target="I-D.zhangb-cats-service-metrics-op"/>.</t>
          <t>Monitoring service instance health and adjusting reported metrics accordingly.</t>
          <t>Receiving client requests from the Egress CATS-Forwarder and dispatching them to the most suitable service instance.</t>
          <t>Reporting aggregated metrics and status to the C-SMA.</t>
        </list>
      </t>

      <section title="Scope and Relationship to Other CATS Documents" anchor="scope">
        <t>This document specifies the functional behavior of the SCI as a CATS component. The scope is deliberately limited in two ways. First, the normative requirements concern the <em>observable behavior</em> of the SCI - the metrics it reports, the semantics of those metrics, the health state machine that drives metric adjustment, and the update policies that control reporting frequency. Second, how the SCI is internally structured (e.g., whether it is implemented as a load balancer, a gateway, or an application-level function) is a local matter; <xref target="architecture"/> provides an illustrative component decomposition that is informative in nature.</t>
        <t>The relationships to other CATS documents are as follows:</t>
        <t>
          <list style="symbols">
            <t><xref target="I-D.ietf-cats-framework"/> defines the CATS components and their interactions, including general implementation considerations on using CATS metrics (Section 5.4 of that document). This document provides the detailed SCI-side behavior that the framework leaves open.</t>
            <t><xref target="I-D.ietf-cats-metric-definition-11"/> defines the CATS metric taxonomy (Level 0/1/2), the CATS metric field template, and the "CATS Metrics" registry. This document uses the service-oriented metrics defined in <xref target="I-D.zhangb-cats-service-metrics-op"/> (registered in that registry) and specifies how an SCI produces and reports them; it does not define new metric categories.</t>
            <t><xref target="I-D.zhangb-cats-service-metrics-op"/> defines the service-oriented metrics (Global Available Slots, Computing Time, Price, Reputation, Security Level) and the C-PS side joint selection workflow. This document specifies the SCI-side execution: how these metrics are derived and maintained at the service site, and how they are reported to the C-SMA.</t>
            <t><xref target="I-D.ietf-cats-oam-fw"/> defines the OAM framework and requirements for CATS. The health monitoring and failure-detection behavior specified in <xref target="health-monitoring"/> of this document is a component-level behavior that complements, and is consistent with, the OAM requirements; the actual OAM mechanisms, measurement frameworks, and management interfaces are out of scope here.</t>
            <t><xref target="I-D.yl-cats-data-model"/> defines a YANG data model for the management of CATS. The report fields specified in <xref target="message-format"/> are semantic requirements; their mapping to configuration and management objects should be aligned with that data model where applicable.</t>
            <t><xref target="I-D.yxl-cats-protocols-applicability"/> discusses the applicability of existing protocols for CATS. The choice of transport for metric reporting (Section 7) should follow that document rather than being re-specified here.</t>
          </list>
        </t>
        <t>The framework states that how a service provider structures its services remains out of the scope of CATS. This document does not standardize the internal structure of a service site; it standardizes the behavior of the SCI that is observable by other CATS components, in particular the C-SMA and, indirectly, the C-PS.</t>
      </section>

      <section title="Requirements Language" anchor="requirements-language">
        <t>The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in BCP 14 <xref target="RFC2119"/> <xref target="RFC8174"/> when, and only when, they appear in all capitals, as shown here.</t>
      </section>
    </section>

    <section title="Terminology" anchor="terminology">
      <t>This document makes use of the terms defined in <xref target="I-D.ietf-cats-framework"/>, <xref target="I-D.ietf-cats-metric-definition-11"/>, and <xref target="I-D.zhangb-cats-service-metrics-op"/>. In particular:</t>
      <t>
        <list style="symbols">
          <t>CS-ID (CATS Service ID): An identifier for a service.</t>
          <t>CSCI-ID (CATS Service Contact Instance ID): An identifier for a service contact instance. In this document, it is used operationally as a locator (e.g., IP address and port) towards which the Egress CATS-Forwarder forwards traffic; the framework notes that a service contact instance is reachable via at least one Egress CATS-Forwarder. The exact encoding and resolution of the CSCI-ID are deployment-specific and to be confirmed with the working group.</t>
          <t>Global Available Slots (GAS): The maximum number of concurrent requests a service site is willing and able to serve for a specific CS-ID at a given time, as defined in <xref target="I-D.zhangb-cats-service-metrics-op"/>.</t>
          <t>Computing Time: The time required for the site to perform one service request, as defined in <xref target="I-D.zhangb-cats-service-metrics-op"/>.</t>
          <t>C-SMA (CATS Service Metric Agent): The functional entity that collects service metrics and advertises them to C-PSes.</t>
        </list>
      </t>
      <t>Additionally, the following terms are used in this document:</t>
      <t>
        <list style="symbols">
          <t>Service Instance (SI): A collection of running resources that are orchestrated following a service logic. An SCI may manage one or more SIs for the same CS-ID.</t>
          <t>Service Instance ID (SI-ID): A local identifier for a service instance within the scope of an SCI.</t>
          <t>Instance Metric: A raw or derived metric specific to a single service instance (e.g., instance CPU utilization, instance queue length).</t>
          <t>Health Status: A qualitative assessment of a service instance's operational state (e.g., healthy, degraded, failed).</t>
        </list>
      </t>
    </section>

    <section title="Service Contact Instance Functional Architecture" anchor="architecture">
      <t>This section provides a decomposition of the SCI into internal functional components. This decomposition is <em>illustrative</em>: it identifies the functions that an SCI implementation needs to provide, but it does not mandate a specific software architecture. The normative requirements of this document are expressed in terms of observable SCI behavior (Sections 4 to 8), in particular the reporting semantics towards the C-SMA in <xref target="reporting"/>. Implementations may realize the functions below in any manner that satisfies those requirements.</t>
      <t>Figure 1 shows a logical decomposition of the SCI:</t>

      <figure anchor="fig-arch" title="SCI Internal Functional Components (Illustrative)">
        <artwork><![CDATA[
+---------------------------------------------------------------+
|                   Service Contact Instance (SCI)              |
|                                                               |
|  +-----------+   +-----------+   +-----------+   +---------+  |
|  | Metric    |   | Health    |   | Dispatcher|   | Session |  |
|  | Collector |   | Monitor   |   | (DP)      |   | Manager |  |
|  | (MC)      |   | (HM)      |   |           |   | (SM)    |  |
|  +-----+-----+   +-----+-----+   +-----+-----+   +----+----+  |
|        |               |               |              |       |
|        +-------+       |       +-------+              |       |
|                |       |       |                      |       |
|  +-------------v-------v-------v----------------------v-----+ |
|  |              Metric Aggregator (MA)                      | |
|  +---------------------------+------------------------------+ |
|                              |                                |
|  +---------------------------v------------------------------+ |
|  |              Reporting Interface (RI)                    | |
|  +---------------------------+------------------------------+ |
|                              |                                |
+------------------------------v--------------------------------+
                               |
                        +------v------+
                        |   C-SMA     |
                        +-------------+
]]></artwork>
      </figure>

      <t>The internal components of an SCI are defined as follows. These definitions are descriptive; the behavioral requirements that each component helps satisfy are given in the corresponding sections.</t>
      <t>
        <list style="hanging">
          <t hangText="Metric Collector (MC):">Responsible for collecting raw metrics from each service instance. This includes CPU utilization, memory usage, GPU utilization, request queue depth, response latency, throughput, and error rates.</t>
          <t hangText="Metric Aggregator (MA):">Responsible for aggregating instance-level metrics into service-oriented metrics (e.g., GAS, Computing Time) as defined in <xref target="I-D.zhangb-cats-service-metrics-op"/>. The MA applies local policy and reference information (e.g., from <xref target="I-D.zhangb-cats-cmas"/>) to derive actionable metrics.</t>
          <t hangText="Health Monitor (HM):">Responsible for actively and passively monitoring the health of each service instance. It performs health checks, detects anomalies, and classifies instance status.</t>
          <t hangText="Failure Handler (FH):">Responsible for reacting to health changes. It updates the metric aggregator when instances fail or recover, and may trigger re-dispatch of in-flight requests.</t>
          <t hangText="Dispatcher (DP):">Responsible for receiving client requests from the Egress CATS-Forwarder and dispatching them to the most suitable service instance based on local load, health status, and affinity requirements.</t>
          <t hangText="Session Manager (SM):">Responsible for maintaining session state, including affinity bindings between client flows and service instances. It tracks active sessions and notifies the Metric Aggregator of slot allocation and release.</t>
          <t hangText="Reporting Interface (RI):">Responsible for formatting and sending aggregated metrics and status updates to the C-SMA.</t>
        </list>
      </t>

      <section title="SCI Position in CATS Framework" anchor="position">
        <t>The SCI sits at the boundary between the CATS network and the service site. Its relationships with other CATS components are:</t>
        <t>
          <list style="symbols">
            <t>The SCI receives traffic from the Egress CATS-Forwarder.</t>
            <t>The SCI reports metrics to the C-SMA, which may be co-located with or adjacent to the SCI.</t>
            <t>The C-SMA advertises the SCI's metrics (along with the CSCI-ID) to the C-PS.</t>
            <t>The C-PS uses these metrics to make traffic steering decisions.</t>
            <t>The SCI does not directly interact with the C-PS or C-NMA; all control-plane communication goes through the C-SMA.</t>
          </list>
        </t>
        <t>The SCI is transparent to the client. The client sees only the CSCI-ID (e.g., an IP address and port) and is unaware of the internal service instances managed by the SCI.</t>
      </section>
    </section>

    <section title="Service Instance Metric Collection" anchor="metric-collection">
      <t>The Metric Collector (MC) gathers the following categories of information from each service instance:</t>

      <section title="Metric Collection Scope" anchor="collection-scope">
        <t>
          <list style="hanging">
            <t hangText="Resource Metrics:">Raw computing resource utilization. Examples include: CPU utilization (%), Memory utilization (%), GPU utilization (%) and GPU memory usage, NPU/TPU utilization (if applicable), Disk I/O and network I/O rates.</t>
            <t hangText="Performance Metrics:">Service-level performance indicators. Examples include: Request queue depth, Average response time, Requests per second (throughput), Error rate (5xx errors, timeout rate).</t>
            <t hangText="Status Metrics:">Operational state indicators. Examples include: Instance health (up/down), Load average, Number of active connections/sessions, Custom application-level status.</t>
          </list>
        </t>
        <t>The MC collects these metrics at a configurable sampling interval (e.g., every 5-10 seconds). The specific collection mechanism (e.g., Prometheus scraping, SNMP, gRPC, HTTP health endpoints) is deployment-specific and outside the scope of this document.</t>
      </section>

      <section title="Collection Methods" anchor="collection-methods">
        <t>The MC SHOULD support multiple collection methods to accommodate different service instance types:</t>
        <t>
          <list style="hanging">
            <t hangText="Pull-based:">The MC periodically queries each service instance via a standardized metrics endpoint (e.g., Prometheus /metrics, HTTP REST API, SNMP). This is suitable for instances that expose metrics via well-known interfaces.</t>
            <t hangText="Push-based:">Service instances push metrics to the MC via a message queue or streaming protocol (e.g., gRPC streaming, MQTT, Kafka). This is suitable for high-frequency updates or event-driven metrics.</t>
            <t hangText="Agent-based:">A lightweight agent runs alongside each service instance and reports metrics to the MC. This is suitable for environments where instances cannot be modified to expose endpoints.</t>
          </list>
        </t>
        <t>The MC MUST support at least one of these methods. In mixed deployments, the MC MAY use different methods for different service instances.</t>
      </section>

      <section title="Aggregation and Derivation" anchor="aggregation">
        <t>The Metric Aggregator (MA) combines instance-level metrics into the service-oriented metrics defined in <xref target="I-D.zhangb-cats-service-metrics-op"/>, which are meaningful for CATS traffic steering. The definitions and semantics of these metrics are normative in <xref target="I-D.zhangb-cats-service-metrics-op"/>; this section specifies the SCI-side execution details for deriving and maintaining them. The key aggregated metrics are:</t>
        <t>
          <list style="hanging">
            <t hangText="Global Available Slots (GAS):">The MA calculates the total GAS for the SCI by summing the available capacity of all healthy service instances. For each instance, the available capacity is derived from: the instance's maximum concurrent request capacity (the initial GAS value, which can be calculated based on resource allocation and the service reference information defined in <xref target="I-D.zhangb-cats-service-metrics-op"/>), the instance's current active session count (from the Session Manager), and the instance's health status (from the Health Monitor). Unhealthy instances contribute 0 to GAS.</t>
          </list>
        </t>
        <t>GAS = SUM_over_instances(max_capacity_i - active_sessions_i) for all healthy instances i</t>
        <t>The MA MAY apply a local policy factor (e.g., a safety margin of 80%, as an example configuration) to prevent over-subscription.</t>
        <t>
          <list style="hanging">
            <t hangText="Computing Time:">The MA estimates the Computing Time for the SCI based on the observed response times of service instances. This can be computed as: the weighted average of instance response times, weighted by instance load; the median or percentile (e.g., p95) of instance response times to account for outliers; or a dynamically adjusted estimate based on current load and historical trends. The MA also reports the min/max computing time and the associated input token counts for reference.
            </t>
            <t hangText="Optional Metrics:">The MA MAY also derive optional service attributes such as Price, Reputation, and Security Level as defined in <xref target="I-D.zhangb-cats-service-metrics-op"/>.</t>
          </list>
        </t>
        <t>The derivation algorithm is a local matter and is not standardized by this document. However, the MA MUST ensure that the reported metrics are consistent and comparable across updates, and that they conform to the semantics and units defined in <xref target="I-D.zhangb-cats-service-metrics-op"/>.</t>
      </section>
    </section>

    <section title="Service Instance Health Monitoring" anchor="health-monitoring">
      <t>The Health Monitor (HM) monitors the health of each service instance and drives metric adjustment and dispatch behavior. The health monitoring behavior specified in this section is a component-level behavior of the SCI; it complements the CATS OAM framework <xref target="I-D.ietf-cats-oam-fw"/>, which defines the OAM layering and requirements for CATS as a whole. The mechanisms used to realize health checks (e.g., in-band OAM, dedicated probes) are out of scope of this document.</t>
      <t>The HM performs the following types of health checks on each service instance:</t>

      <section title="Health Check Types" anchor="health-check-types">
        <t>
          <list style="hanging">
            <t hangText="Liveness Check:">Verifies that the service instance process is running and responsive. This is typically done via a simple TCP connection or HTTP GET to a /health endpoint. Failure indicates the instance is down.</t>
            <t hangText="Readiness Check:">Verifies that the service instance is ready to accept new requests. This may check: whether the instance has finished initialization, whether critical dependencies are available, and whether the instance is not in a maintenance mode.</t>
            <t hangText="Performance Check:">Verifies that the service instance is performing within acceptable bounds. This may check: response time against a threshold, error rate against a threshold, and resource utilization (e.g., CPU &gt; 90% for 30 seconds).</t>
          </list>
        </t>
        <t>The HM SHOULD perform these checks at regular intervals (e.g., every 5-10 seconds for liveness, every 30 seconds for readiness and performance). The intervals and thresholds SHOULD be configurable.</t>
      </section>

      <section title="Failure Detection and Recovery" anchor="failure-detection">
        <t>The HM classifies each service instance into one of the following health states:</t>
        <t>
          <list style="symbols">
            <t>HEALTHY: The instance is fully operational and accepting requests.</t>
            <t>DEGRADED: The instance is operational but experiencing performance issues (e.g., high latency, high error rate). New requests MAY be steered away, but existing sessions are maintained.</t>
            <t>UNHEALTHY: The instance is not operational or not ready. No new requests are dispatched to this instance. Existing sessions MAY be migrated or terminated based on policy.</t>
            <t>UNKNOWN: The HM cannot determine the instance's status (e.g., network partition). The instance is treated as UNHEALTHY until status is confirmed.</t>
          </list>
        </t>
        <t>Figure 2 shows the health state machine and the transitions between states:</t>
        <figure anchor="fig-health-statemachine" title="Service Instance Health State Machine">
          <artwork><![CDATA[
   +------------+
   |  UNKNOWN   |
   +-----+------+
         | status confirmed (treated as UNHEALTHY meanwhile)
         v
+--------+  fail xN   +----------+  fail xN   +-----------+
| HEALTHY| ---------> | DEGRADED | ---------> | UNHEALTHY |
|        | <--------- |          | <--------- |           |
+--------+  pass xM   +----------+  pass xM   +-----------+
    ^          |            ^          |            ^
    |          |            |          |            |
    +----------+------------+----------+------------+
       recovery (pass xM) at any degraded/unhealthy state
]]></artwork>
        </figure>
        <t>State transitions trigger the following actions:</t>
        <t>
          <list style="symbols">
            <t>HEALTHY -&gt; DEGRADED: The HM notifies the MA to reduce the instance's contribution to GAS. The DP reduces or stops sending new requests to the instance.</t>
            <t>DEGRADED -&gt; UNHEALTHY: The HM notifies the MA to set the instance's contribution to GAS to 0. The DP stops sending new requests. The SM initiates session migration or graceful termination for affected sessions.</t>
            <t>UNHEALTHY -&gt; HEALTHY: The HM notifies the MA to restore the instance's contribution to GAS. The DP resumes sending requests.</t>
          </list>
        </t>
        <t>The HM MUST implement a hysteresis mechanism (e.g., require N consecutive failed checks before marking UNHEALTHY, and M consecutive passed checks before marking HEALTHY; N and M are example configuration parameters) to avoid flapping.</t>
      </section>

      <section title="Metric Adjustment on Health Changes" anchor="metric-adjustment">
        <t>When the HM detects a health state change, the MA MUST adjust the aggregated metrics accordingly:</t>
        <t>
          <list style="symbols">
            <t>When an instance becomes DEGRADED, the MA MAY reduce its max_capacity by a configured degradation factor (e.g., 50%) or set it to 0 based on local policy.</t>
            <t>When an instance becomes UNHEALTHY, the MA MUST set its contribution to GAS to 0.</t>
            <t>When an instance recovers to HEALTHY, the MA MUST restore its contribution to GAS based on current load.</t>
            <t>The Computing Time estimate MUST be recalculated to exclude UNHEALTHY instances and weight DEGRADED instances lower.</t>
          </list>
        </t>
        <t>These adjustments are reflected in the next metric report to the C-SMA. The MA SHOULD batch rapid health changes to avoid excessive updates.</t>
      </section>
    </section>

    <section title="Request Dispatch and Load Balancing" anchor="dispatch">
      <t>The Dispatcher (DP) receives client requests from the Egress CATS-Forwarder and selects a service instance to handle each request. The dispatch decision is based on:</t>

      <section title="Dispatch Decision Logic" anchor="dispatch-logic">
        <t>
          <list style="numbers">
            <t>Affinity Requirements: If the request belongs to an existing session with affinity, the DP MUST dispatch to the same service instance (if healthy).</t>
            <t>Health Status: The DP MUST NOT dispatch to UNHEALTHY instances. It SHOULD avoid DEGRADED instances unless no HEALTHY instances are available.</t>
            <t>Load Balancing Policy: The DP selects among HEALTHY instances using a local load balancing algorithm. Supported algorithms include round-robin, least-connections, weighted response time, resource-aware (lowest CPU/memory/GPU utilization), and slot-based (most available slots, i.e., max_capacity - active_sessions). The choice of algorithm is a local matter.</t>
            <t>Local Policy: The DP MAY apply additional policies such as price optimization, security level requirements, or instance preference.</t>
          </list>
        </t>
        <t>The DP MUST handle the case where no HEALTHY instances are available. In this case, it MAY:</t>
        <t>
          <list style="symbols">
            <t>Return an error to the client (e.g., HTTP 503 Service Unavailable).</t>
            <t>Queue the request and retry after a timeout.</t>
            <t>Dispatch to a DEGRADED instance as a last resort.</t>
          </list>
        </t>
      </section>

      <section title="Affinity Handling" anchor="affinity">
        <t>The Session Manager (SM) maintains affinity bindings between client flows and service instances. Affinity is identified by a flow key (e.g., 5-tuple: source IP, destination IP, source port, destination port, protocol).</t>
        <t>When a new request arrives:</t>
        <t>
          <list style="symbols">
            <t>The SM checks if the flow key has an existing binding.</t>
            <t>If yes, and the bound instance is HEALTHY, the DP dispatches to that instance.</t>
            <t>If yes, but the bound instance is UNHEALTHY, the SM removes the binding and the DP selects a new instance.</t>
            <t>If no binding exists, the DP selects an instance and the SM creates a new binding.</t>
          </list>
        </t>
        <t>Affinity bindings have a configurable timeout. After the timeout expires with no activity, the SM removes the binding.</t>
        <t>The SM MUST support affinity at the flow level. It MAY also support affinity at the session level (e.g., for HTTP sessions identified by cookies or session IDs).</t>
      </section>

      <section title="Session Lifecycle Management" anchor="session-lifecycle">
        <t>The SM tracks the lifecycle of each session:</t>
        <t>
          <list style="hanging">
            <t hangText="Allocation:">When a request is dispatched, the SM increments the active session count for the selected instance and records the flow binding.</t>
            <t hangText="Renewal:">For long-lived sessions, the SM may refresh the binding timeout on each packet or keepalive.</t>
            <t hangText="Release:">When the session ends (e.g., TCP FIN/RST, timeout, explicit logout), the SM decrements the active session count and removes the binding.</t>
          </list>
        </t>
        <t>The SM notifies the MA of session allocation and release events so that GAS can be updated. However, per-session changes do not necessarily trigger immediate reports to the C-SMA; the MA applies local aggregation and threshold policies.</t>
      </section>
    </section>

    <section title="Metric Reporting to C-SMA" anchor="reporting">
      <t>The Reporting Interface (RI) communicates with the C-SMA to report aggregated metrics and status updates. This section is normative: it specifies the semantics of what an SCI reports, when it reports, and how update frequency is controlled. The transport mechanisms used to carry the reports are deployment-specific and should follow the protocol applicability analysis in <xref target="I-D.yxl-cats-protocols-applicability"/>.</t>

      <section title="Reporting Interface" anchor="reporting-interface">
        <t>The RI SHOULD use a reliable transport (e.g., TCP, QUIC, or HTTP/2) to ensure metric delivery. The specific protocol is deployment-specific.</t>
        <t>In centralized deployments (e.g., SDN controller), the RI MAY use a RESTful API (e.g., RESTCONF <xref target="RFC8040"/>) or gRPC to push metrics to the C-SMA.</t>
        <t>In distributed deployments, the RI MAY use a routing protocol extension (e.g., BGP-LS <xref target="RFC8571"/>, GRASP <xref target="RFC8990"/>) or a dedicated CATS metric distribution protocol. The applicability of these mechanisms is discussed in <xref target="I-D.yxl-cats-protocols-applicability"/>; this document does not select among them.</t>
      </section>

      <section title="Update Policies and Thresholds" anchor="update-policies">
        <t>The RI applies the following policies to control update frequency. These policies directly address the question of the appropriate frequency and scope of metric distribution, which the CATS framework leaves open:</t>
        <t>
          <list style="hanging">
            <t hangText="Periodic Heartbeat:">A full report is sent at a fixed interval (e.g., every 60 seconds, as an example configuration) to maintain soft state and prevent stale metrics.</t>
            <t hangText="Threshold-based Trigger:">A delta report is sent when any metric crosses a configured threshold. Examples include: GAS changes by more than 10% or an absolute value of 50, Computing Time deviates by more than 20%, or Health status changes for any instance. The numerical values above are examples; the thresholds MUST be configurable.</t>
            <t hangText="Event-based Trigger:">A delta report is sent immediately upon critical events: all instances become UNHEALTHY (GAS = 0), an instance recovers from UNHEALTHY to HEALTHY, or security-related events (e.g., detected attack).</t>
            <t hangText="Rate Limiting:">The RI MUST implement rate limiting to prevent excessive updates during rapid fluctuations (e.g., during a thundering herd or DDoS attack).</t>
          </list>
        </t>
        <t>The specific thresholds and intervals SHOULD be configurable via the CATS Management Plane, and their alignment with the CATS data model <xref target="I-D.yl-cats-data-model"/> SHOULD be provided.</t>
      </section>

      <section title="Report Message Format" anchor="message-format">
        <t>The report message contains the following fields. The semantics of these fields are normative; their encoding and wire format depend on the chosen transport and are out of scope of this document:</t>
        <table anchor="tbl-report-fields">
          <name>Report Message Fields</name>
          <thead>
            <tr>
              <th>Field</th>
              <th>Description</th>
            </tr>
          </thead>
          <tbody>
            <tr><td>CSCI-ID</td><td>Identifier of the reporting SCI</td></tr>
            <tr><td>CS-ID</td><td>Identifier of the service</td></tr>
            <tr><td>Timestamp</td><td>Time of metric generation (Unix epoch)</td></tr>
            <tr><td>Sequence Number</td><td>Monotonically increasing sequence number</td></tr>
            <tr><td>GAS</td><td>Global Available Slots, as defined in <xref target="I-D.zhangb-cats-service-metrics-op"/></td></tr>
            <tr><td>Computing Time</td><td>Estimated processing time in milliseconds, as defined in <xref target="I-D.zhangb-cats-service-metrics-op"/></td></tr>
            <tr><td>Health Summary</td><td>Number of HEALTHY/DEGRADED/UNHEALTHY instances</td></tr>
            <tr><td>Optional Metrics</td><td>Price, Reputation, Security Level (if applicable), as defined in <xref target="I-D.zhangb-cats-service-metrics-op"/></td></tr>
            <tr><td>Instance Details</td><td>Per-instance status (optional, for debugging)</td></tr>
          </tbody>
        </table>
        <t>The encoding details, including field lengths and wire format, are TBD and depend on the chosen transport protocol, following the applicability analysis in <xref target="I-D.yxl-cats-protocols-applicability"/>.</t>
      </section>
    </section>

    <section title="External Communication" anchor="external">
      <t>The SCI exposes a client-facing interface that is the CSCI-ID (e.g., an IP address and port). Clients send service requests to this interface, and the Egress CATS-Forwarder forwards traffic to it.</t>

      <section title="Client-Facing Interface" anchor="client-interface">
        <t>The SCI MUST support the following on the client-facing interface:</t>
        <t>
          <list style="symbols">
            <t>Accept incoming connections/requests from the Egress CATS-Forwarder.</t>
            <t>Maintain transport-layer state (e.g., TCP connections) for the duration of the session.</t>
            <t>Support the service protocol (e.g., HTTP, gRPC, RTP) required by the CS-ID.</t>
          </list>
        </t>
        <t>The SCI MUST NOT expose internal service instance details (e.g., SI-IDs, internal IP addresses) to the client.</t>
      </section>

      <section title="Egress CATS-Forwarder Interface" anchor="egress-interface">
        <t>The SCI communicates with the Egress CATS-Forwarder via the underlay network. The Egress CATS-Forwarder is responsible for:</t>
        <t>
          <list style="symbols">
            <t>Decapsulating CATS overlay traffic and forwarding it to the SCI's CSCI-ID.</t>
            <t>Receiving responses from the SCI and encapsulating them for return to the Ingress CATS-Forwarder.</t>
          </list>
        </t>
        <t>The SCI does not need to be CATS-aware; it operates as a standard service endpoint from the network perspective. However, the SCI MAY support CATS-specific signaling (e.g., for affinity or session state synchronization) if defined by future documents.</t>
      </section>
    </section>

    <section title="Security Considerations" anchor="security">
      <t>The SCI handles client requests and manages internal service instances, making it a critical security boundary. The following security measures are REQUIRED:</t>
      <t>
        <list style="hanging">
          <t hangText="Authentication:">The SCI MUST authenticate the Egress CATS-Forwarder to prevent unauthorized traffic injection. This may use mutual TLS, IPsec, or network-layer authentication.</t>
          <t hangText="Authorization:">The SCI MUST verify that incoming requests are authorized for the requested CS-ID. This may involve validating tokens, certificates, or network-level policies.</t>
          <t hangText="Isolation:">Service instances managed by the same SCI MUST be isolated from each other to prevent cross-tenant attacks. This includes network isolation (e.g., separate VLANs or namespaces) and resource isolation (e.g., CPU/memory limits).</t>
          <t hangText="Metric Integrity:">The RI MUST protect the integrity of metric reports sent to the C-SMA. This may use TLS, message authentication codes (MACs), or digital signatures.</t>
          <t hangText="Confidentiality:">Metric reports SHOULD be encrypted to prevent disclosure of internal service topology and capacity.</t>
          <t hangText="DDoS Protection:">The SCI SHOULD implement rate limiting and connection throttling to protect against denial-of-service attacks that could exhaust GAS.</t>
          <t hangText="Health Check Security:">Health check endpoints exposed by service instances SHOULD be protected to prevent spoofing or manipulation.</t>
        </list>
      </t>
    </section>

    <section title="IANA Considerations" anchor="iana">
      <t>This document defines the semantics of the metrics reported by an SCI, but does not define a new protocol or a new wire format for metric reporting; the encoding depends on the chosen transport (see <xref target="reporting-interface"/>). Therefore, this document has no IANA actions at this time.</t>
      <t>If a dedicated CATS metric reporting protocol or a registry for report fields is defined in the future, the corresponding IANA registrations would be specified in that document, in coordination with the CATS working group.</t>
    </section>

  </middle>

  <back>

    <references title="Normative References">
      <reference anchor="RFC2119">
        <front><title>Key words for use in RFCs to Indicate Requirement Levels</title>
        <author initials="S." surname="Bradner" fullname="Scott Bradner"/>
        <date year="1997" month="March"/></front>
        <seriesInfo name="BCP" value="14"/>
        <seriesInfo name="RFC" value="2119"/>
        <format type="TXT" target="https://www.rfc-editor.org/rfc/rfc2119.txt"/>
      </reference>

      <reference anchor="RFC8174">
        <front><title>Ambiguity of Uppercase vs Lowercase in RFC 2119 Key Words</title>
        <author initials="B." surname="Leiba" fullname="Barry Leiba"/>
        <date year="2017" month="May"/></front>
        <seriesInfo name="BCP" value="14"/>
        <seriesInfo name="RFC" value="8174"/>
        <format type="TXT" target="https://www.rfc-editor.org/rfc/rfc8174.txt"/>
      </reference>

      <reference anchor="I-D.ietf-cats-framework">
        <front><title>A Framework for Computing-Aware Traffic Steering (CATS)</title>
        <author initials="C." surname="Li" fullname="Cheng Li"/>
        <author initials="Z." surname="Du" fullname="Zongpeng Du"/>
        <author initials="M." surname="Boucadair" fullname="Mohamed Boucadair"/>
        <author initials="L. M." surname="Contreras" fullname="Luis M. Contreras"/>
        <author initials="J." surname="Drake" fullname="John Drake"/>
        <date year="2026" month="April" day="2"/></front>
        <seriesInfo name="Internet-Draft" value="draft-ietf-cats-framework-24"/>
        <format type="TXT" target="https://www.ietf.org/archive/id/draft-ietf-cats-framework-24.txt"/>
      </reference>

      <reference anchor="I-D.ietf-cats-metric-definition-11">
        <front><title>CATS Metrics Definition</title>
        <author initials="Y." surname="Kehan" fullname="Yao Kehan"/>
        <author initials="C." surname="Li" fullname="Cheng Li"/>
        <author initials="L. M." surname="Contreras" fullname="Luis M. Contreras"/>
        <author initials="J." surname="Ros-Giralt" fullname="Jordi Ros-Giralt"/>
        <author initials="G." surname="Zeng" fullname="Guanming Zeng"/>
        <date year="2026" month="September" day="4"/></front>
        <seriesInfo name="Internet-Draft" value="draft-ietf-cats-metric-definition-11"/>
        <format type="TXT" target="https://www.ietf.org/archive/id/draft-ietf-cats-metric-definition-11.txt"/>
      </reference>

      <reference anchor="I-D.zhangb-cats-service-metrics-op">
        <front><title>Computing Service Metrics Operation and Joint Service Selection under CATS</title>
        <author initials="B." surname="Zhang" fullname="Bin Zhang"/>
        <author initials="Y." surname="Dai" fullname="Yina Dai"/>
        <author initials="Z." surname="Du" fullname="Zongpeng Du"/>
        <author initials="G." surname="Zeng" fullname="Guanming Zeng"/>
        <author initials="C." surname="Miao" fullname="Chuanyang Miao"/>
        <date year="2026" month="September" day="14"/></front>
        <seriesInfo name="Internet-Draft" value="draft-zhangb-cats-service-metrics-op-05"/>
        <format type="TXT" target="https://www.ietf.org/archive/id/draft-zhangb-cats-service-metrics-op-05.txt"/>
      </reference>
    </references>

    <references title="Informative References">
      <reference anchor="I-D.zhangb-cats-cmas">
        <front><title>Public Service Platform for Computing-Aware Traffic Steering (CATS)</title>
        <author initials="B." surname="Zhang" fullname="Bin Zhang"/>
        <author initials="Y." surname="Dai" fullname="Yina Dai"/>
        <author initials="Z." surname="Du" fullname="Zongpeng Du"/>
        <date year="2026" month="August" day="18"/></front>
        <seriesInfo name="Internet-Draft" value="draft-zhangb-cats-cmas-06"/>
        <format type="TXT" target="https://www.ietf.org/archive/id/draft-zhangb-cats-cmas-06.txt"/>
      </reference>

      <reference anchor="I-D.ietf-cats-oam-fw">
        <front><title>OAM Framework and Requirements for Computing-Aware Traffic Steering (CATS)</title>
        <author initials="H." surname="Fu" fullname="Huakai Fu"/>
        <author initials="Q." surname="Xiong" fullname="Quan Xiong"/>
        <author initials="Z." surname="Du" fullname="Zongpeng Du"/>
        <author initials="B." surname="Liu" fullname="Bo Liu"/>
        <date year="2026" month="July" day="21"/></front>
        <seriesInfo name="Internet-Draft" value="draft-ietf-cats-oam-fw-01"/>
        <format type="TXT" target="https://www.ietf.org/archive/id/draft-ietf-cats-oam-fw-01.txt"/>
      </reference>

      <reference anchor="I-D.yl-cats-data-model">
        <front><title>Data Model for Computing-Aware Traffic Steering (CATS)</title>
        <author initials="C." surname="Lin" fullname="Changwang Lin"/>
        <author initials="Z." surname="Li" fullname="Zhenqiang Li"/>
        <author initials="Q." surname="Xiong" fullname="Quan Xiong"/>
        <author initials="L. M." surname="Contreras" fullname="Luis M. Contreras"/>
        <date year="2026" month="July" day="6"/></front>
        <seriesInfo name="Internet-Draft" value="draft-yl-cats-data-model-07"/>
        <format type="TXT" target="https://www.ietf.org/archive/id/draft-yl-cats-data-model-07.txt"/>
      </reference>

      <reference anchor="I-D.yxl-cats-protocols-applicability">
        <front><title>Protocols Applicability for Computing-Aware Traffic Steering (CATS)</title>
        <author initials="H." surname="Yao" fullname="Huijuan Yao"/>
        <author initials="L." surname="Dunbar" fullname="Linda Dunbar"/>
        <author initials="Q." surname="Xiong" fullname="Quan Xiong"/>
        <author initials="C." surname="Lin" fullname="Changwang Lin"/>
        <date year="2026" month="July" day="2"/></front>
        <seriesInfo name="Internet-Draft" value="draft-yxl-cats-protocols-applicability-01"/>
        <format type="TXT" target="https://www.ietf.org/archive/id/draft-yxl-cats-protocols-applicability-01.txt"/>
      </reference>

      <reference anchor="RFC8040">
        <front><title>RESTCONF Protocol</title>
        <author initials="A." surname="Bierman" fullname="Andy Bierman"/>
        <author initials="M." surname="Bjorklund" fullname="Martin Bjorklund"/>
        <author initials="K." surname="Watsen" fullname="Kent Watsen"/>
        <date year="2017" month="January"/></front>
        <seriesInfo name="RFC" value="8040"/>
        <format type="TXT" target="https://www.rfc-editor.org/rfc/rfc8040.txt"/>
      </reference>

      <reference anchor="RFC8571">
        <front><title>BGP - Link State (BGP-LS) Advertisement of IGP Traffic Engineering Performance Metric Extensions</title>
        <author initials="L." surname="Ginsberg" fullname="Les Ginsberg"/>
        <author initials="S." surname="Previdi" fullname="Stefano Previdi"/>
        <author initials="Q." surname="Wu" fullname="Qin Wu"/>
        <date year="2019" month="March"/></front>
        <seriesInfo name="RFC" value="8571"/>
        <format type="TXT" target="https://www.rfc-editor.org/rfc/rfc8571.txt"/>
      </reference>

      <reference anchor="RFC8990">
        <front><title>GeneRic Autonomic Signaling Protocol (GRASP)</title>
        <author initials="C." surname="Bormann" fullname="Carsten Bormann"/>
        <author initials="B." surname="Carpenter" fullname="Brian Carpenter"/>
        <author initials="B." surname="Liu" fullname="Bing Liu"/>
        <date year="2021" month="March"/></front>
        <seriesInfo name="RFC" value="8990"/>
        <format type="TXT" target="https://www.rfc-editor.org/rfc/rfc8990.txt"/>
      </reference>
    </references>

    <section title="Acknowledgments" anchor="ack">
      <t>The authors thank the CATS working group for their valuable feedback and contributions to this document.</t>
    </section>

  </back>

</rfc>
<!-- （注：内容由AI生成） -->
