<?xml version="1.0"?> <!-- -*- fill-column:100; -*- -->
<?xml-stylesheet type="text/xsl" href="lib/rfc2629.xslt"?>
<?rfc toc="yes" ?>
<?rfc symrefs="yes" ?>
<?rfc sortrefs="yes" ?>
<?rfc compact="yes"?>
<?rfc subcompact="no" ?>
<?rfc linkmailto="no" ?>
<?rfc editing="no" ?>
<?rfc comments="yes"?>
<?rfc inline="yes"?>
<?rfc rfcedstyle="yes"?>
<?rfc-ext allow-markup-in-artwork="yes" ?>
<?rfc-ext include-index="no" ?>

<rfc ipr="trust200902"
     category="exp"
     submissionType="IETF"
     docName="draft-thierry-bulk-08"
     xmlns:xi="http://www.w3.org/2001/XInclude">
  <front>
    <title abbrev="BULK1">Binary Universal Language Kit 1.0</title>

    <author initials="P." surname="Thierry" fullname="Pierre Thierry">
      <organization>Comonad Dev</organization>
      <address>
        <email>pierre@comonad.dev</email>
      </address>
    </author>

    <date day="27" month="09" year="2026" />
    <keyword>binary</keyword>

    <abstract>
      <t>
        This specification describes a simple, decentrally extensible and efficient format for data
        serialization.
      </t>
    </abstract>

  </front>

  <middle>
    <section anchor="intro" title="Introduction">
      <section title="Rationale">
        <t>
          This specification aims at finding an original trade-off between transparency, syntax
          complexity, generality, extensibility, decentralization, discoverability, compactness,
          streamability, safety, processing speed and processing footprint for a data format (see
          <xref target="concepts" format="none">definitions</xref>). It is our opinion that every
          widely used existing format occupy a different position than this one in the solution
          space for formats, that none is better on all axes, and that this one is the current best
          on several axes, hence this new design. It is also our opinion that some of those existing
          formats constitute an optimal solution for their specific use case, either in a absolute
          sense, or at least at the time of their design. But the ever-changing field of IT now
          faces new challenges that call for a new approach.
        </t>
	<t>
	  In particular, whereas the previous trend for Internet and Web standards and programming
	  tools has been to create human-readable syntaxes for data and protocols, the advent of
	  technologies like <xref target="protobuf">protocol buffers</xref>, <xref
	  target="RFC8949">CBOR</xref>, <xref target="Thrift">Thrift</xref>, the various binary
	  serializations for JSON like <xref target="Avro">Avro</xref> or <xref
	  target="Smile">Smile</xref>, or the binary <xref target="RFC7540">HTTP/2</xref> seem to
	  indicate that the time is ripe for a generalized use of binary, reserved until now for the
	  low-level protocols. The lessons about flexibility learnt in the previous switch from
	  binary to plain text can now be applied to efficient binary syntaxes.
	</t>
	<section anchor="concepts" title="Definitions">
	  <t>
	    By transparency, we mean the property of a format that can be parsed even by an
	    application that doesn't understand the semantics of every part of the processed data.
	  </t>
	  <t>
	    By syntax complexity, we mean the number of different syntactic structures of the
	    format.
	  </t>
	  <t>
	    High transparency and low syntax complexity mostly have value in the face of extension,
	    as a fixed format doesn't need either, but even in that case, they might make the
	    overall format and its implementation simpler, which can have a lot of value by reducing
	    the risk of incompatible implementations and the likelihood that complex or opaque
	    elements open up security vulnerabilities.
	  </t>
	  <t>
	    Almost all extensible formats have a relatively high transparency and low syntax
	    complexity for their extensible part. The goal was thus to achieve a low syntax
	    complexity for the whole format, while still having high generality (i.e. that extending
	    the format without a new syntactic structure is as easy and practical as possible).
	  </t>
	  <t>
	    A good counter-example is found in most programming languages. Adding a new branching
	    construct cannot be done in a terse way without modifying the underlying
	    implementation. Such a construct either cannot be defined by user code (because of
	    evaluation rules) or can in a terribly verbose and inconvenient way (with lots of
	    boilerplate code). Notable exceptions to this limitation of programming languages are
	    Lisp languages (e.g. Common Lisp, Scheme or Clojure), languages with lazy evaluation
	    (e.g. Haskell, Purescript or Agda) and stack (or concatenative) languages (e.g. Forth,
	    Postscript or Factor).
	  </t>
	  <t>
	    On the other hand, stack languages are the canonical examples of non-transparent
	    formats. Each operator takes a number of operands from the stack. Not knowing the arity
	    of an operator makes it impossible to continue parsing, even when its evaluation was
	    optional to the final processing. In the design space, stack languages completely
	    sacrifice transparency to achieve one of the highest combination of extensibility,
	    compactness and speed of processing.
	  </t>
	  <t>
	    By generality, we mean the ability of a format to describe any type of data with a
	    reasonable (or better yet, high) level of compactness and simplicity. By analogy with
	    data structures, while both arrays and linked lists are both able to store any kind of
	    data, they actually do at the cost of complexity and transparency for arrays (they need
	    the embedding of data structure in the data or in the processing logic) and size for
	    linked lists (in-memory linked lists can waste as much as half or two third of the space
	    for the overhead of the data structure).
	  </t>
	  <t>
	    By extensibility, we mean the ability of a format to encode types and values that were
	    not anticipated when the syntax was designed.
	  </t>
	  <t>
	    By decentralization, we mean the ability to encode new types and values while avoiding
	    name collisions, but without the need of coordination. Note that the DNS, as we use it
	    (e.g. in domain names in module names in some programming languages, or in URIs in XML
	    Namespaces), is <em>not</em> decentralized in this sense, but distributed, as it cannot
	    work without its root servers and prior knowledge of their location.
	  </t>
	  <t>
	    By discoverability, we mean the ability for a processing application to automatically
	    discover new extensions when it encounters their use in data, with no prior knowledge of
	    them beforehand.
	  </t>
	  <t>
	    By compactness, we mean the ability of a format to encode as many diverse types of data
	    as possible with an overall size as small as possible.
	  </t>
	  <t>
	    By streamability, we mean two levels. The first, being streamed, is the ability of a
	    format to be transported in pieces that can be processed before the next piece is
	    available. The second, being broadcasted, is, when a stream of data is cut in two at an
	    arbitrary place, the ability of the second halt to be processed successfully.
	  </t>
	  <t>
	    By safety, we mean the ability for a format to have a parser implementing the full
	    specification while having a default behaviour that doesn't expose the system where it
	    runs to attacks triggered by malicious input.
	  </t>
	  <t>
	    By processing speed, we mean the property of a format that lends itself to benefit from
	    current computing architectures to be processed at high speed overall (e.g. processing
	    can be fast in itself, some shortcuts can be taken, or parallelization is possible).
	  </t>
	  <t>
	    By processing footprint, we mean the ability of a format to be processed while using a
	    low, possibly constrained quantity of memory.
	  </t>
	</section>
	<section title="State of the art">
	  <t>
	    Transparency, generality and extensibility are usually highly-valued traits in formats
	    design. Programming languages obviously feature them foremost, although their generality
	    usually stops at what they are supposed to express: procedures. Most of them are
	    ill-suited to represent arbitrary data, but notable exceptions include Lisp (where "code
	    is data") and Javascript, from which a subset has been extracted to exchange data, JSON,
	    which has seen a tremendous success for this purpose. JSON may have some caveats with
	    regards to generality and a relatively low compactness, but its design makes its parsing
	    really straightforward and fast. All of them, though, lack decentralization and
	    discoverability. Some of them make it possible to extend them in a distributed way if
	    some discipline is followed (for example, by naming modules after domain names), but the
	    discipline is not mandatory (and even with domain names, a change of ownership makes it
	    possible for name collisions).
	  </t>
	  <t>
	    The SGML/XML family of formats also feature good transparency, syntax complexity,
	    generality and extensibility and actually fare much better than programming languages on
	    those axes. XML namespaces also make XML naming distributed and there have been attempts
	    at making it compact (e.g. EXI from W3C, Fast Infoset from ISO/ITU or EBML).
	  </t>
	  <t>
	    All the previously cited formats clearly lack compactness, although just applying
	    standard compression techniques would sacrifice only very little processing time to gain
	    huge size reductions on most of their intended use cases, but compression may not
	    address their ineffectiveness at storing arbitrary bytes (and compression of the base64
	    encoding of arbitrary bytes can be less efficient than compression of the arbitrary
	    bytes).
	  </t>
	  <t>
	    Neither JSON nor XML are suitable for streaming and the whole document must usually be
	    parsed entirely before its content can be processed, which impacts processing speed and
	    footprint.
	  </t>
	  <t>
	    So-called binary formats pretty much exhibit the opposite trade-offs. Most of them have
	    high syntax complexity, low generality and low extensibility to achieve better
	    compactness. Some are specifically designed for a great generality, but many lack
	    extensibility. When they are extensible, it's never in a decentralized way nor are they
	    discoverable, both for reasons that have to do with compactness. They are usually
	    extremely fast to parse, and while some are designed to be streamed, few can be
	    broadcasted.
	  </t>
	  <t>
	    Actually, many binary formats are not so much formats as they are formats frameworks,
	    and exclude extensibility by design. For each use case, an IDL compiler creates a brand
	    new format that is essentially incompatible with all other formats created by the same
	    compiler (EBML specifically cites this property among its own disadvantages). If the IDL
	    compiler and framework are well designed, such a format can represent an optimum in
	    compactness and speed of processing, as the compiler can also automatically generate an
	    ad-hoc optimized parser.
	  </t>
	  <t>
	    Where extensibility has been planned in existing binary formats, it often doesn't get
	    used that much or at all because of the complications around it. Many binary formats
	    include reserved values meant to extend them to future uses, like the <tt>CM</tt> field
	    in the ZIP format. A case like this one faces an chicken-and-egg problem: if you don't
	    write and get a specification officially adopted, implementations might not want to
	    include your extension, but if your extension is purely theoretical and hasn't been
	    tested in the wild, you may face resistance to get it officially adopted. This is
	    probably why even though most compression or compressed archive formats include the
	    ability to later encode other compression methods, each new compression method usually
	    comes with its own new format.
	  </t>
	  <t>
	    When extensions are managed with any form of registry, another issue is that you usually
	    need to reserve a large set of values for free experimentation, and once an extension
	    gains any traction while in experimentation, its authors face the difficulty to switch
	    all existing implementations to the definitive values they'll get. And how experimenters
	    choose their temporary values makes them vulnerable to conflicts with
	    others. Furthermore, the process of switching between the experimental and registered
	    versions of the format or protocol might be error-prone and add a significant editorial
	    workload (<xref target="I-D.bormann-cbor-draft-numbers"/> details some of the pitfalls
	    and suggests a process to deal with this transition).
	  </t>
	</section>
	<section title="Use cases">
	  <t>
	    Here are some cases where the use of BULK formats and protocols would make software
	    engineers' and users' lives easier.
	  </t>
	  <table>
	    <thead><tr><th>Without BULK</th><th>With BULK</th></tr></thead>
	    <tbody>
	      <tr>
		<td>
		  <t>
		    If a user has a huge collection of pictures but many of them contain extended
		    metadata that her image software doesn't support yet, there is no safe and easy
		    way for her to access it.
		  </t>
		</td>
		<td>
		  <t>
		    The user's BULK image software likely is able to show the existence of different
		    metadata in every image file and, with safely auto-discovered data, can
		    visualize the extended metadata, at least in a raw form but with human-readable
		    labels or, better yet, mapped into a form it knows.
		  </t>
		</td>
	      </tr>
	      <tr>
		<td>
		  <t>
		    If a user has a collection of pictures with some in a file format her image
		    software doesn't support yet, she cannot access the kind of common metadata that
		    she can expect most images to contain (like author or copyright information).
		  </t>
		</td>
		<td>
		  <t>
		    Generic BULK tools can let the user make queries about her whole image
		    collection, and even her broader file collection, making use of either common
		    metadata formats between different image or file formats, or the mapping between
		    different metadata formats into a single one. Queries could be for all objects,
		    whole files or entries inside files, that have been flagged "confidential" or
		    belong to some entity, or all images or videos where some person has been tagged
		    as visible, for example.
		  </t>
		</td>
	      </tr>
	      <tr>
		<td>
		  <t>
		    If a new image or video compression has been designed, its designers usually
		    need to create a whole custom container format, with custom metadata format. A
		    lot of work is needed for the many image or video software to be able to
		    accommodate reading or writing this new container, new metadata and new
		    codec. Because this new work involves parsing binary data, it often is a source
		    of security vulnerabilities.
		  </t>
		</td>
		<td>
		  <t>
		    If a new image or video compression has been designed, its designers don't need
		    to create anything more to embed it in BULK. If it has unusual features, it is
		    still pretty easy to make what is backward-compatible with existing data models
		    fit into the existing container structure. Even new kinds of structures leverage
		    existing parsing code, minimizing attack surface.
		  </t>
		</td>
	      </tr>
	      <tr>
		<td>
		  <t>
		    If a new compression, signing or encryption algorithm has been designed, and its
		    designers hope to see it used in existing formats, they have to work separately
		    for each format where it may be used and then for each software project
		    implementing them and in each case, they will face a chicken-and-egg problem
		    where software projects may be reluctant to put effort into something users may
		    not ultimately want or benefit from. The designers may also need to interact
		    with one or several registries to get their algorithm registered, which may be a
		    necessary upfront work. Users need to wait for each software to be updated to
		    include the new algorithm and sometimes suffer from incompatible implementations
		    at the format level (e.g. early adopters producing files with obsolete numbers
		    allocated to the algorithm).
		  </t>
		</td>
		<td>
		  <t>
		    If a new compression, signing or encryption algorithm has been designed, and its
		    designers hope to see it used in existing BULK formats, their BULK vocabulary is
		    guaranteed to uniquely describe the data produced by their algorithm and can be
		    readily embedded in any existing BULK format. The only work needed may be to
		    implement their algorithm as a pure function that can be plugged into the BULK
		    evaluation mechanism, for the relevant programming languages, so that software
		    projects just have to register the link between BULK vocabulary and software
		    plugin (or may not even need to intervene, if a safe mechanism is implemented to
		    auto-discover software plugins providing pure functions).
		  </t>
		</td>
	      </tr>
	      <tr>
		<td>
		  <t>
		    If a user encounters a file of an unknown type, their system might give them
		    useful media type identification or not, and based on that or file extension, an
		    advanced user might search in her system's software registry or on the Internet
		    and find some software that can visualize some or all of the file's content.
		  </t>
		</td>
		<td>
		  <t>
		    If a user encounters a BULK file of a yet unknown type, generic BULK tools can
		    safely auto-discover human-readable information about the structure of the file,
		    and parts of the file that the user can readily use can be easily found and
		    viewed or extracted.
		  </t>
		</td>
	      </tr>
	      <tr>
		<td>
		  <t>
		    When a communication protocol is designed, its designers have to choose a
		    trade-off between latency caused by message size, ease of debugging and
		    complexity of the parser. It's deceptively easy to make mistakes in the syntax
		    that make it harder to use or implement, but using an existing syntax ties the
		    protocol with this syntax' limitations.
		  </t>
		</td>
		<td>
		  <t>
		    When a BULK-based communication protocol is designed, its designers can make it
		    both very compact and easy to debug with very little effort. Generic tools can
		    present data on the wire with the support of safely auto-discovered data about
		    the protocol.
		  </t>
		</td>
	      </tr>
	    </tbody>
	  </table>
	  <t>
	    In a few of those cases, without BULK, a registry of executable software patches or
	    plugins that can be automatically discovered could have been a solution, with the caveat
	    that almost all of our current software architectures make that possibility more
	    dangerous than it's worth (because of the broad ambient authority given to most code).
	  </t>
	  <t>
	    With BULK, there is immediately a strong incentive to provide discoverability, at first
	    for just data, not executable code, with safeties already in place on that process. The
	    data model of BULK then makes it more likely that code plugins are provided that are
	    designed to work with extremely limited privileges, to act as transformers of BULK
	    expressions.
	  </t>
	</section>
      </section>
      <section title="Format overview">
	<t>
	  A BULK stream is a stream of 8-bit bytes, in big-endian order. Parsing a BULK stream
	  yields a sequence of expressions, which can be either atoms or forms, which are sequences
	  of expressions.
	</t>
	<t>
	  Forms have a simple syntax: a starting byte marker, a sequence of expressions and an
	  ending byte marker.
	</t>
	<t>
	  Atoms each have a special syntax, for compactness purposes: they start with a marker byte,
	  followed by a static or dynamic number of bytes, depending on the type. But there are only
	  5 kinds: the nil atom, generic arrays, small arrays, small unsigned integers and
	  references.
	</t>
	<t>
	  Even booleans and floating-point numbers use this existing syntax without the need for
	  special cases.
	</t>
	<t>
	  References consist of a namespace marker (in almost all cases, a single byte) followed by
	  an identifier within this namespace (a single byte). All in all, a very little sacrifice
	  is made in compactness for the benefit of a very simple and resilient syntax: apart from
	  nil and small integers, nothing is smaller than 2 bytes, and as most forms involve a
	  reference followed by some content, a form is usually 4 bytes + its content.
	</t>
	<t>
	  A namespace marker in a BULK stream is associated to a namespace identified by some
	  identifier guaranteed to be unique without coordination (like a UUID or cryptographical
	  hash), thus ensuring decentralized extensibility. The stream can be processed even if the
	  application doesn't recognize the namespace. Parsing remains possible thanks to the
	  transparent syntax.
	</t>
	<t>
	  Combination of BULK namespaces, BULK streams and even other formats doesn't need any
	  content transformation to work. Here are some examples:
	  <list style="symbols">
	    <t>
	      The content of a BULK stream, enclosed in form starting and ending byte markers,
	      constitute a valid BULK expression. Thus BULK streams can be aggregated or annotated
	      within a BULK stream without modification.
	    </t>
	    <t>
	      A BULK format could specify in its syntax the place for a metadata expression. Whether
	      the specification provides its own metadata forms or not, an application could use a
	      BULK serialization for MARC, TEI Header, XML or RDF for this metadata expression. The
	      vocabulary selected would be univocally expressed by the namespace and every
	      vocabulary would be parsed by the same mechanisms.
	    </t>
	    <t>
	      Whenever a content must be stored as-is instead of serialized, or a highly-optimized
	      ad hoc serialization exists for some data, anything can always be stored within an
	      array. They can contain arbitrary bytes and there is no limit to their size.
	    </t>
	  </list>
	</t>
	<t>
	  Furthermore, BULK expressions can be evaluated. Many expressions evaluate to themselves,
	  but others evaluate to the result of executing a pure function, making it possible to
	  serialize data in an even more compact form, by eliminating boilerplate data and repeated
	  patterns.
	</t>
      </section>
      <section title="Conventions and Terminology">
        <t>
          The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD
          NOT", "RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as
          described in <xref target="BCP14"/>.
        </t>
        <t>
          Literal numerical values are provided in decimal or hexadecimal as appropriate.
          Hexadecimal literals are prefixed with <tt>0x</tt> to distinguish them from decimal
          literals.
        </t>
	<t>
	  The text notation of the BULK stream uses mnemonics for some bytes sequences. Mnemonics
	  are sequences of characters, excluding all capital letters and white space, like
	  <tt>this-is-one-mnemonic</tt> or <tt>what-the-%§!?#-is-that?</tt>. They are always
	  separated by white space. Outside the use of mnemonics, a sequence of bytes (of one or
	  more bytes) can be represented by its hexadecimal value as an unsigned integer prefixed by
	  <tt>0x</tt> (e.g. <tt>0x3F</tt> or <tt>0x3A0B770F</tt>). Such a sequence of bytes can
	  include dashes to make it more readable
	  (e.g. <tt>0xDDA37D36-85E6-4E6D-9B51-959E1CCE366C</tt>). Some types in this specification
	  define a special syntax for their representation in the text notation.
	</t>
	<t>
	  In the grammar, a shape is a pattern of bytes, following the rules of the text notation
	  for a BULK stream. Apart from mnemonics and fixed sequences of bytes, a shape can contain:
	  <list style="symbols">
	    <t>
	      an arbitrary sequence of a fixed number of bytes, represented by its size, i.e. a
	      number of bytes in decimal immediately followed by a B uppercase letter
	      (e.g. <tt>4B</tt>)
	    </t>
	    <t>
	      a typed sequence of bytes, represented by the name of its type, a capitalized word
	      (e.g.  <tt>Foo</tt>); this means a sequence of bytes whose specific yield (see <xref
	      target="parsing" format="title"/>) has this type
	    </t>
	    <t>
	      a named sequence of bytes (of zero or more bytes), represented by a sequence of any
	      character excluding '{}' between '{' and '}' (e.g. <tt>{quux}</tt>); a named sequence
	      can be typed or sized, in which case it is immediately followed by ':' and a type or
	      size (e.g. <tt>{quux}:Bar</tt> or <tt>{quux}:12B</tt>)
	    </t>
	  </list>
	</t>
	<t>
	  The shape that describes the byte sequence of an atom is called its parsing shape. When a
	  shape is given for a form, it merely describes the semantics of evaluating forms of that
	  shape. A reference used in such a shape can be used in different shapes, with unrelated
	  semantics.
	</t>
	<t>
	  For example, this specification defines a way do encode a string with explicit encoding
	  with forms of the shape <tt>( string {enc}:Expr {string}:Expr )</tt>. But the shapes <tt>(
	  string {arg1}:Int {arg2}:Int )</tt> or <tt>( {arg1}:Int string {arg2}:Int )</tt> are
	  syntactically valid. They just evaluate to themselves as lists of three expressions, as
	  far as this specification is concerned.
	</t>
      </section>
    </section>

    <section title="BULK syntax">
      <t>
	A BULK stream is a sequence of 8-bit bytes. Bits and bytes are in big-endian order. The
	result of parsing a BULK stream is a list of abstract data, called the abstract yield. BULK
	parsing is injective: a BULK stream has only one abstract yield, but different BULK streams
	can have the same abstract yield (if they associate namespaces to different markers, see
	<xref target="nss-pkgs" format="title"/>).
      </t>
      <t>
	A processing application is not expected to actually produce the abstract yield, but an
	adaptation of the abstract yield to its own implementation, called the concrete yield. Also,
	some expressions in a BULK stream may have the semantics of a transformation of the abstract
	yield. A processing application MAY thus not produce or retain the concrete yield but the
	result of its transformation. This specification deals mainly with the byte sequence and the
	abstract yield and occasionally provides guidelines about the concrete yield. Of course, a
	processing application MAY not produce any concrete yield at all but produce various data
	structures and side effects from parsing the BULK stream (for example, an event sourced
	application may read its event log from a BULK stream and build its application state by
	applying the events, discarding each of them as soon as it has been applied).
      </t>
      <t>
	The abstract yield is a list of expressions. Expressions can be atoms or forms. Forms are
	lists of expressions. If a byte sequence is parsed as an expression, this byte sequence is
	said to encode this expression.
      </t>
      <t>
	When a sequence of bytes is named in a shape, its name can be used in this specification to
	designate either the byte sequence, the expression or sequence of expressions it encodes, or
	the result of evaluating those expressions. When there could be ambiguity, this
	specification specifies which is designated.
      </t>

      <section anchor="parsing" title="Parsing algorithm">
	<t>
	  The parser operates with a context, which is a list of expressions. Each time an
	  expression is parsed, it is appended at the end of the context. The initial context is the
	  abstract yield.
	</t>
	<t>
	  At the beginning of a BULK stream and after having consumed the byte sequence encoding a
	  complete expression, the parser is at the dispatch stage. At this stage, the next byte is
	  a marker byte, which tells the parser what kind of expression comes next (the marker byte
	  is the first byte of the sequence that encodes an expression). The expression appended to
	  the context after reading a byte sequence is called the specific yield of the byte
	  sequence.
	</t>
	<t>
	  The <tt>0x01</tt> and <tt>0x02</tt> marker bytes are special cases. When the parser reads
	  <tt>0x01</tt>, it immediately appends an empty list to the current context. This list
	  becomes the new context. This new context has the previous context as parent. Then the
	  parser returns to its dispatch stage. When the parser reads <tt>0x02</tt>, it appends
	  nothing to the context, but instead the parent of the current context becomes the new
	  context and the parser returns to the dispatch stage. Thus it is a parsing error to read
	  <tt>0x02</tt> when the context is the abstract yield. It is also a parsing error when the
	  stream ends and the parser is not at the dispatch stage in the abstract yield.
	</t>
	<t>
	  When a form contains at least one expression, the first expression is called the operator
	  of that form. When a form contains more than one expression, all but the first expression
	  are called the operands of that form.
	</t>
	<t>
	  Some forms have side-effects in their semantics. Those side-effects MUST NOT affect the
	  parsing of any expression. They can affect evaluation, in which case they MUST only affect
	  the evaluation of expressions in the scope of the form. The scope of an expression is the
	  part of its context that follows the expression. This makes BULK lexically scoped.
	</t>
	<t>
	  Whenever a parsing error is encountered, parsing of the BULK stream MUST stop.
	</t>
	<t>
	  The version of the BULK stream can affect parsing, see <xref target="version"/>.
	</t>
	<section title="Summary of marker bytes">
	  <table>
	    <thead><tr><th>marker</th><th>shape</th><th></th></tr></thead>
	    <tbody>
	      <tr><td><tt>00</tt></td><td><xref target="nil" format="none"><tt>nil</tt></xref></td><td></td></tr>
	      <tr><td><tt>01</tt></td><td><xref target="form" format="none"><tt>(</tt></xref></td><td></td></tr>
	      <tr><td><tt>02</tt></td><td><xref target="form" format="none"><tt>)</tt></xref></td><td></td></tr>
	      <tr><td><tt>03</tt></td><td><xref target="array" format="none"><tt># Nat {content}</tt></xref></td><td></td></tr>
	      <tr><td><tt>04–0F</tt></td><td></td><td><xref target="reserved" format="none">reserved</xref></td></tr>
	      <tr><td><tt>10–7F</tt></td><td><xref target="ref" format="none"><tt>Ref</tt></xref></td><td></td></tr>
	      <tr><td><tt>80–BF</tt></td><td><xref target="smallint" format="none"><tt>w6[value]</tt></xref></td><td></td></tr>
	      <tr><td><tt>C0–FF</tt></td><td><xref target="smallarray" format="none"><tt>#[size] {content}</tt></xref></td><td></td></tr>
	    </tbody>
	  </table>
	</section>
	<section anchor="eval" title="Evaluation">
	  <t>
	    A processing application MAY implement evaluation of BULK expressions and streams. When
	    evaluating a BULK stream, when the parser gets to the dispatch stage and the context is
	    the abstract yield (this is called a <em>dispatch point</em>), the last expression in
	    the context is replaced by what it evaluates to. (of course, this description is
	    supposed to provide the semantics of BULK evaluation, but a processing application MAY
	    implement evaluation with a different algorithm as long as it provides the same
	    semantics)
	  </t>
	  <t>
	    Evaluation cannot affect parsing, which means that parsing and evaluation can be done in
	    parallel.
	  </t>
	  <t>
	    The default evaluation rule is that an expression evaluates to itself. A name within a
	    namespace can have a value, which is what a reference associated to this name evaluates
	    to. A reference whose marker value is associated to no namespace or whose name has no
	    value evaluates to itself. How self-evaluating BULK expressions are represented in the
	    concrete yield is application-dependent, but future specifications MAY define a standard
	    API to access it, similar to the Document Object Model for XML.
	  </t>
	  <t>
	    The evaluation of a form obeys a special rule, though: if the operator of the form has
	    type <tt>Function</tt>, that function is called with an argument list and the form
	    evaluates to the return value if it's an atom or the evaluation of the return value if
	    it is a form. If the function has type <tt>LazyFunction</tt>, the argument list is the
	    operands of the form. If the function has type <tt>EagerFunction</tt>, the argument list
	    is the result of evaluating the operands of the form, from left to right. Any expression
	    that has type <tt>LazyFunction</tt> or <tt>EagerFunction</tt> also has type
	    <tt>Function</tt>.
	  </t>
	  <t>
	    In the abstract yield, if the first expression produced as a value by evaluation has
	    type <tt>EagerFunction</tt>, then after the whole abstract yield has been evaluated, the
	    resulting sequence of expressions MAY be evaluated as a form. This is called whole
	    stream evaluation, and MAY be a configuration option. In particular, a processing
	    application MAY choose to disable whole stream evaluation during <xref
	    target="streaming">streaming</xref>.
	  </t>
	  <t>
	    When this specification describes the evaluation of a form starting with a
	    <tt>LazyFunction</tt>, by default, named shapes designate the unevaluated
	    expressions. When this specification describes the evaluation of a form starting with an
	    <tt>EagerFunction</tt>, by default, named shapes designate the result of evaluating
	    expressions.
	  </t>
	  <t>
	    A form whose operand doesn't have the type <tt>Function</tt> evaluates to a form
	    containing the result of evaluating each expression of the form, from left to right.
	  </t>
	  <t>
	    When an application evaluates a BULK expression, it MUST verify that evaluation
	    terminates in a finite number of evaluation steps. An application MAY verify finite
	    termination statically or dynamically. For example, an application MAY stop evaluation
	    in error after a predetermined number of steps.
	  </t>
	  <t>
	    Whenever this specification describes the semantics of an expression, it describes the
	    result of evaluating that expression, either in terms of the value returned or the
	    side-effects executed by evaluation of the expression.
	  </t>
	  <t>
	    When an evaluation error is encountered, a processing application MAY not stop
	    processing with an error. If the processing application produces a partially evaluated
	    concrete yield, it MUST convey enough information for the using agent to know the order
	    of all expressions in the original BULK stream, which expressions were successfully
	    evaluated, and which were not because of evaluation errors. A processing application MAY
	    stop evaluation at the first evaluation error and produce the concrete yield in two
	    separated sections, the successfully evaluated part, followed by the unevaluated one. A
	    processing application MAY continue evaluation and produce the concrete yield as a
	    sequence of expressions, each tagged with the fact that it was successfully evaluated or
	    not.
	  </t>
	  <t>
	    The version of the BULK stream can affect evaluation, see <xref target="version"/>.
	  </t>
	</section>
      </section>

      <section anchor="form" title="Forms">
	<t>
	  <list style="hanging">
	    <t hangText="starting marker"><tt>0x01</tt><br/>mnemonic: <tt>(</tt></t>
	    <t hangText="ending marker"><tt>0x02</tt><br/>mnemonic: <tt>)</tt></t>
	  </list>
	</t>

	<section title="Difference between sequence and form">
	  <t>
	    There is a difference between a byte sequence encoding several expressions among the
	    current context and a byte sequence encoding a form (i.e. a single expression that is a
	    list of expressions). As an example, let's examine several forms of the shape <tt>( foo
	    {bar} )</tt>.
	  </t>
	  <t>
	    <list style="symbols">
	      <t>
		In the form <tt>( foo nil nil nil )</tt>, <tt>{bar}</tt> encodes 3 expressions, and
		they are three atoms in the yield.
	      </t>
	      <t>
		In the form <tt>( foo nil )</tt>, <tt>{bar}</tt> is a single expression in the
		yield, and that expression is an atom.
	      </t>
	      <t>
		In the form <tt>( foo ( nil nil nil ) )</tt>, <tt>{bar}</tt> is also a single
		expression in the yield, and that expression is a form, a list in the yield.
	      </t>
	    </list>
	  </t>
	  <t>
	    In a shape, when a byte sequence must yield a single expression, it has the type
	    <tt>Expr</tt>. So the last two examples fit the shape <tt>( foo {bar}:Expr )</tt> but
	    not the first.
	  </t>
	</section>
      </section>

      <section title="Atoms">
	<section anchor="nil" title="nil">
	  <t>
	    <list style="hanging">
	      <t hangText="marker"><tt>0x00</tt><br/>mnemonic: <tt>nil</tt></t>
	      <t hangText="parsing shape"><tt>nil</tt></t>
	    </list>
	  </t>
	  <t>
	    Apart from being a possible short marker value, the fact that the <tt>0x00</tt> byte
	    represents a valid atom means that a sequence of null bytes is a valid part of a BULK
	    stream, thus making the format less fragile. In a network communication, nil atoms can
	    be sent to keep the channel open. They can also be used as padding at the end of a form
	    or between forms.
	  </t>
	</section>

	<section title="Arrays">
	  <t>
	    Arrays can be used to store arbitrary bytes.
	  </t>
	  <t>
	    An array can be interpreted either as a bits sequence or as an unsigned integer in
	    binary notation. The choice depends on the context and the application. Actually, many
	    processing applications may not need make any choice, as most programming language
	    implementations actually also confuse unsigned integers and bits sequences to some
	    extent. Expressions that are unsigned integers (that is, natural numbers) have type
	    <tt>Nat</tt> (whether they are encoded as an array or not).
	  </t>
	  <t>
	    Big arrays typically store the content of a file or a binary message of another
	    format. They can also be used to store a vector or matrix of fixed-size elements.
	  </t>
	  <t>
	    In any case, the semantics of the content must be inferred by the processing
	    application; where ambiguity can appear, an application SHOULD enclose the array in a
	    form that makes the semantics explicit (e.g. <xref target="string"
	    format="none"><tt>string</tt></xref>, <xref target="blob"
	    format="none"><tt>blob</tt></xref>, or <xref target="unsigned-int"
	    format="none"><tt>unsigned-int</tt></xref>).
	  </t>
	  <t>
	    Because BULK arrays have no end markers, the payload of a BULK array can constitute the
	    end of the stream.
	  </t>
	  <t>
	    The start and end of an array are known without reading its content, which means that
	    its content can be skipped in constant time and mapped in memory (or read lazily by any
	    other means).
	  </t>
	  <t>
	    Because BULK can use integers with arbitrary size to store the size of an array, BULK
	    arrays have no limit in size.
	  </t>
	  <t>
	    Any array also has the type <tt>String</tt> if its contents can be decoded as a string
	    in the current encoding.
	  </t>
	  <t>
	    When this specification mentions "the bytes contained in the expression <tt>{foo}</tt>",
	    it never means the bytes encoding that BULK expression, but the bytes that are the
	    payload of that expression, which could be an array or some other byte container defined
	    in a BULK vocabulary.
	  </t>
	  
	  <section anchor="array" title="Generic array">
	    <t>
	      <list style="hanging">
		<t hangText="marker"><tt>0x03</tt><br/>mnemonic: <tt>#</tt></t>
		<t hangText="parsing shape"><tt># Nat {content}</tt></t>
	      </list>
	    </t>
	    <t>
	      After consuming the marker byte, the parser returns to the dispatch stage. It is a
	      parsing error if the parsed expression is not of type <tt>Nat</tt> or if its value
	      cannot be recognized. This integer is not added to any context, but the parser
	      consumes as many bytes as this integer and they constitute the content of this array.
	    </t>
	    <t>
	      In the text notation, a quoted string is the notation for a generic array containing
	      the encoding of that string in the <xref target="stringenc" format="none">current
	      encoding</xref>, except if the size of the encoding is below 64 bytes, cf. <xref
	      target="smallarray" format="none">small arrays</xref>.
	    </t>
	    <t>
	      In the text notation, some text notation enclosed between balanced <tt>([</tt> and
	      <tt>])</tt> is the notation for a generic array containing the encoding of that text
	      notation, except if the size of the encoding is below 64 bytes, cf. <xref
	      target="smallarray" format="none">small arrays</xref>.
	    </t>
	    <t>Types: <tt>Bytes</tt>, <tt>Nat</tt></t>
	  </section>

	  <section anchor="smallarray" title="Small array">
	    <t>
	      <list style="hanging">
		<t hangText="marker"><tt>0xC0–0xFF</tt><br/>mnemonic: <tt>#[size]</tt></t>
		<t hangText="parsing shape"><tt>#[size] {content}</tt></t>
	      </list>
	    </t>
	    <t>
	      The 6 least significant bits of the marker byte are treated as an unsigned
	      integer. This integer is not added to any context, but the parser consumes as many
	      bytes as this integer and they constitute the content of this array.
	    </t>
	    <t>
	      In the text notation, the notation of the marker byte of a small array of size X is
	      <tt>#[X]</tt>. For example, <tt>#[2] 0x1234</tt> is a notation for the bytes
	      <tt>0xC2-1234</tt>.
	    </t>
	    <t>
	      In the text notation, a quoted string is the notation for a small array containing the
	      encoding of that string in the current encoding if the size of the encoding is below
	      64 bytes. For example, <tt>"abc"</tt> and <tt>#[3] 0x616263</tt> are the notation for
	      the same byte sequence if the current encoding is UTF-8.
	    </t>
	    <t>
	      In the text notation, some text notation enclosed between balanced <tt>([</tt> and
	      <tt>])</tt> is the notation for a small array containing the encoding of that text
	      notation if the size of the encoding is below 64 bytes. For example, <tt>([ nil 0 1
	      256 ])</tt> and <tt>#[6] nil w6[0] w6[1] #[2] 0x0100</tt> are the notation for the
	      same byte sequence.
	    </t>
	    <t>Types: <tt>Bytes</tt>, <tt>Nat</tt></t>
	  </section>

	  <section anchor="smallint" title="Small unsigned integers">
	    <t>
	      <list style="hanging">
		<t hangText="marker"><tt>0x80–0xBF</tt><br/>mnemonic: <tt>w6[value]</tt></t>
		<t hangText="parsing shape"><tt>w6[value]</tt></t>
	      </list>
	    </t>
	    <t>
	      The 6 least significant bits of the marker byte are the value encoded by this byte (as
	      bits or as an unsigned integer in binary notation).
	    </t>
	    <t>
	      In the text notation, the notation of the marker byte of a small unsigned integer of
	      value X is <tt>w6[X]</tt>. For example, <tt>w6[11]</tt> is a notation for the byte
	      <tt>0x8B</tt> (as is <tt>11</tt>, cf. <xref target="arithmetic"/>).
	    </t>
	    <t>Types: <tt>Bytes</tt>, <tt>Nat</tt></t>
	  </section>

	  <section anchor="nat-enc" title="Encoding natural numbers">
	    <t>
	      When the syntax of a BULK form mandates that an expression can only be a <tt>Nat</tt>,
	      an application SHOULD encode it as the smallest possible array using one of the
	      following sizes: 6, 8, 16, 32, or any multiple of 64 bits.
	    </t>
	  </section>
	</section>

	<section anchor="reserved" title="Reserved marker bytes">
	  <t>
	    Marker bytes <tt>0x04−0x0F</tt> are reserved for future major versions of BULK. It is a
	    parsing error if a BULK stream with version 1.0 contains such a marker byte (see <xref
	    target="version"/>).
	  </t>
	</section>

	<section anchor="ref" title="References">
	  <t><list style="hanging">
	    <t hangText="marker"><tt>0x10−0x7F</tt></t>
	    <t hangText="parsing shape">
	      <tt>{ns}:1B {name}:1B</tt>
	      <vspace/>
  	      <tt>0x7F {ns'} {name}:1B</tt>
	    </t>
	  </list>
	  </t>
	  <t>
	    The <tt>{ns}</tt> byte is a value associated with a namespace, called the namespace
	    marker. Values <tt>0x10−0x13</tt> are reserved for standard namespaces defined by BULK
	    specifications. Greater values can be associated with namespaces identified by a unique
	    identifier.
	  </t>
	  <t>
	    The <tt>{name}</tt> byte is the name index within the namespace. Vocabularies with more
	    than 256 names thus need to be spread across several namespaces.
	  </t>
	  <t>
	    The specification of a namespace SHOULD include a mnemonic for the namespace and for
	    each defined name. When descriptions use several namespaces, the mnemonic of a reference
	    SHOULD be the concatenation of the namespace mnemonic, ":" and the name mnemonic if
	    there can be an ambiguity. For example, the <tt>fft</tt> name in namespace <tt>math</tt>
	    becomes <tt>math:fft</tt>.
	  </t>
	  <t>Type: <tt>Ref</tt></t>
	  <section title="Special case">
	    <t>
	      References have a second parsing rule. In case a BULK stream needs an important number
	      of namespaces, if the marker byte is <tt>0x7F</tt>, the parser continues to read bytes
	      until it finds a byte different than 0xFF. The sum of each of those bytes taken as
	      unsigned integers is the namespace marker. For example, the reference encoded by the
	      bytes <tt>0x7F 0xFF 0x8C 0x1A</tt> is the name 26 in the namespace associated with
	      522.
	    </t>
	  </section>
	</section>
      </section>
    </section>

    <section anchor="kinds" title="Kinds of namespaces">
      <t>
	Standard namespaces have a fixed marker value and are not identified by a unique
	identifier.
      </t>
      <t>
	Standard namespaces are immutable. It is an evaluation error when the reference in a name
	definition is in a standard namespace.
      </t>
      <t>
	Extension namespaces are defined with a unique identifier, to be associated to a marker
	value.
      </t>
      <t>
	By its decentralized nature, as far as a processing application is concerned, while standard
	namespaces are a special case, there is no difference between an extension namespace defined
	as part of the official BULK suite and any other one.
      </t>
    </section>

    <section title="BULK core namespace">
      <t>
	<list style="hanging">
	  <t hangText="marker"><tt>0x10</tt><br/>namespace mnemonic: <tt>bulk</tt></t>
	</list>
      </t>
      <table>
	<thead><tr><th>name</th><th>mnemonic</th><th>type</th></tr></thead>
	<tbody>
	  <tr><td><tt>00</tt></td><td><spanx style="verb"><xref target="version"
	  format="none">version</xref></spanx></td><td><tt>LazyFunction</tt></td></tr>
	  <tr><td><tt>01</tt></td><td><tt>import</tt> (<xref target="import-ns"
	  format="none">namespace</xref>, <xref target="import-pkg"
	  format="none">package</xref>)</td><td><tt>LazyFunction</tt></td></tr>
	  <tr><td><tt>02</tt></td><td><tt>namespace</tt></td><td/></tr>
	  <tr><td><tt>03</tt></td><td><tt>package</tt></td><td/></tr>
	  <tr><td><tt>04</tt></td><td><spanx style="verb">define (<xref target="def-ns"
	  format="none">namespace</xref>, <xref target="def-pkg" format="none">package</xref>, <xref
	  target="def-name"
	  format="none">name</xref>)</spanx></td><td><tt>LazyFunction</tt></td></tr>
	  <tr><td><tt>05</tt></td><td><spanx style="verb"><xref target="mnemonic"
	  format="none">mnemonic</xref></spanx></td><td><tt>LazyFunction</tt></td></tr>
	  <tr><td><tt>06</tt></td><td><spanx style="verb"><xref target="explain"
	  format="none">explain</xref></spanx></td><td><tt>LazyFunction</tt></td></tr>
	  <tr><td><tt>07</tt></td><td><spanx style="verb"><xref target="string"
	  format="none">string</xref></spanx></td><td><tt>EagerFunction</tt></td></tr>
	  <tr><td><tt>08</tt></td><td><spanx style="verb"><xref target="iana-charset"
	  format="none">iana-charset</xref></spanx></td><td><tt>EagerFunction</tt></td></tr>
	  <tr><td><tt>09</tt></td><td><spanx style="verb"><xref target="nested-bulk"
	  format="none">bulk</xref></spanx></td><td><tt>EagerFunction</tt></td></tr>
	  <tr><td><tt>0A</tt></td><td><spanx style="verb"><xref target="blob"
	  format="none">blob</xref></spanx></td><td><tt>EagerFunction</tt></td></tr>
	  <tr><td><tt>0B</tt></td><td><spanx style="verb"><xref target="concat"
	  format="none">concat</xref></spanx></td><td><tt>EagerFunction</tt></td></tr>
	  <tr><td><tt>0C</tt></td><td><spanx style="verb"><xref target="indexed-bulk"
	  format="none">indexed-bulk</xref></spanx></td><td><tt>EagerFunction</tt></td></tr>
	  <tr><td><tt>0D</tt></td><td><spanx style="verb"><xref target="indexed-array"
	  format="none">indexed-array</xref></spanx></td><td><tt>EagerFunction</tt></td></tr>
	  <tr><td><tt>0E</tt></td><td><spanx style="verb"><xref target="booleans"
	  format="none">true</xref></spanx></td><td><tt>Boolean</tt></td></tr>
	  <tr><td><tt>0F</tt></td><td><spanx style="verb"><xref target="booleans"
	  format="none">false</xref></spanx></td><td><tt>Boolean</tt></td></tr>
	  <tr><td><tt>10</tt></td><td><spanx style="verb"><xref target="subst"
	  format="none">subst</xref></spanx></td><td><tt>LazyFunction</tt></td></tr>
	  <tr><td><tt>11</tt></td><td><spanx style="verb"><xref target="arg"
	  format="none">arg</xref></spanx></td><td/></tr>
	  <tr><td><tt>12</tt></td><td><spanx style="verb"><xref target="rest"
	  format="none">rest</xref></spanx></td><td/></tr>
	  <tr><td><tt>13</tt></td><td><spanx style="verb"><xref target="values"
	  format="none">values</xref></spanx></td><td><tt>LazyFunction</tt></td></tr>
	  <tr><td><tt>14</tt></td><td><spanx style="verb"><xref target="unsigned-int"
	  format="none">unsigned-int</xref></spanx></td><td><tt>EagerFunction</tt></td></tr>
	  <tr><td><tt>15</tt></td><td><spanx style="verb"><xref target="signed-int"
	  format="none">signed-int</xref></spanx></td><td><tt>EagerFunction</tt></td></tr>
	  <tr><td><tt>16</tt></td><td><spanx style="verb"><xref target="fraction"
	  format="none">fraction</xref></spanx></td><td><tt>EagerFunction</tt></td></tr>
	  <tr><td><tt>17</tt></td><td><spanx style="verb"><xref target="binary-float"
	  format="none">binary-float</xref></spanx></td><td><tt>EagerFunction</tt></td></tr>
	  <tr><td><tt>18</tt></td><td><spanx style="verb"><xref target="decimal-float"
	  format="none">decimal-float</xref></spanx></td><td><tt>EagerFunction</tt></td></tr>
	  <tr><td><tt>19</tt></td><td><spanx style="verb"><xref target="prefix"
	  format="none">prefix</xref></spanx></td><td><tt>LazyFunction</tt></td></tr>
	  <tr><td><tt>1A</tt></td><td><spanx style="verb"><xref target="postfix"
	  format="none">postfix</xref></spanx></td><td><tt>LazyFunction</tt></td></tr>
	  <tr><td><tt>1B</tt></td><td><spanx style="verb"><xref target="arity"
	  format="none">arity</xref></spanx></td><td><tt>LazyFunction</tt></td></tr>
	</tbody>
      </table>

      <section anchor="version" title="Version">
	<t>
	  <list style="hanging">
	    <t hangText="shape"><tt>( version {major}:Nat {minor}:Nat )</tt></t>
	  </list>
	</t>
	<t>
	  When parsing a BULK stream, a processing application MUST determine explicitly the major
	  and minor version of the BULK specification that the stream obeys. This information MAY be
	  exchanged out-of-band, if BULK is used to exchange a number a very small messages, where
	  repeated headers of 6 bytes might become too big an overhead. A processing application
	  MUST NOT assume a default version.
	</t>
	<t>
	  If the version is expressed within a BULK stream, this form MUST be the first in the
	  stream, where its semantics is the side-effect of declaring the version. In any other
	  place, this form evaluates to itself. This specification defines BULK 1.0. When writing a
	  BULK stream's version, an application MUST encode <tt>{major}</tt> and <tt>{minor}</tt> by
	  the smallest byte sequence as described in <xref target="nat-enc"></xref>.
	</t>
	<t>
	  An application writing a BULK stream to long-term storage (e.g. in a file or a database
	  record) SHOULD include a <tt>version</tt> form.
	</t>
	<t>
	  Two BULK versions with the same major version MUST share the same parsing rules and the
	  same definitions of marker bytes used by both. Changing the syntax or semantics of
	  existing marker bytes warrants a new major version. Changing the syntax or semantics of
	  existing standard names (meaning names in standard namespaces) also warrants a new major
	  version.
	</t>
	<t>
	  It is a parsing error when a processing application encounters a major version that it
	  doesn't explicitly support.
	</t>
	<t>
	  Using marker bytes in the reserved interval, adding standard names, or adding new
	  syntactic uses of existing standard names that don't overlap with existing uses warrants a
	  new minor version.
	</t>
	<t>
	  If version A and version B of BULK have a different default profile, and the change of
	  profile would only affect evaluation of BULK streams that use new marker bytes, new
	  standard names or new syntactic uses of existing standard names, then the profile
	  difference warrants a difference of minor version between A and B. Otherwise, the profile
	  change warrants a difference of major version between A and B.
	</t>
	<t>
	  Overall, the goal is that if an application writes a BULK stream of version 1.2 that only
	  contains syntactic and semantic elements from BULK 1.1, a processing application that only
	  supports BULK 1.1 will be able to parse and evaluate that stream. The writing application
	  doesn't need to know the features that are specific to BULK 1.0 and BULK 1.1, because as
	  soon as the processing application encounters a reserved marker byte, an unknown standard
	  name or an unknown syntactic use of a known standard name, it can detect that it cannot
	  correctly evaluate the stream.
	</t>
	<t>
	  For that reason, it is an evaluation error for a processing application when there is an
	  unknown syntactic use of a standard name, in a BULK stream with a major version supported
	  by the processing application but a minor version not supported.
	</t>
	<t>
	  As far as BULK 1.0 is concerned, using <tt>import</tt>, <tt>define</tt>,
	  <tt>mnemonic</tt>, <tt>explain</tt>, and <tt>arity</tt>, when none of their operands is
	  either a standard name or a form with a standard name as operator, or using
	  <tt>version</tt> somewhere else than as first expression, don't constitute an unknown
	  syntactic use. This makes it possible to overload those names in ways that don't interfere
	  with forward-compatibility.
	</t>
      </section>

      <section anchor="nss-pkgs" title="Namespaces and packages">
	<section anchor="import-ns" title="Importing a namespace">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( import {marker}:Nat ( namespace {id}:Expr ) )</tt></t>
	    </list>
	  </t>
	  <t>
	    The semantics of this form is the side-effect of associating the namespace identified by
	    <tt>{id}</tt> to the namespace marker <tt>{marker}</tt>, within the scope of this
	    expression.
	  </t>
	  <t>
	    <tt>{id}</tt> can be a form using one or several names whose namespace marker has the
	    value of <tt>{marker}</tt>. This is called importing a bootstrapping namespace, see
	    <xref target="bootstrap"/>.
	  </t>
	  <t>
	    It is not an evaluation error if the namespace is unknown to the processing
	    application. A processing application MAY produce warnings when it encounters this case.
	  </t>
	</section>

	<section anchor="import-pkg" title="Importing a package">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( import {base}:Nat ( package {id}:Expr ) {count}:Nat )</tt></t>
	    </list>
	  </t>
	  <t>
	    The semantics of this form is the side-effect of associating the first <tt>{count}</tt>
	    namespaces in the package identified by <tt>{id}</tt> with a continuous range of marker
	    bytes starting at <tt>{base}</tt>, within the scope of this expression.
	  </t>
	  <t>
	    It is not an evaluation error if the package is unknown to the processing application,
	    or if it is known but not all namespaces packaged inside are. A processing application
	    MAY produce warnings when it encounters those cases.
	  </t>
	  <t>
	    This forms needs an explicit number of namespaces to import to preserve transparency: if
	    the number was implicit, there could be issues when a processing application that
	    doesn't known how many namespaces are in the package wanted to modify the BULK
	    stream. This makes BULK overall simpler and more predictable.
	  </t>
	  <t>
	    Example: <tt>( import 21 ( package {foo} ) 3 )</tt> associates the first 3 namespaces of
	    the package identified by <tt>{foo}</tt> to the markers 21, 22 and 23.
	  </t>

	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( import {base}:Nat ( package {id}:Expr ) {count}:Nat
	      {increment}:Nat )</tt></t>
	    </list>
	  </t>
	  <t>
	    The semantics of this form is the side-effect of associating the first <tt>{count}</tt>
	    namespaces in the package identified by <tt>{id}</tt> with a range of marker bytes
	    starting at <tt>{base}</tt> at <tt>{increment}</tt> increments, within the scope of this
	    expression.
	  </t>
	  <t>
	    Example: <tt>( import 21 ( package {foo} ) 3 2 )</tt> associates the first 3 namespaces
	    of the package identified by <tt>{foo}</tt> to the markers 21, 23 and 25.
	  </t>
	</section>
	  
	<section anchor="canonid" title="Canonical identifiers">
	  <t>
	    A processing application MAY use any BULK expression as a namespace or package
	    identifier, including atoms like numbers or arrays, but it is RECOMMENDED to use one
	    kind of expression called a <em>canonical identifier</em>. Canonical identifiers have
	    the shape <tt>( {type}:Ref {content}:Bytes )</tt>. The role of <tt>{type}</tt> is to
	    describe how to interpret <tt>{content}</tt>: is it a URI, a UUID, a SHA-3 checksum?
	  </t>
	  <t>
	    Canonical identifiers can be used to efficiently encode IPLD's Content IDentifiers in
	    BULK, but they have a broader use. Where IPLD's CIDs can only encode content-adressing,
	    canonical identifiers can encode arbitrary identifiers, including some that can be
	    created independently of the content, like UUIDs.
	  </t>
	</section>

	<section anchor="def-ns" title="Namespace definition">
	  <t>
	    <list style="hanging">
	      <t hangText="shape">
		<tt>( define ( namespace {id}:Expr ) {def}:Bytes )</tt>
	      </t>
	    </list>
	  </t>
	  <t>
	    The semantics of this form is the side-effect of defining a new namespace. The bytes
	    contained in <tt>{def}</tt> MUST be a BULK stream with a version form, and have the
	    shape <tt>( version {major}:Nat {minor}:Nat ) {config}:Expr
	    {definitions}</tt>. <tt>{config}</tt> MUST be a form and its first element MUST have the
	    shape <tt>( namespace {marker}:Nat )</tt>. The state of the namespace associated to
	    <tt>{marker}</tt> after evaluating all expressions in <tt>{definitions}</tt> (starting
	    with a namespace with no definitions, no mnemonics, and no documentations) is made the
	    definition of the namespace identified by <tt>{id}</tt>, within the scope of this
	    expression. This also associates that namespace to the namespace marker
	    <tt>{marker}</tt>, within the scope of this expression.
	  </t>
	  
	  <section anchor="verifiable" title="Verifiable namespace definition">
	    <t>
	      When a processing application recognizes that <tt>{id}</tt> designates a digest that
	      matches the bytes contained in <tt>{def}</tt>, this creates a verifiable namespace.
	    </t>
	    <t>
	      If more data than <tt>{id}</tt> is needed to verify <tt>{id}</tt> against the bytes
	      contained in <tt>{def}</tt> (like the salt of a hash function, or the namespace of a
	      UUID), this data MUST be provided in <tt>{config}</tt>.
	    </t>
	    <t>
	      It is an evaluation error if the processing application recognizes that <tt>{id}</tt>
	      designates a digest but it doesn't match the bytes contained in <tt>{def}</tt>.
	    </t>
	    <t>
	      Verifiable namespaces are meant to be immutable, but that would be circumvented if
	      they were built upon namespaces that aren't. A verifiable namespace that only uses
	      names from immutable namespaces is an immutable namespace (see <xref
	      target="kinds"/>).
	    </t>
	    <t>
	      A processing application SHOULD only consider digest algorithms that are currently
	      known to be cryptographically secure for the determination of namespace and package
	      immutability. For example, a processing application could check that a namespace with
	      an MD5 identifier is verifiable, but it SHOULD NOT determine it to be immutable.
	    </t>
	    <t>
	      The BULK stream in <tt>{def}</tt> has a version form to prevent the possibility that a
	      verified namespace definition could be evaluated to different results by applications
	      using different BULK versions. The definitions are in a nested BULK stream because if
	      the definition could use immutable namespaces imported outside of the definition, the
	      same verifiable definition could be used in different contexts, also defeating the
	      immutability.
	    </t>

	    <section anchor="bootstrap" title="Bootstrapping verification">
	      <t>
		When using canonical identifiers (see <xref target="canonid"/>), a verifiable
		namespace will use a form to express its digest. For immutable namespaces to exist,
		this means that at least one namespace needs to express its own digest with a name
		from within itself. This is called a bootstrapping namespace. A bootstrapping
		namespace MUST use a canonical identifier.
	      </t>
	      <t>
		When importing a bootstrapping namespace, the processing application looks up in its
		known namespaces if there is a namespace identified by a form whose operator is a
		name in itself with the same name index as the operator in the identifier in the
		import. If that name is associated with a digest algorithm and the logic of the
		digest algorithm determines that the digest form in the import matches the digest in
		the namespace definition (along with additional configuration data provided there),
		then the bootstrapping is successful (meaning that the definition has been found and
		can be imported).
	      </t>
	      <t>
		This process doesn't rely just on the digest in the import being identical to the
		digest in the known definition, because some digests algorithms
		(e.g. extended-output functions) have been designed to retain some collision
		resistance when using a prefix of the digest, so a definition could contain a 256 or
		512 bits wide digest, but some small BULK streams could import it with the first 64
		or 128 bits, when it is worth the trade-off between stream size and collision
		resistance.
	      </t>
	      <t>
		It is not an evaluation error if bootstrapping fails. The only consequence is that
		the bootstrapping namespace is not known, and any other namespace using this
		namespace for its identifier will not be known either. A processing application MAY
		produce warnings when it encounters this case.
	      </t>
	      <t>
		The same logic can be used to import a bootstrapping package, which is a package
		identified with a form from one of its own namespaces.
	      </t>
	      <t>
		When using immutable namespaces, bootstrapping packages are likely to be norm, as
		using almost any immutable namespace requires importing the bootstrapping namespace
		used to identify the target namespace, then that namespace. Every definition of an
		immutable namespace can be accompanied by the definition of a bootstrapping package
		packaging that namespace, its bootstrapping namespace and any other namespaces that
		are likely to be used with it.
	      </t>
	    </section>
	  </section>
	</section>

	<section anchor="def-pkg" title="Package definition">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( define ( package {id}:Expr ) {def}:Bytes )</tt></t>
	    </list>
	  </t>
	  <t>
	    The semantics of this form is the side-effect of creating a package identified by
	    <tt>{id}</tt>. The bytes contained in <tt>{def}</tt> MUST be a BULK stream with a
	    version form and have the shape <tt>( version {major}:Nat {minor}:Nat ) {config}:Expr
	    {preamble} {namespaces}:Expr</tt>. <tt>{config}</tt> MUST be a
	    form. <tt>{namespaces}</tt> MUST be a form containing a sequence of expressions each
	    identifying a BULK namespace.
	  </t>
	  <t>
	    When a processing application recognizes that <tt>{id}</tt> designates a digest that
	    matches the bytes contained in <tt>{def}</tt>, this creates a verifiable package.
	  </t>
	  <t>
	    If more data than <tt>{id}</tt> is needed to verify <tt>{id}</tt> against the bytes
	    contained in <tt>{def}</tt> (like the salt of a hash function, or the namespace of a
	    UUID), this data MUST be provided in <tt>{config}</tt>.
	  </t>
	  <t>
	    It is an evaluation error if the processing application recognizes that <tt>{id}</tt>
	    designates a digest but it doesn't match the bytes contained in <tt>{def}</tt>.
	  </t>
	  <t>
	    Packages are meant to be immutable, but that would be circumvented if they were built
	    upon namespaces that aren't. A verifiable package that only contains immutable
	    namespaces is an immutable package.
	    </t>
	</section>

	<section anchor="def-name" title="Name definitions">
	  <t>
	    To define a reference is to make a value the semantics of any reference with the same
	    associated namespace and the same name, in the scope of that definition.
	  </t>
	  <t>
	    This change of semantics operates on the namespace as identified by its unique
	    identifier, not the marker value. This means that if a namespace Foo is associated to
	    markers 21 and 22 and a reference with namespace marker 21 and name 0 is defined to
	    <tt>true</tt>, then in the scope of that definition, a reference with namespace marker
	    22 ans name 0 will be evaluated as <tt>true</tt>.
	  </t>
	  <t>
	    When a BULK stream containing definitions for a namespace comes from a trusted source
	    (i.e. in configuration files of the application, or in the communication with an agent
	    that has been granted the relevant authority), an application MAY give those definitions
	    long-lasting semantics (i.e. keep the values of the names at the end of parsing). This
	    is the RECOMMENDED mechanism for bulk namespace definition when the semantics of the
	    defined expressions can be expressed completely by BULK expressions (see <xref
	    target="robustNS"/>).
	  </t>

	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( define {ref}:Ref {value}:Expr )</tt></t>
	    </list>
	  </t>
	  <t>
	    The semantics of this form is the side-effect of defining the reference <tt>{ref}</tt>
	    to the value of evaluating <tt>{value}</tt>.
	  </t>
	</section>

	<section anchor="mnemonic" title="Mnemonic">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( mnemonic ( namespace {marker}:Nat ) {mnemonic}:Expr
	      )</tt></t>
	    </list>
	  </t>
	  <t>
	    This shape declares the value of evaluating <tt>{mnemonic}</tt> to be the mnemonic for
	    the namespace associated with the marker <tt>{marker}</tt>.
	  </t>
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( mnemonic Ref {mnemonic}:Expr )</tt></t>
	    </list>
	  </t>
	  <t>
	    This shape declares the value of evaluating <tt>{mnemonic}</tt> to be the mnemonic for
	    the name designated by the reference.
	  </t>
	</section>

	<section anchor="explain" title="Explain">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( explain ( namespace {marker}:Nat ) {doc}:Expr )</tt></t>
	    </list>
	  </t>
	  <t>
	    This shape declares the value of evaluating <tt>{doc}</tt> to be the documentation for
	    the namespace associated with the marker <tt>{marker}</tt>.
	  </t>
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( explain Ref {doc}:Expr )</tt></t>
	    </list>
	  </t>
	  <t>
	    This shape declares the value of evaluating <tt>{doc}</tt> to be the documentation for
	    the name designated by the reference.
	  </t>
	  <t>
	    Documentation expressions can be strings with plain text or use any BULK vocabulary to
	    use a richer documentation format (including BULK forms to make a format of plain text
	    explicit).
	  </t>
	</section>
      </section>

      <section title="Strings and other typed byte arrays">
	<section anchor="stringenc" title="Current encoding">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( define string {encoding}:Expr )</tt></t>
	    </list>
	  </t>
	  <t>
	    The semantics of this form is the side-effect that, in the scope of this expression, the
	    default encoding for expressions that are understood by the application as character
	    strings is the encoding designated by <tt>{encoding}</tt>.
	  </t>
	  <t>
	    As the abstract yield doesn't contain strings but expressions that will be used as
	    strings by the application, it is not a parsing error if the application doesn't
	    recognize <tt>{encoding}</tt>. In this situation, it is only a parsing error when the
	    application actually needs to decode a byte sequence as a string with that encoding. It
	    is not a parsing error when a processing application only transmits a byte sequence
	    encoding a string, if it can accurately convey the encoding to the receiving
	    application.
	  </t>
	</section>

	<section anchor="string" title="String">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( string {string}:Expr )</tt></t>
	    </list>
	  </t>
	  <t>
	    This form indicates that the bytes contained in the expression <tt>{string}</tt> are
	    meant to be interpreted as a string encoded with the current default string encoding.
	  </t>
	</section>

	<section anchor="string*" title="String with explicit encoding">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( string {encoding}:Expr {string}:Epxr )</tt></t>
	    </list>
	  </t>
	  <t>
	    This form indicates that the bytes contained in the expression <tt>{string}</tt> are
	    meant to be interpreted as a string encoded with the encoding designated by
	    <tt>{encoding}</tt>.
	  </t>
	</section>

	<section anchor="iana-charset" title="IANA registered character set">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( iana-charset {id}:Expr )</tt></t>
	    </list>
	  </t>
	  <t>
	    This designates the string encoding registered among the <xref
	    target="IANA-Charsets">IANA Character Sets</xref> whose MIBenum is <tt>{id}</tt>.
	  </t>
	  <t>
	    Type: <tt>Encoding</tt>.
	  </t>
	</section>

	<section anchor="nested-bulk" title="Nested BULK stream">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( bulk {bulk}:Expr )</tt></t>
	    </list>
	  </t>
	  <t>
	    This form indicates that the bytes contained in the expression <tt>{bulk}</tt> are meant
	    to be interpreted as a BULK stream. If the stream doesn't start with a <tt>version</tt>
	    form, the stream explicitly has the same version as the parent stream.
	  </t>
	  <t>
	    The semantics of this form is the same as the evaluation of the BULK stream in
	    <tt>{bulk}</tt> taken as a form. For example, these two forms have the same evaluation:
            <list style="symbols">
	      <t><tt>( 4 5 )</tt></t>
	      <t><tt>( bulk true #[2] 4 5 )</tt></t>
	    </list>
	  </t>
	  <t>
	    This form can be useful to let the application reading a BULK stream skip parsing a
	    large section. In that case, it MUST be enclosed in a <xref target="values"
	    format="none">values</xref> form to prevent the security issue described in <xref
	    target="skippable"/>.
	  </t>
	</section>

	<section anchor="blob" title="Blob">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( blob {blob}:Expr )</tt></t>
	    </list>
	  </t>
	  <t>
	    This form indicates that the bytes contained in <tt>{blob}</tt> are meant be interpreted
	    as just a raw sequence of bytes, not to be decoded.
	  </t>
	</section>
      </section>

      <section title="Array operations">
	<section anchor="concat" title="Array concatenation">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( concat {arrays} )</tt></t>
	    </list>
	  </t>
	  <t>
	    The value of this form is an array that contains the bytes contained in the first
	    expression in <tt>{arrays}</tt> followed by the bytes contained in the second
	    expression, and so on for each expression. It is an evaluation error if any expression
	    in <tt>{arrays}</tt> doesn't evaluate to a container of bytes.
	  </t>
	  <t>
	    <tt>( concat )</tt> evaluates to an empty array like <tt>#[0]</tt>.
	  </t>
	</section>

	<section title="Indexed data">
	  <t>
	    When writing a stream containing a big number of expressions where an application may
	    want to access one of those expression without parsing all expressions before, one could
	    imagine as a solution to use pointer-like references that each use the offset of some
	    expression in the stream. This solution creates a security risk, because if reading
	    according to the pointers doesn't produce the same result as parsing the stream without
	    using them, an attacker might use this inconsistency to their advantage, when they can
	    expect one application to use pointers and another application to use normal parsing,
	    especially when the stream is big enough that verifying the consistency of the pointers
	    might be costly enough that it might not be done or not in time to prevent the attack.
	  </t>
	  <t>
	    Because of that risk, whenever a stream includes indexed BULK expressions, that is,
	    expressions that are meant to be accessed by their byte position, indexed reading SHOULD
	    be the only way used to access them. To that end, indexed data SHOULD be stored in
	    arrays.
	  </t>
	  <t>
	    When the goal of indexed data is to selectively parse only part of the BULK stream, a
	    <xref target="values" format="none">values</xref> form MUST be used to prevent the
	    security issue described in <xref target="skippable"/>.
	  </t>

	  <section anchor="indexed-bulk" title="Indexed BULK expression">
	    <t>
	      <list style="hanging">
		<t hangText="shape"><tt>( indexed-bulk {container}:Expr {start}:Expr )</tt></t>
	      </list>
	    </t>
	    <t>
	      The semantics of this form is the value of evaluating the BULK expression starting at
	      offset <tt>{start}</tt> in the bytes contained in expression <tt>{container}</tt>.
	    </t>
	    <t>
	      Beware that, although evaluations of all other BULK definitions follow lexical
	      scoping, any definition used inside an indexed expression that isn't defined inside
	      that same indexed expression follows dynamic scoping with respect to any places where
	      it's used. Any definition made inside an indexed expression still follows lexical
	      scoping.
	    </t>
	  </section>
	  <section anchor="indexed-array" title="Indexed array">
	    <t>
	      <list style="hanging">
		<t hangText="shape"><tt>( indexed-array {container}:Expr {start}:Expr {size}:Expr
		)</tt></t>
	      </list>
	    </t>
	    <t>
	      The semantics of this form is the value of an array whose content are <tt>{size}</tt>
	      bytes, starting at offset <tt>{start}</tt> in the bytes contained in the expression
	      <tt>{container}</tt>.
	    </t>
	    <t>
	      <list style="hanging">
		<t hangText="shape"><tt>( indexed-array {container}:Expr {start}:Expr )</tt></t>
	      </list>
	    </t>
	    <t>
	      The semantics of this form is the value of an array whose content are the bytes
	      starting at offset <tt>{start}</tt> in the array <tt>{container}</tt> until its end.
	    </t>
	    <t>
	      Compared to <tt>indexed-bulk</tt>, which can reference an array expression,
	      <tt>indexed-array</tt> is useful when several different but overlapping sections of
	      the same byte sequence are needed as arrays, or when reversing <xref
	      target="packing">packing</xref> through evaluation (to avoid packing the marker
	      bytes).
	    </t>
	  </section>
	</section>
      </section>

      <section anchor="booleans" title="Booleans">
	<t>
	  <list style="hanging">
	    <t hangText="shape"><tt>true</tt></t>
	    <t hangText="shape"><tt>false</tt></t>
	  </list>
	</t>
	<t>
	  Type: <tt>Boolean</tt>.
	</t>
      </section>

      <section title="Substituton">
	<section anchor="subst" title="Substitution function">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( subst {code} )</tt></t>
	    </list>
	  </t>
	  <t>
	    The semantics of this form is an expression of type <tt>LazyFunction</tt>, called the
	    substitution function. The semantics of the substitution function are the semantics of
	    <tt>{code}</tt>, but where the names <tt>arg</tt> and <tt>rest</tt> are substituted as
	    such:
	    <list style="symbols">
	      <t anchor="arg"><tt>( arg {n}:Nat )</tt> is replaced by the element number
	      <tt>{n}</tt> (starting at zero) of the substitution function's arguments list.</t>
	      <t anchor="rest"><tt>( rest {n}:Nat )</tt> is replaced by the substitution function's
	      arguments list without its first <tt>{n}</tt> elements.</t>
	    </list>
	  </t>
	  <t>
	    It is an evaluation error if the substitution function is called with too few arguments
	    with respect to the <tt>arg</tt> and <tt>rest</tt> forms in <tt>{code}</tt>.
	  </t>
	</section>
	<section title="Examples">
	  <t>
	    Here is a definition of the inverse followed by the numbers 1/2, 1/3 and 1/4:
	  </t>
	  <t><figure><artwork>( define inverse ( subst ( fraction 1 ( arg 0 ) ) ) )
( inverse 2 )
( inverse 3 )
( inverse 4 )</artwork></figure></t>
          <t>
	    Substitution will splice multiple expressions in place:
	  </t>
	  <t>
	    The evaluation of:
	  </t>
	  <t><figure><artwork>( define foo ( subst 20 ( rest 0 ) 50 ) )
( 10 ( foo 30 40 ) 60 )</artwork></figure></t>
          <t>
	    must produce: <tt>( 10 20 30 40 50 60 )</tt>
	  </t>
	</section>

	<section anchor="values" title="Isolated scope">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( values {exprs} )</tt></t>
	    </list>
	  </t>
	  <t>
	    This form's semantics is the values produced by the evaluation of <tt>{exprs}</tt>. This
	    means that the side-effects in <tt>{exprs}</tt> can affect its own evaluation, but not
	    the scope of the <tt>values</tt> form. It makes it possible to isolate evaluation
	    side-effects.
	  </t>
	  <t>
	    This MUST be used for selectively parseable data, to prevent the security issue
	    described in <xref target="skippable"/>. Because the semantics of this form is only the
	    values produced by the evaluation of its operands, this evaluation can wait until those
	    values are actually needed.
	  </t>
	  <t>
	    This property means that if a processing application uses lazy evaluation of BULK
	    expressions, every expression in a <tt>values</tt> form is selectively evaluated only
	    when needed, which in turns means that any nested BULK stream within a <tt>values</tt>
	    form is selectively parsed only when needed.
	  </t>
	</section>

      </section>

      <section anchor="arithmetic" title="Arithmetic">
	<t>
	  A processing application must recognize the type of all expressions defined in this
	  specification that have the type Nat, but an application MAY consider a number as having
	  an unknown value if it can't decode its value or has no adequate data type to store it. It
	  is only a parsing error if the number is needed by the parsing algorithm. It is only an
	  evaluation error if the number is needed by the evaluation algorithm.
	</t>
	<t>
	  In the text notation of a BULK stream, a decimal integer is the notation for the smallest
	  byte sequence that yields this integer as described in <xref target="nat-enc"></xref>. For
	  example, <tt>( 31 256 )</tt> is a notation for the bytes <tt>0x01 0x9F 0xC2-0100
	  0x02</tt>.
	</t>

	<section anchor="unsigned-int" title="Unsigned integer">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( unsigned-int {bits}:Expr )</tt></t>
	    </list>
	  </t>
	  <t>
	    The semantics of this form is the value of the unsigned integer represented in binary
	    notation in the bits contained in <tt>{bits}</tt>. This form exists in case
	    disambiguation of the semantics of an array (or another bit container) is necessary.
	  </t>
	  <t>
	    Type: <tt>Number</tt>, <tt>Real</tt>, <tt>Int</tt>, <tt>Nat</tt>.
	  </t>
	</section>

	<section anchor="signed-int" title="Signed integer">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( signed-int {bits}:Expr )</tt></t>
	    </list>
	  </t>
	  <t>
	    The semantics of this form is the value of the signed integer represented in
	    two's-complement notation in the bits contained in <tt>{bits}</tt>.
	  </t>
	  <t>
	    Type: <tt>Number</tt>, <tt>Real</tt>, <tt>Int</tt>.
	  </t>
	</section>

	<section anchor="fraction" title="Fraction">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( fraction {num}:Expr {div}:Expr )</tt></t>
	    </list>
	  </t>
	  <t>
	    The semantics of this form is the fraction with denominator <tt>{num}</tt> and divisor
	    <tt>{div}</tt>.
	  </t>
	  <t>
	    Type: <tt>Number</tt>.
	  </t>
	  <section title="Fixed-point numbers">
	    <t>
	      <tt>fraction</tt> makes it possible to express fixed-point numbers in BULK:
	      <list style="symbols">
		<t>
		  <tt>( fraction 15 4 )</tt> has value <tt>11.11<sub>2</sub></tt>
		  (<tt>3.75<sub>10</sub></tt>)
		</t>
		<t>
		  <tt>( fraction 123 100 )</tt> has value <tt>1.23</tt>
		</t>
	      </list>
	    </t>
	    <t>
	      A more extensive arithmetic vocabulary could define forms to express fixed-point
	      numbers according to a given base, as well as predefined fixed points (e.g. the use of
	      fixed-point numbers with two decimals is pretty common with financial data).
	    </t>
	  </section>
	</section>

	<section anchor="binary-float" title="Binary floating-point number">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( binary-float {bits}:Expr )</tt></t>
	    </list>
	  </t>
	  <t>
	    The semantics of this form is the floating-point number expressed in IEEE 754-2008
	    binary interchange format by the bits contained in <tt>{bits}</tt>. <tt>{bits}</tt> can
	    be of size 16, 32, 64, 128 or any bigger multiple of 32 bits, as per IEEE 754-2008
	    rules.
	  </t>
	  <t>
	    Types: <tt>Number</tt>, <tt>Real</tt>, <tt>Float</tt>.
	  </t>
	</section>

	<section anchor="decimal-float" title="Decimal floating-point number">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( decimal-float {bits}:Expr )</tt></t>
	    </list>
	  </t>
	  <t>
	    The semantics of this form is the floating-point number expressed in IEEE 754-2008
	    decimal interchange format by the bits contained in <tt>{bits}</tt>. <tt>{bits}</tt> can
	    be of size 32, 64, 128 or any bigger multiple of 32 bits, as per IEEE 754-2008 rules.
	  </t>
	  <t>
	    Types: <tt>Number</tt>, <tt>Real</tt>, <tt>Float</tt>.
	  </t>
	</section>
      </section>

      <section title="Bytecodes">
	<t>
	  This specification and other official BULK specifications use forms with a reference
	  operator as their basic building blocks. Basically, these are a binary representation of
	  an abstract syntax tree. As noted previously, this means that most representations weigh 4
	  bytes plus their actual content, which will in turn have some overhead because of one or
	  several marker bytes.
	</t>
	<t>
	  But when there is a special need for compactness, BULK makes it possible to design
	  protocols and formats with different trade-offs, while retaining its property of being
	  parseable by processing applications not knowing the protocol in its entirety.
	</t>
	<t>
	  On one end of the spectrum, a format might choose to use an array to encapsulate an ad hoc
	  binary format. An extreme use of this scheme would be to use BULK just to make explicit
	  the binary format used and for nothing else. With a known <xref
	  target="profiles">profile</xref> (for example with a file extension and/or media type for
	  such explicitly typed BLOBs), such a BULK stream can consist solely of the version form, a
	  reference that describes the binary format and an array, which would amount to an overhead
	  between 11 bytes and 20 bytes depending on the size of the content (11, 13, 14, 16 and 20
	  bytes for contents of no more than 63B, 255B, 65kB, 4GB and 18EB respectively). Without a
	  profile, with the namespaces associations in a package, the minimum overhead is only
	  between 32 and 41 bytes (the difference is a single <tt>import</tt> form, assuming a
	  digest of 64 bits).
	</t>
	<t>
	  Still, even this extreme in the design space retains the ability to insert expressions in
	  the BULK stream, whatever their type. Thus metadata can be added about data that is
	  represented in a format that doesn't allow for metadata or that allows only for limited
	  metadata. <xref target="typing-medias"/> gives a few examples of what encapsulating
	  existing media types in BULK could bring.
	</t>
	<t>
	  In-between these two extremes, several options are available to produce a format that
	  leverages the BULK parser a lot more while being more compact than a basic BULK
	  format. The following forms provide a standard way to create such formats, called BULK
	  bytecodes.
	</t>
	<t>
	  A BULK bytecode is a flat sequence of expressions. The evaluation of a bytecode form
	  transforms that sequence to an abstract syntax tree of its contents (and then the
	  resulting expression can be evaluated with the normal BULK evaluation rules). The
	  expressions of the bytecode are divided among bytecode operators and bytecode
	  operands. Operators are references that will end up as form operators in the abstract
	  syntax tree. Operands are all other expressions. Prefix bytecodes are those where
	  operators come before their operands, postfix bytecodes are those where operators come
	  after their operands. In the following forms, operators MUST be references.
	</t>
	<t>
	  When evaluating a bytecode, it is an evaluation error when the processing application
	  encounters a reference for which it cannot determine if it is an operator or its arity
	  (the number of operands it will have). An expression that is not a reference is always an
	  operand.
	</t>

	<section anchor="prefix" title="Prefix bytecode">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( prefix {bytecode} )</tt></t>
	    </list>
	  </t>
	  <t>
	    This is a prefix bytecode form. The bytecode to be transformed is the sequence of
	    expressions in <tt>{bytecode}</tt>.
	  </t>
	  <t>
	    To transform a prefix bytecode, a processing application creates an alternate
	    context. If the first expression of the bytecode is an operand, it is removed from the
	    beginning of the bytecode and appended at the end of the alternate context. If the first
	    expression of the bytecode is an operator, it is removed from the beginning of the
	    bytecode and a list is created with the operator as the first expression, then as many
	    next expressions as its arity are removed from the beginning of the bytecode and
	    appended at the end of this list. Then that resulting list is appended at the end of the
	    alternate context. The transformation continues until the bytecode is empty, in which
	    case the transformation is complete and the alternate context is the value of evaluating
	    the bytecode form. The resulting form can then be evaluated in turn.
	  </t>
	  <t>Example: the evaluation of</t>
	  <t><figure><artwork>( define ( arity prefix )
  ( nil game ) ( 2 black ) )
( prefix game black 1 2 black 3 4 black 5 6 )
	  </artwork></figure></t>
	  <t>is</t>
	  <t><figure><artwork>( game
 ( black 1 2 )
 ( black 3 4 )
 ( black 5 6 ) )
	  </artwork></figure></t>
          <t>
	    It is an evaluation error when there are less expressions remaining in the bytecode than
	    the arity of the current operator. In the error information, a processing application
	    MAY provide the alternate context and the remaining bytecode. The alternate context
	    after a failed transformation MUST NOT appear in the abstract yield as if evaluation had
	    successfully transformed the bytecode.
	  </t>
	</section>
	<section anchor="postfix" title="Postfix bytecode">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( postfix {bytecode} )</tt></t>
	    </list>
	  </t>
	  <t>
	    This is a postfix bytecode form. The bytecode to be transformed is the sequence of
	    expressions in <tt>{bytecode}</tt>.
	  </t>
	  <t>
	    To transform a postfix bytecode, a processing application creates a data stack. If the
	    first expression of the bytecode is an operand, it is removed from the beginning of the
	    bytecode and pushed on top of the stack. If the first expression of the bytecode is an
	    operator, it is removed from the beginning of the bytecode and a list is created with
	    the operator as the first expression, then as many next expressions as its arity are
	    popped from the stack and appended at the end of this list (with the top of the stack as
	    the last element). Then that resulting list is pushed on top of the stack. The
	    transformation continues until the bytecode is empty, in which case the transformation
	    is complete and the list of expressions on the stack (with the top of the stack as the
	    last element) is the value of evaluating the bytecode form. The resulting form can then
	    be evaluated in turn.
	  </t>
	  <t>
	    Example: the evaluation of
	  </t>
	  <t><figure><artwork>( define ( arity postfix )
 ( nil game ) ( 2 black white comment alternative ) )
( postfix
  game
  1 2 black
  "white tried an unorthodox opening" 3 4 white comment
  "a more classical opening would be" 8 9 white comment
  alternative
  2 3 black
  4 5 white )</artwork></figure></t>
          <t>is</t>
	  <t><figure><artwork>( game
  ( black 1 2 )
  ( alternative
    ( comment "white tried an unorthodox opening" ( white 3 4 ) )
    ( comment "a more classical opening would be" ( white 8 9 ) ) )
  ( black 2 3 )
  ( white 4 5 ) )</artwork></figure>
	  </t>
          <t>
	    The obvious advantage of postfix bytecode is that it makes it possible to compact nested
	    forms when they have a known arity. When a reference in a vocabulary can be used in a
	    form containing a variable number of expressions, if some arity is used frequently
	    enough, an application can define a specific form for it. The trade-offs for this are
	    explained in <xref target="arityForm"/>
	  </t>
          <t>
	    It is an evaluation error when there are less expressions remaining on the data stack
	    than the arity of the current operator. In the error information, a processing
	    application MAY provide the data stack and the remaining bytecode. The data stack after
	    a failed transformation MUST NOT appear in the abstract yield as if evaluation had
	    successfully transformed the bytecode.
	  </t>
	</section>
	<section anchor="arity" title="Arity definition">
	  <t>
	    <list style="hanging">
	      <t hangText="shape"><tt>( define ( arity {contexts} ) {arities} )</tt></t>
	    </list>
	  </t>
	  <t>
	    This form defines the arity of references in the context of bytecodes.
	  </t>
	  <t>
	    <tt>{contexts}</tt> can contain <tt>prefix</tt>, <tt>postfix</tt>, or any other
	    reference, to specify in the context of which kind of bytecode the arities are
	    modified. If <tt>{contexts}</tt> is empty, the arities are modified in the context of
	    all kinds of bytecodes.
	  </t>
	  <t>
	    <tt>{arities}</tt> is a sequence of expressions that each can be shaped as such:
	    <list style="symbols">
	      <t><tt>nil</tt>: meaning all known arities should be forgotten</t>
	      <t>
		<tt>( {kind}:Expr {target} )</tt>:
		<list style="symbols">
		  <t>
		    if <tt>{kind}</tt> is <tt>nil</tt>, it sets all references designated by
		    <tt>{target}</tt> as operands
		  </t>
		  <t>
		    if <tt>{kind}</tt> is typed <tt>Nat</tt>, it sets all references designated by
		    <tt>{target}</tt> as operators of arity <tt>{kind}</tt>
		  </t>
		  <t>
		    if <tt>{target}</tt> is <tt>nil</tt>, it designates all references with unknown
		    arity
		  </t>
		  <t>
		    if <tt>{target}</tt> is a sequence of references, it designates each of those
		  </t>
		</list>
	      </t>
		
	    </list>
	  </t>
	</section>
      </section>
    </section>

    <section title="Optimizing compactness">
      <section anchor="packing" title="Packing">
        <t>
	  If the overhead of several marker bytes in some operands is too much, more compactness can
	  be achieved by packing together small operands. For example, instead of an operator with
	  two integers as its operands, one could specify an operator to take a single array as
	  operand and extract the integers from it. When the processing application does this
	  extraction, the format retains the ability to operate on many sizes of integers, because
	  the processing application can still deduce the size of the integers by dividing the size
	  of the array by two. This can be used outside or inside of bytecodes (to stack the
	  compacting effects of both).
	</t>
	<t>
	  For example, a BULK format representing player moves with a pair of coordinates on a large
	  board might represent a single move with the following shapes:
	</t>
	<t>
	  <list style="hanging">
	    <t hangText="basic (8 bytes)"><tt>( move/2 #[1] 0x41 #[1] 0x5A )</tt></t>
	    <t hangText="packed basic (7 bytes)"><tt>( move/1 #[2] 0x41 0x5A )</tt></t>
	    <t hangText="bytecode (6 bytes)"><tt>move/2 #[1] 0x41 #[1] 0x5A</tt></t>
	    <t hangText="packed bytecode (5 bytes)"><tt>move/1 #[2] 0x41 0x5A</tt></t>
	  </list>
	</t>
	<t>
	  Packing can also be done without adding its burden on the logic of the processing
	  application, by using evaluation to transform packed forms into simpler forms (but they
	  need to be created separately for each size of operands). For example, the following:
	</t>
	<t><figure><artwork>( define move/1-16
  ( subst ( move/2 ( indexed-array ( arg 0 ) 0 1 )
                   ( indexed-array ( arg 0 ) 1 1 ) ) ) )
( move/1-16 #[2] 0x415A )
</artwork></figure></t> 
	<t>Would be evaluated into:</t>
	<t><figure><artwork>( move/2 #[1] 0x41 #[1] 0x5A )</artwork></figure></t>
	<t>
	  More complex packing can be encoded as well. For example, this would be a form that packs
	  two 16-bits unsigned integers, one 32-bits signed integer, one reference, and one
	  variable-length string:
	</t>
	<t><figure><artwork>( define pack-224-ref-str
  ( subst ( bar ( indexed-array ( arg 0 ) 0 2 )
                    ( indexed-array ( arg 0 ) 2 2 )
                    ( signed-int ( indexed-array ( arg 0 ) 4 4 ) )
                    ( indexed-bulk ( arg 0 ) 8 )
                    ( indexed-array ( arg 0 ) 10 ) ) )</artwork></figure></t>
	<t>
	  In essence, packing makes it possible to embed very simple ad hoc binary formats described
	  within the BULK framework.
	</t>
      </section>
      <section title="Mixing literals">
	<t>
	  The transformation defined for the bytecode forms makes it possible to mix literal
	  expressions and operations represented by a sequence of operators and operands. A typical
	  example would be that instead of explicitly encoding which player is playing in turn, to
	  only encode the move and let the player information be implicit, when the order of players
	  is dictated by the game logic. In the previous go example, for instance, one might
	  represent each alternating move by the two players as two integers, lowering the weight of
	  each normal move to 2 bytes as coordinates are below 64:
	</t>
	<t><figure><artwork>( define ( arity postfix )
 ( nil game ) ( 2 black white comment alternative ) )
( postfix
  game
  1 2
  "white tried an unorthodox opening" 3 4 white comment
  "a more classical opening would be" 8 9 white comment
  alternative
  2 3
  4 5 )</artwork></figure></t> 
        <t>
	  The difference between all these schemes and an array containing fixed-size elements is
	  that you keep the ability to insert other forms, like here to represent comments on the
	  game or variants.
	</t>
      </section>
      <section title="Trade-offs">
	<t>
	  The most visible cost of the bytecode format is that if it contains operators whose arity
	  is unknown to a processing application, the whole list after the first occurrence of them
	  is unreadable to that processing application, whereas in the basic format, the processing
	  application can still process all the forms it understands, and that requires no
	  anticipation by the application creating the BULK stream.
	</t>
	<t>
	  The only case where operators could have an unknown arity is when the application writing
	  the stream didn't include the arities of every operator used in the stream to avoid the
	  redundancy with their previous definition (typically in the definition of their respective
	  namespaces). That redundancy would be offset by the space reduction of postfix bytecode
	  for streams containing a few dozens forms. At that point, with all arities explicit, with
	  packing and with literals for the most used forms, postfix bytecode gets on the Pareto
	  front for size and generality, retaining the full generality of BULK while saving a lot of
	  space.
	</t>
	<t>
	  It is RECOMMENDED that whenever explicit arities would be a small fraction of the total
	  stream size, they all be given. A processing application MAY choose to never include
	  explicit arities for the names of the main namespace of the format, if that namespace's
	  definition includes arities, when processing the stream without knowledge of that
	  definition wouldn't make sense.
	</t>
	<t>
	  But another cost of the bytecode format is the loss of resilience. If a BULK stream ends
	  up with errors in the content, whether during creation or transit, but those errors don't
	  affect the syntactic structure of the stream, then those errors will only prevent
	  processing the form they are in. If errors occur within a metadata form, the data is still
	  readable. If errors occur within one entry in an archive, other entries are still
	  readable. But a far wider class or errors will make a bytecode impossible to evaluate, and
	  the blast radius is everything after the error.
	</t>
	<t>
	  When a protocol needs the exchange of messages where every byte spared counts and there is
	  no sense in trying to recover a partial message after corruption, then using a packed
	  bytecode is probably an excellent solution. When writing large quantities of small data
	  elements to long-term storage, when overhead adds up significantly, it is RECOMMENDED to
	  use a mechanism to limit the blast radius of possible data corruption.
	</t>
	<t>
	  There are several possible solutions to limit where bytecode errors propagate: one is
	  chunking the bytecode into several bytecode forms, another is inserting at regular
	  intervals a beacon expression that can never be an operand, and that can "reset" the
	  bytecode transformation process that had been corrupted before. The former is simpler but
	  the latter might be better suited when data is streamed (see <xref target="streaming"/>).
	</t>
      </section>
    </section>


    <section anchor="profiles" title="Profiles">
      <t>
	A profile is a byte sequence parsed by a processing application just after the
	<tt>version</tt> form or before the first expression if there is no <tt>version</tt>
	form. Thus a parser SHOULD look ahead at the beginning of a stream to see if the first three
	bytes are <tt>( bulk:version</tt>. With respect to the BULK stream, the profile is an
	out-of-band information, usually implicit.
      </t>
      <t>
	A processing application doesn't need to actually parse the profile or include the profile's
	yield in the concrete yield, as long as the semantics of the abstract yield are maintained.
      </t>
      <t>
	The same BULK stream might be processed with different profiles.
      </t>
      <t>
	A processing application MUST NOT deduce the profile from the content of a BULK stream.
      </t>

      <section title="Profile redundancy">
	<t>
	  A processing application SHOULD only rely on the use of a profile when it is a safe
	  assumption that the profile is known, for example within a communication where the
	  protocol dictates the profile.
	</t>
	<t>
	  In particular, long-term storage of a BULK stream SHOULD preserve profile information, for
	  example with a media type that dictates the profile.
	</t>
	<t>
	  Otherwise, an application writing a BULK stream in a long-term storage SHOULD include the
	  profile after the version form. For this reason, the expressions in a profile SHOULD have
	  idempotent semantics.
	</t>
      </section>

      <section title="Standard profile">
	<t>
	  This specification defines the default profile that a processing application MUST use when
	  it is not using a specific profile:
	</t>
	<t>
	  <tt>( define string ( iana-charset 106 ) )</tt>
	</t>
	<t>
	  This means that the default string encoding in a BULK stream is UTF-8.
	</t>
      </section>

      <section title="Fixed BULK: all profile, no evaluation">
	<t>
	  Fixed BULK is a mode of processing for a format or protocol where evaluation has been
	  deemed detrimental (it could be that it's too expensive computationally, or that it adds
	  too much complexity in the processing application's code or the protocol). In Fixed BULK
	  mode, a processing application uses a profile that contains one or several namespace
	  associations and possibly definitions. Because evaluation is disabled, no namespace
	  association or definition will be executed during processing and namespaces are "fixed" to
	  their markers as per the profile.
	</t>
	<t>
	  Fixed BULK mode lets a format or protocol use BULK's syntax while operating more like
	  binary format frameworks, like ASN.1, Protocol Buffers, or CBOR. An interesting difference
	  is that the concatenation of the profile and the Fixed BULK stream is a normal BULK
	  stream.
	</t>
      </section>
    </section>

    <section anchor="discover" title="Discovery of namespaces and packages">
      <t>
	When a processing application encounters an unknown namespace or package identifier, it MAY
	ask several sources to provide the definition. This is called <em>discovery</em>. The
	possible sources include the agent that made use of the unknown identifier, known BULK
	registries (registries of definitions for BULK namespaces or packages), or protocols based
	on content-addressing. When discovery is done without user intervention while processing a
	BULK stream, in order to evaluate it fully, it is called <em>immediate discovery</em>. When
	missing identifiers are collected to be retrieved later, under human supervision, it is
	called <em>deferred discovery</em>.
      </t>
      <t>
	The fundamental security risk in this mechanism is when an attacker manages to be the first
	to use some identifier and poison the processing application by giving it a malicious
	definition, or managed to poison one or several BULK registries. To avoid that risk, a
	processing application MUST only accept definitions for immutable namespaces and immutable
	packages during immediate discovery. It is also RECOMMENDED that a BULK registry only store
	definitions for immutable namespaces and immutable packages.
      </t>
      <t>
	During deferred discovery, a processing application MUST NOT accept namespace or package
	definitions from an untrusted source when they are not immutable (including when the digest
	doesn't match the data, or when the identifier form is not known to be a digest by the
	processing application). A BULK registry MUST NOT store namespace or package definitions
	from an untrusted source when they are not immutable.
      </t>
      <t>
	While discovery of bootstrapping namespaces and packages can be done, the digest algorithm
	used in the identifier of a bootstrapping namespace or package MUST have been known before
	discovery (this can only be checked after discovery has retrieved and evaluated the
	definitions). This is possible when a bootstrapping package uses a digest from a namespace
	that was known by the processing application before discovery, or when the digest name comes
	from a namespace where it is aliased to a digest name that was known by the processing
	application before discovery.
      </t>
      <t>
	With immutable namespaces and immutable packages, though, automatic discovery can be a safe
	mechanism if some risks are mitigated:
	<list style="symbol">
	  <t>
	    Asking untrusted registries might expose when other agents communicate with the
	    processing application. Making the request asynchronously, with noise in the timing, can
	    mitigate that risk. Application operators need to consider the trade-off between latency
	    and privacy, or give the option to agents to make that determination. Making the request
	    in ways that hide the processing application's identity, like using The Onion Router, is
	    another option.
	  </t>
	  <t>
	    Asking untrusted registries might expose what kind of data is sent by other agents to
	    the processing application. Making the request in ways that hide the processing
	    application's identity, like using The Onion Router, is an option.
	  </t>
	  <t>
	    Asking the agent or not might expose the fact that a namespace was already known or
	    not. Some processing applications SHOULD provide configurable policies so that operators
	    can choose one that is relevant to the sensitivity of the services they
	    operate. Examples could include: always asking for namespaces that haven't been
	    explicitly marked safe, always or never asking to some classes of agents.
	  </t>
	</list>
      </t>
    </section>

    <section anchor="streaming" title="Streaming">
      <t>
	The parsing and evaluation algorithms allow for reading a BULK stream while it has not been
	completely received by the processing application. Once the parser reaches a dispatch point
	(see <xref target="eval"/>), the fully parsed expression can be evaluated if evaluation is
	enabled.
      </t>
      <t>
	This makes BULK usable for streaming data, including in the case of a communication protocol
	with a long-lived connection where requests and responses must be processed immediately,
	including a full-duplex communication (see <xref target="full-duplex"/>).
      </t>
      <section title="Broadcasting BULK">
	<t>
	  It is possible to stream BULK data in a broadcast setting, meaning that a processing
	  application could start receiving data mid-stream and never see the data sent before it
	  started listening to the broadcasted stream.
	</t>
	<t>
	  A broadcasted BULK stream MUST NOT contain expressions whose evaluation have side-effects
	  when their scope is the abstract yield. If it contained such expressions, the same
	  processing application that started listening at different points in the stream could
	  produce different results for a common portion of the concrete yield.
	</t>
	<t>
	  This means that a protocol that employs BULK broadcasting and uses any extension namespace
	  MUST provide a profile, either in the protocol specification, or during connection
	  establishment. For the latter, when a BULK stream is broadcasted by HTTP, the server can
	  use the <tt>bulk-profile</tt> link relation in headers (see <xref target="link"/>).
	</t>
	<t>
	  When broadcasting BULK, one issue is stream capture, to discover an offset in the stream
	  that is a dispatch point. This specification describes four ways: server clipping, beacon
	  expressions, Ogg encapsulation and Magrat encapsulation, but others are possible.
	</t>
	<section title="Server clipping">
	  <t>
	    Conceptually, the simplest solution to stream capture is just for the server to always
	    start sending data from a dispatch point. If the server streaming data can be aware of
	    the internal structure of the BULK stream, it can make new clients wait until the next
	    dispatch point to start sending them data.
	  </t>
	</section>
	<section title="Beacon expressions">
	  <t>
	    To signal some of the dispatch points, the abstract yield of the BULK stream contains an
	    expression whose byte pattern is unique in the stream, repeated frequently enough to
	    minimize the length of bytes that a processing application must go through before
	    achieving stream capture. After reading that byte pattern, the processing application
	    can start parsing BULK expressions.
	  </t>
	  <t>
	    The problem with the idea of a beacon expression is that because BULK arrays can contain
	    arbitrary bytes, no beacon expression exists that cannot appear in a BULK stream. There
	    are two solutions. First, in a variety of situations, an application could know or
	    preclude a byte pattern to appear in the stream, and chose that expression as
	    beacon. Second, an application could choose a byte pattern that can only appear in a
	    BULK array and, whenever a BULK array contains that byte pattern, represent it in the
	    broadcasted BULK stream as the concatenation of its split across the beacon pattern.
	  </t>
	  <t>
	    For example, if the beacon expression is <tt>0xC3AABBCC</tt>, the expression <tt>#[8]
	    0x0000-C3AABBCC-0000</tt> would become <tt>( concat #[4] 0x0000-C3AA #[4] 0xBBCC-0000
	    )</tt>.
	  </t>
	</section>
	<section anchor="ogg" title="Ogg encapsulation">
	  <t>
	    The Ogg format<xref target="RFC3533"/> already provides an efficient mechanism for
	    stream capture with a relatively low overhead. It can stream and multiplex data from
	    multiple media types.
	  </t>
	  <t>
	    The BULK stream MUST begin with a version form, and the beginning of that form
	    constitutes the codec identifier: <tt>( version 1</tt> (see <xref target="iana-ogg"/>).
	  </t>
	  <t>
	    Ogg packets provided to the Ogg encoder MUST start and end at dispatch points. This
	    ensures than any complete Ogg packet provided to the processing application by the Ogg
	    decoder is a valid BULK stream and can be parsed into zero or more expressions. Granule
	    position SHOULD be the number of expressions parsed in the abstract yield after parsing
	    the Ogg packet.
	  </t>
	</section>
	<section title="Magrat encapsulation">
	  <t>
	    The Magrat encapsulation is inspired by the Ogg format and has similar properties,
	    except that it is less concerned with audio and video, is less powerful, and takes less
	    space (hence the name).
	  </t>
	  <t>
	    A Magrat broadcast stream is a sequence of Magrat forms. The BULK stream that is
	    encapsulated inside a Magrat broadcast stream is called the Magrat embedded stream. This
	    specification documents basic Magrat encapsulation. In this version, a Magrat form has
	    the following shape:
	  </t>
	  <sourcecode>( values ( bulk {chunk}:Bytes ) {checksum}:Expr )</sourcecode>
	  <t>
	    This means that a processing application can look for the sequence of six bytes
	    <tt>0x01-1013-01-1009</tt> to mark the beginning of a Magrat form. In case those bytes
	    were in a BULK array, there are several features that work to verify that they actually
	    started a Magrat form: first, they must be followed by an array <tt>{chunk}</tt>, a form
	    end, a single expression <tt>{checksum}</tt> and another form end, second,
	    <tt>{checksum}</tt> MUST be a digest that matches the bytes contained in
	    <tt>{chunk}</tt>. Those bytes are called the Magrat chunk and they MUST be a valid BULK
	    stream. A protocol using Magrat encapsulation MAY specify which kind of checksum can be
	    used.
	  </t>
	  <t>
	    In the basic Magrat encapsulation, the concatenation of chunks is the embedded
	    stream. An extension of basic Magrat encapsulation MAY add metadata forms inside the
	    chunk that are removed before the chunks are concatenated as the embedded stream.
	  </t>
	</section>
      </section>
      <section anchor="full-duplex" title="Full-duplex communication">
	<section title="Separate channels">
	  <t>
	    At the BULK level, the simplest way to do full-duplex communication is when the
	    underlying protocol can create separate channels for each direction. Each agent streams
	    its own BULK stream that is fully independent from the other agent's stream. There is no
	    restriction on the semantics used in the streams and an agent can use namespace
	    associations and definitions to build a complex data model.
	  </t>
	  <t>
	    Each agent is faced with the usual security considerations of BULK processing (see <xref
	    target="sec"/>). An agent faced with BULK data that crosses a safety threshold SHOULD
	    stop processing the other agent's stream. For that reason, a protocol using separate
	    channels for BULK full-duplex communication SHOULD provide a way for an agent to end a
	    channel. This SHOULD include the reason for ending the channel, and if the agent offers
	    the option to restart the channel from scratch. It MAY also include the option to
	    restart the channel up to a previous safe dispatch point.
	  </t>
	</section>
	<section title="Shared channel">
	  <t>
	    When two agents want to communicate over a bidirectional channel and reference data sent
	    by each other, one naive way to do it would be to consider a virtual BULK stream acting
	    as a kind of shared whiteboard. Each agent sending a BULK expression would add that
	    expression to the white board. When the agents have a mechanism to ensure some
	    transactional safety, meaning that one agent cannot write without having properly read
	    what the other agent had written, this is a valid option. This can even work for more
	    than two agents.
	  </t>
	  <t>
	    For when this safety is not available, this specification defines BULK's basic
	    full-duplex protocol. In the basic protocol, the fact that it is used is an out-of-band
	    information, which could be part of the underlying protocol statically or conveyed
	    during connection establishment. Another BULK full-duplex protocol could instead define
	    a namespace to convey its use and some configuration parameters.
	  </t>
	  <t>
	    In the basic protocol, the agent that initiates the communication is called the
	    initiating agent and the other agent is called the responding agent. The initiating
	    agent and the responding agents each have a set of namespace markers designated for
	    their use. The initiating agent's designated markers are even numbers. The responding
	    agent's designated markers are odd numbers. Agents MUST only associate immutable
	    namespaces and MUST only associate them to their designated markers and MUST NOT
	    associate a namespace to a marker that already is associated to a namespace. Agents MUST
	    only define names that don't already have a definition, and only to references whose
	    namespace marker is in their designated markers. It is a protocol error when any of
	    those rules is broken by either agent.
	  </t>
	  <t>
	    There is a separate virtual BULK stream associated with each agent. Each agent has a
	    <em>pull position</em>, which is an offset in the other agent's virtual stream. At the
	    beginning of the exchange, each agent's pull position is 0, the beginning of the stream.
	  </t>
	  <t>
	    Let Alice and Bob be two agents in a full-duplex communication. When Alice streams a
	    BULK expression A1 that only contains references with standard namespaces or Alice's
	    designated markers, this expression A1 is appended to Alice's virtual stream. But when
	    Alice streams a BULK expression A2 that contains one or several references with Bob's
	    designated markers that got a definition in Bob's virtual stream, through namespace
	    association or by definition, after Alice's pull position, those definitions are
	    <em>pulled</em>, i.e. the content of Bob's virtual stream between Alice's pull position
	    and the first dispatch point where all those references have a definition gets appended
	    to Alice's virtual stream, Alice's pull position becomes that dispatch point, then A2 is
	    appended to Alice's virtual stream. When Alice streams a BULK expression A3 that
	    contains one or several references with Bob's designated markers, but all those
	    references got a definition before the pull position, only A3 is appended to Alice's
	    virtual stream.
	  </t>
	  <t>
	    If both agents have associated a single namespace to correct designated markers for each
	    one, and both agents have independently defined a name that wasn't defined in the
	    namespaces's immutable definition, it is a protocol error when either agent pulls the
	    other agent's definition of that name.
	  </t>
	  <t>
	    Whenever an agent sees a protocol error, it MUST end the connection. Before ending the
	    connection, the agent MAY stream an expression shaped <tt>( explain false
	    {reason}:String )</tt>, in which case <tt>{reason}</tt> MUST contain a human-readable
	    description of the issue that triggered the disconnect.
	  </t>
	</section>
      </section>
    </section>

    <section title="Security Considerations" anchor="sec">
      <section title="Parsing">
	<t>
	  Parsing a BULK stream is designed to be free of side-effects for the processing
	  application, apart from storing the parsed results.
	</t>
	<t>
	  Arrays in BULK carry their size, to avoid the need for escaping their content. A malicious
	  software, however, may announce an array with a size chosen to get an application to
	  exhaust its available memory. When a BULK stream has been completely received, an array
	  bigger than the remaining data is a parsing error. When a BULK stream's size is not known
	  in advance, the application SHOULD use a growable data structure.
	</t>
	<t>
	  Evaluation opens up some known attacks that appear whenever a format provides a way to
	  express abstraction, like the billion laughs attack. As it is explained in <xref
	  target="eval" format="title"/>, an implementation MAY stop evaluation after a predefined
	  number of evaluation steps. As this has been demonstrated not to be sufficient to prevent
	  attacks based on expansion, an implementation SHOULD also put predefined limits on the
	  space that the concrete yield can take on disk or in memory.
	</t>
	<t>
	  A processing application SHOULD use lazy immutable data structures to represent array
	  concatenation and array indexing, as a defence against evaluaton attacks. For example, in
	  the billion laughs attack, the resulting concatenation would produce 9 lists of 10
	  pointers and one actual array of 3 characters, instead of an array of 3 billion
	  characters.
	</t>
	<t>
	  Applications MAY use out-of-band information to select size limits (like HTTP
	  attributes), or a BULK namespace MAY provide hints.
	</t>
      </section>
      <section title="Forwarding">
	<t>
	  When a processing application forwards all or part of the data in a BULK stream to
	  another application, care must be taken if part of the forwarded data was not entirely
	  recognized, as it could be used by an attacker to benefit from the authority the
	  forwarding application has on the recipient of the data.
	</t>
	<t>
	  If a protocol deems it necessary for applications to be able to forward data they don't
	  fully understand, a known protection from that threat is the use of capability security,
	  where the agent that provides the data to be forwarded must also provide the explicit
	  authority that will be used after forwarding. If the authority of the forwarding
	  application is not used, it cannot be abused.
	</t>
      </section>
      <section title="Definitions">
	<t>
	  The architecture of a processing application SHOULD ensure that a malicious agent cannot
	  abuse authority given to it to define a namespace in order to modify associations in
	  other namespaces. Depending on the use of data structures storing BULK expressions, this
	  could amount to giving an attacker a way to manipulate the application's state. See <xref
	  target="robustNS"/> for an example of architecture that is resistant to that kind of
	  attack.
	</t>
      </section>
      <section anchor="skippable" title="Selectively parseable content">
	<t>
	  It could be a security risk if a single BULK stream could be parsed into two different
	  abstract yields by two conformant applications, so the evaluation of the whole stream
	  cannot change whether some part that is designed to be selectively parseable is decoded or
	  not. For that reason, any side-effects in the selectively parseable expressions that
	  affect how BULK expressions are evaluated (like namespace associations or definitions)
	  MUST be isolated.
	</t>
	<t>
	  For that security reason, there isn't a <tt>( bulk-with-size Nat Expr )</tt> form to make
	  the expression skippable, because it would open up that risk when the size given is not
	  the actual size of the enclosed expression, accidentally or maliciously.
	</t>
	<t>
	  Whenever BULK data is selectively parseable, it MUST be enclosed in a <xref
	  target="values" format="none">values</xref> form.
	</t>
      </section>
      <section title="BULK formats and protocols">
	<t>
	  This specification doesn't address in too much detail the security considerations that a
	  BULK format or protocol would need to include, because of the wide diversity of use cases
	  for BULK. But <xref target="design"/> gives some hints and resources on the subject.
	</t>
      </section>
    </section>

    <section title="IANA Considerations">
      <section title="Media type">
	<t>
	  This specification defines two new media types, <tt>application/bulk</tt> and
	  <tt>text/bulk</tt>. Here are the informations for its registration to IANA <xref
	  target="BCP13"/>:
	</t>
	<section title="application/bulk">
	  <t>
	    <list style="hanging">
	      <t hangText="Type name">application</t>
	      <t hangText="Subtype name">bulk</t>
	      <t hangText="Required parameters">N/A</t>
	      <t hangText="Optional parameters">N/A</t>
	      <t hangText="Encoding considerations">none, content is self-describing</t>
	      <t hangText="Security considerations">cf. <xref target="sec"/></t>
	      <t hangText="Interoperability considerations">N/A</t>
	      <t hangText="Published specification">this document</t>
	      <t hangText="Applications that use this media type">the BARK manifest prototype</t>
	      <t hangText="Fragment identifier considerations">this specification defines no
	      semantics for addressing the data with a fragment identifier; a future specification
	      MAY define fragment identifier syntaxes to address the content by byte offset or the
	      parsed results by their position in the abstract yield</t>
	      <t hangText="Additional information"><br/>
	      <list style="hanging">
		<t hangText="Magic numbers">
		  the constraint to start any BULK file with a version form has the side-effect that
		  classes of BULK streams can be identified by a sequence of bytes acting as "magic
		  number", at offset 0:
		  <list style="hanging">
		    <t hangText="0x011000">any BULK stream</t>
		    <t hangText="0x01100081">a BULK stream of major version 1</t>
		    <t hangText="0x011000818002">a BULK stream of version 1.0</t>
		  </list>
		</t>
		<t hangText="File extensions">.bulk</t>
		<t hangText="Structured type name suffix"> <xref target="RFC6839"/><br/>this
		specification defines a suffix <tt>+bulk</tt> for naming media types that use BULK
		as their core syntax</t>
	      </list>
	      </t>
	    </list>
	  </t>
	</section>
	<section title="text/bulk">
	  <t>
	    <list style="hanging">
	      <t hangText="Type name">text</t>
	      <t hangText="Subtype name">bulk</t>
	      <t hangText="Required parameters">N/A</t>
	      <t hangText="Optional parameters">N/A</t>
	      <t hangText="Encoding considerations">content MUST be encoded in UTF-8 <xref target="STD63"/></t>
	      <t hangText="Security considerations">cf. <xref target="sec"/></t>
	      <t hangText="Interoperability considerations">N/A</t>
	      <t hangText="Published specification">this document, <xref target="bulktext"/></t>
	      <t hangText="Applications that use this media type">the BARK manifest prototype</t>
	      <t hangText="Fragment identifier considerations">this specification defines no
	      semantics for addressing the data with a fragment identifier; a future specification
	      MAY define fragment identifier syntaxes to address the content by byte offset or the
	      parsed results by their position in the abstract yield</t>
	      <t hangText="Additional information"><br/>
	      <list style="hanging">
		<t hangText="Magic numbers">
		  The text notation allows for arbitrary number and kinds of whitespaces around
		  lexical elements, so there are no "magic numbers" as such but, in most cases, BULK
		  streams in text notation will not have leading whitespace and use a single space
		  within the first version form, so the first characters can identify classes of
		  BULK streams:
		  <list style="symbol">
		    <t>'<spanx style='verb'>( version </spanx>' or '<spanx style='verb'>( bulk:version </spanx>': any BULK stream</t>
		    <t>'<spanx style='verb'>( version 1 </spanx>' or '<spanx style='verb'>( bulk:version 1 </spanx>': a BULK stream of major version 1</t>
		    <t>'<spanx style='verb'>( version 1 0 )</spanx>' or '<spanx style='verb'>( bulk:version 1 0 )</spanx>': a BULK stream of version 1.0</t>
		  </list>
		</t>
		<t hangText="File extensions">.bulktext</t>
	      </list>
	      </t>
	    </list>
	  </t>
	</section>
      </section>
      <section anchor="link" title="Link relation">
	<t>
	  This specification defines a new link relation type, <tt>bulk-profile</tt>. Here are the
	  informations for its registration to IANA <xref target="RFC8288"/>:
	</t>
	<t>
	  <list style="hanging">
	    <t hangText="Relation name">bulk-profile</t>
	    <t hangText="Description">This link target is the BULK stream that a processing
	    application SHOULD use as the profile of the link's context.</t>
	    <t hangText="Reference">this document</t>
	  </list>
	</t>
      </section>
      <section anchor="iana-ogg" title="Ogg media mapping">
	<t>
	  This specification defines a new Ogg logical bitstream type<xref target="RFC5334"/>, for
	  the media type <tt>application/bulk</tt>:
	</t>
	<t>
	  <list style="hanging">
	    <t hangText="Codec identifier"><tt>char[4]: '\x01\x10\x00\x81'</tt></t>
	    <t hangText="Codecs parameter">bulk1</t>
	  </list>
	</t>
	<t>
	  For more details, see <xref target="ogg"/>.
	</t>
	<t>
	  A BULK-aware Ogg decoder could anticipate future BULK versions and recognize any version
	  form conformant with <xref target="version"/> as codec identifier with the encompassing
	  codecs parameter "bulk".
	</t>
      </section>
    </section>

    <section title="Acknowledgements">
      <t>
	The original author of this specification read <eref
	target="http://www.schnada.de/grapt/eriknaggum-xmlrant.html">Erik Naggum's famous rant about
	XML</eref> several years before, and while forgotten as such for a time, it definitively was
	the seed that slowly bloomed into the design of BULK. This format is dedicated to Erik.
      </t>
      <t>
	Unknowingly, work on BULK started just as CBOR<xref target="RFC8949"/> was in <em>Request
	for Last Call</em> at IETF. The early design goals of BULK and the design goals of CBOR had
	both significant differences and a large common ground. It felt like an implicitly obvious
	choice to make the parser a simple state machine that could be implemented with a jump table
	but CBOR's inspiration was to make fast processing speed and low processing footprint
	explicit requirements.
      </t>
      <t>
	The idea to store together marking bits and a small argument in a marker byte was a direct
	inspiration from both CBOR and MessagePack<xref target="MsgPack"/> and it made BULK's syntax
	and implementation both drastically simpler.
      </t>
    </section>
  </middle>

  <back>
    <references title="Normative References">
      <xi:include href="https://bib.ietf.org/public/rfc/bibxml-rfcsubseries/reference.BCP.13.xml"/>
      <xi:include href="https://bib.ietf.org/public/rfc/bibxml-rfcsubseries/reference.BCP.14.xml"/>
      <xi:include href="https://bib.ietf.org/public/rfc/bibxml-rfcsubseries/reference.BCP.18.xml"/>
      <xi:include href="https://bib.ietf.org/public/rfc/bibxml-rfcsubseries/reference.STD.63.xml"/>
      <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.6839.xml"/>
      <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.8288.xml"/>
      <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.3533.xml"/>
      <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.5334.xml"/>

      <reference anchor="IANA-Charsets" target="http://www.iana.org/assignments/character-sets">
        <front>
          <title>
	    IANA Charset Registry (archived at):
          </title>
	  <author/>
	  <date/>
        </front>
      </reference>

    </references>


    <references title="Informative references">

      <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.5234.xml"/>
      <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.7540.xml"/>
      <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.8610.xml"/>
      <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.8949.xml"/>
      <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.9839.xml"/>
      <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.8264.xml"/>
      <xi:include href="https://bib.ietf.org/public/rfc/bibxml3/reference.I-D.bormann-cbor-draft-numbers.xml"/>

      <reference anchor="Avro" target="http://avro.apache.org/docs/1.7.4/spec.html">
	<front>
	  <title>Apache Avro™ 1.7.4 Specification</title>
	  <author initials="D." surname ="Cutting" fullname="Doug Cutting">
            <organization>Cloudera</organization>
	  </author>
	  <date month="February" year="2013"/>
	</front>
      </reference>

      <reference anchor="protobuf" target="https://developers.google.com/protocol-buffers/">
	<front>
	  <title>Protocol Buffers</title>
	  <author/>
	  <date month="July" year="2008"/>
	</front>
      </reference>

      <reference anchor="Smile" target="https://github.com/FasterXML/smile-format-specification">
	<front>
	  <title>Smile Data Format</title>
	  <author initials="T." surname ="Saloranta" fullname="Tatu Saloranta">
	    <address><email>tsaloranta@gmail.com</email></address>
	  </author>
	  <date month="September" year="2010"/>
	</front>
      </reference>

      <reference anchor="Thrift" target="http://thrift.apache.org/static/files/thrift-20070401.pdf">
	<front>
	  <title>Thrift: Scalable Cross-Language Services Implementation</title>
	  <author initials="M." surname ="Slee" fullname="Mark Slee">
	    <organization>Facebook</organization>
	    <address><email>mcslee@facebook.com</email></address>
	  </author>
	  <author initials="A." surname ="Agarwal" fullname="Aditya Agarwal">
	    <organization>Facebook</organization>
	    <address><email>aditya@facebook.com</email></address>
	  </author>
	  <author initials="M." surname ="Kwiatkowski" fullname="Marc Kwiatkowski">
	    <organization>Facebook</organization>
	    <address><email>marc@facebook.com</email></address>
	  </author>
	  <date month="April" year="2007"/>
	</front>
      </reference>

      <reference anchor="MsgPack" target="https://msgpack.org/">
	<front>
	  <title>MessagePack</title>
	  <author initials="S." surname="Furuhashi" fullname="Sadayuki Furuhashi"/>
	</front>
      </reference>

      <reference anchor="dhall-sec"
		 target="https://docs.dhall-lang.org/discussions/Safety-guarantees.html">
	<front>
	  <title>Dhall Safety Guarantees</title>
	  <author initials="G." surname="Gonzalez" fullname="Gabriella Gonzalez"/>
	</front>
      </reference>

    </references>


    <section anchor="bulktext" title="Using the text notation as a format">
      <t>
	BULK's text notation can be used as a full-fledged format alongside BULK's binary
	syntax.
      </t>
      <t>
	The text format has a different trade-off. On one hand, it is readily human-readable and it
	is easy to author in the absence of any tooling, because it is plain text. On the other
	hand, it is less robust in several ways: it is more dependent on the processing application
	knowing the definitions of namespaces and packages, it is rigid with respect to encoding (it
	MUST be encoded in UTF-8), and it is limited and inefficient in its representation of
	arbitrary bytes and strings. While parsing BULK text notation is a bit more involved than
	parsing binary BULK, the syntax of the text notation is still deliberately simple so as to
	limit even that part's processing footprint.
      </t>
      <t>
	Conceptually, parsing text notation involves translating the notation into binary BULK and
	then processing that. Each lexical element of BULK's text notation is separated from other
	elements by whitespace. Apart from <tt>([</tt> and <tt>])</tt>, delimiting an array
	containing a BULK stream, every lexeme can be immediately translated into its binary
	representation.
      </t>
      <t>
	Translating reference mnemonics involves knowing the mnemonics of previously imported
	namespaces, which means that as complete expressions are produced by the parser, they need
	to be evaluated. Reference mnemonics can appear without a namespace mnemonic if there aren't
	two imported namespaces that both use that mnemonic for a name. When a reference mnemonic
	appears with a known namespace mnemonic but an unknown name mnemonic, a processing
	application MUST associate that mnemonic with the first name in that namespace that doesn't
	already have a mnemonic. This makes it easy to author a namespace definition without having
	to manually number names.
      </t>
    </section>

    <section anchor="robustNS" title="Robust namespace definition">
      <t>
	This constitutes a suggestion of architecture for a BULK processing application. It has the
	advantage that an agent cannot modify the values of names to which it has not specifically
	been given authority. This architecture doesn't ensure this property by checking the
	validity of definitions but by adhering to the Principle Of Least Authority, thus ensuring
	no false positives or TOCTOU race conditions.
      </t>
      <t>
	For each new context (including the abstract yield when parsing starts), the parser creates
	a new copy of each known namespace. These copies are available in this context to retrieve
	and define values. It implements the lexical scoping of definitions on top of providing the
	robustness properties discussed here.
      </t>
      <t>
	By default, all namespaces created in a context are discarded at the end of this context.
      </t>
      <t>
	Of course, an implementation of the architecture presented here can be optimized compared
	to the abstract algorithm, for example by using copy-on-demand.
      </t>
      <t>
	Any namespace that is not a copy for its context but the object retained by the application
	afterwards, gives authority to make long-lasting definitions. A namespace that is stored by
	the processing application after evaluating a BULK stream is called a lasting namespace.
      </t>
      <t>
	Note that there are two ways to define a namespace in a BULK stream: using only the
	definition in the <tt>( define ( namespace {…} )</tt> form, or using this (possibly empty)
	definition as modified by other definitions (in the same BULK stream or not). The former is
	called the <em>initial definition</em>, the latter the <em>amended definition</em>.
      </t>
      <section title="Complete authority">
	<t>
	  When the amended definitions of all namespaces constitute lasting namespaces, it means
	  that the evaluation of the BULK stream can modify any existing namespace. This level of
	  authority might be useful for a handful of privileged BULK streams acting as configuration
	  of the application (e.g. to achieve reverse aliasing, see <xref target="fwdCompat"/>).
	</t>
      </section>
      <section title="Selective authority">
	<t>
	  A number of lasting namespaces are included for the abstract yield. Their unique
	  identifiers are agreed out-of-band. The disadvantage of this solution is that it needs
	  prior agreement on the definable namespaces. This may be a safer way than complete
	  authority to achieve reverse aliasing (see <xref target="fwdCompat"/>).
	</t>
      </section>
      <section title="Open authority">
	<t>
	  Any namespace definition for a unique identifier unknown to the processing application
	  triggers the creation of a lasting namespace.
	</t>
	<t>
	  The disadvantage of this solution is that it opens a denial of service vulnerability. If
	  Bob is a processing application and Carol and Dave are agents communicating with Bob with
	  an open authority, Dave can prevent Carol from defining a namespace if it manages to know
	  the unique identifier and to start a communication with Bob before Carol.
	</t>
	<t>
	  If an agent uses a secure way to create unique identifiers, this solution is both
	  flexible and safe (the burden is not on the BULK processing application). This
	  specification thus encourages the use of open authority restricted to verifiable
	  namespaces (in which case several agents can present the same definition to a processing
	  application without conflict).
	</t>
	<t>
	  A processing application could have a configuration setting to select what will generate
	  a lasting namespace:
	  <list style="hanging">
	    <t hangText="unrestricted open authority">any time a namespace is defined with an
	    identifier that was previously unknown, either its initial or amended definition
	    constitutes a lasting namespace; this is the most lax open authority, most flexible but
	    also most open to misuse and issues</t>
	    <t hangText="collision-free open authority">any time a namespace is defined with an
	    identifier that was previously unknown and the identifier relies explicitly on an
	    algorithm that the processing application deems giving a high enough guarantee that
	    identifiers are unique, either the initial or amended definition of that namespace
	    constitutes a lasting namespace</t>
	    <t hangText="immutable open authority">any time an immutable namespace is defined, its
	    initial definition constitutes a lasting namespace</t>
	  </list>
	</t>
	<t>
	  It is RECOMMENDED to use immutable open authority by default, as several agents can safely
	  present the same definition to a processing application without conflict.
	</t>
      </section>
    </section>

    <section anchor="fwdCompat" title="Forward compatibility">
      <t>
	BULK makes it possible to create new versions of vocabularies that encompass previous
	versions, in a way that minimizes implementation complexity.
      </t>

      <t>
	The first tool is aliasing: reuse names and values from existing namespaces, even in <xref
	target="bootstrap" format="none">bootstrapping namespaces</xref>:
      </t>

      <t><figure><artwork>( define ( namespace ( newhash:shake128 {newhashid} ) 20 )
  ([ ( version 1 0 )
  ( ( namespace 20 ) )
  ( mnemonic ( namespace 20 ) "newhash" )    
  ( explain ( namespace 20 ) "The new, shiny hash namespace!" )
  ( mnemonic newhash:shake128 "shake128" )
  ( import 21 ( namespace ( oldhash:shake128 {oldhashid} ) ) )
  ( define newhash:shake128 oldhash:shake128 ) ]) )</artwork></figure></t>
	  
      <t>
	With this, new namespaces can be created and applications don't need to change the existing
	code.
      </t>

      <t>
	One possible downside with aliasing is that if the number of aliasing namespaces grow, you
	might end up with the implementation of an important namespace scattered across a bunch of
	aliased legacy namespaces. Also, the definition of the new namespace is tied to the old one,
	which means that you need to keep the old definition around for the lifetime of the new
	one. To prevent those issues, a second tool is to reverse the direction of aliasing: all
	the implementation lives in the current namespace, cohesively, and its definition can be
	used on its own, and the old namespace is aliased to the new:
      </t>

      <t><figure><artwork>( import 20 ( namespace ( oldhash:shake128 {oldhashid} ) ) )
( import 21 ( namespace ( newhash:shake128 {newhashid} ) ) )
( define oldhash:shake128 newhash:shake128 )</artwork></figure></t>

      <t>
	Following the Principle of Least Authority, it should not be possible by default for the
	evaluation of any BULK stream to make lasting modifications to existing namespaces.
      </t>

      <t>
	One obvious design would be for the application to have a privileged storage for reverse
	aliasing namespace definitions, with those namespaces given complete authority or, better
	yet, each being given selective authority for a specific existing namespace (see <xref
	target="robustNS"/>). Where this could still not be deemed safe enough, reverse aliasing of
	namespaces could be defined in the application's code.
      </t>
    </section>

    <section anchor="arityForm" title="Arity-carrying forms">
      <t>
	Sometimes a vocabulary will include forms that can contain an arbitrary number of
	expressions. When such a form is used in postfix bytecode, the simplest solution is just to
	use a nested <tt>postfix</tt> form:
      </t>
      <t><figure><artwork>( define ( arity ) ( 2 black white comment ) )
( postfix
  game
  1 2 black
  ( postfix alternative
    "white tried an unorthodox opening" 3 4 white comment
    "a more classical opening would be" 8 9 white comment )
  2 3 black
  ( postfix alternative
    "white played a bad move" 4 5 white comment
    "white could have played a decent move" 5 6 white comment
    "white could have played a great move" 5 7 white comment ) )</artwork></figure></t>
      <t>
        The nested <tt>postfix</tt> form costs 4 bytes, compared to an equivalent postfix bytecode.
      </t>
      <t>
	If those 4 bytes add up to too much space through repetition, an application could define a
	form for the sole purpose of assigning it an arity, while the evaluation of the
	arity-carrying form would just replace it with the original one. For example, after
	evaluating the postfix bytecode transformation and the resulting form of the last expression
	of
      </t>
      <t><figure><artwork>( define alt/2 alternative )
( define alt/3 alternative )
( define ( arity ) ( 2 black white comment alt/2 ) ( 3 alt/3 ) )
( postfix
  game
  1 2 black
  "white tried an unorthodox opening" 3 4 white comment
  "a more classical opening would be" 8 9 white comment
  alt/2
  2 3 black
  "white played a bad move" 4 5 white comment
  "white could have played a decent move" 5 6 white comment
  "white could have played a great move" 5 7 white comment
  alt/3
  )</artwork></figure></t>
      <t>it would be transformed into</t>
      <t><figure><artwork>( game
  ( black 1 2 )
  ( alternative
    ( comment "white tried an unorthodox opening" ( white 3 4 ) )
    ( comment "a more classical opening would be" ( white 8 9 ) ) )
  ( black 2 3 )
  ( alternative
    ( comment "white played a bad move" ( white 4 5 ) )
    ( comment "white could have played a decent move" ( white 5 6 ) )
    ( comment "white could have played a great move" ( white 5 7 ) ) )
  ( white 4 5 ) )</artwork></figure>
      </t>
      <t>
	Such an arity-carrying form costs 10 or 13 bytes to be usable when it is added to an
	existing form defining arities. Which means that compared to the nested <tt>postfix</tt>
	form, it pays for itself if it is used only 3 or 4 times.
      </t>
    </section>

    <section title="The difference between BULK and BULK formats" anchor="design">
      <t>
	BULK aims at being a useful framework for a wide variety of formats, including low-level
	communication protocols, higher-level RPC or REST APIs, media files, rich documents,
	archives and efficient serialization of existing data models (like XML, JSON or RDF).
      </t>
      <t>
	This had several implications on its design.
      </t>

      <section title="Purposefully open: for generality">
	<t>
	  As such, BULK imposes no constraints on what kind of data can be represented. One is free
	  to design a BULK format that mandates the use of EBCDIC. In keeping with <xref
	  target="BCP18"/>, BULK chooses UTF-8 as the default string encoding and its core namespace
	  only permits designating encodings from the IANA charset registry<xref
	  target="IANA-Charsets"/>. But a BULK vocabulary would be free to define a new
	  <tt>windows-codepage</tt> form to use Windows Codepages instead.
	</t>
	<t>
	  Any protocol designed to transport human-readable text should be aware of <xref
	  target="RFC9839"/>, but BULK could be used to transport text emitted in a terminal, making
	  full use of control characters, or even to create a file with examples of ill-formed UTF-8
	  strings containing surrogates. As such, BULK doesn't limit what code points are allowed in
	  UTF-8 strings, or any other Unicode encoding.
	</t>
	<t>
	  BULK doesn't include a grammar to define BULK formats, but a BULK grammar vocabulary, that
	  could encode ABNF<xref target="RFC5234"/> or CDDL<xref target="RFC8610"/> in BULK, would
	  do well to add the ability to express "Unicode Scalars", "XML Characters" and "Unicode
	  Assignables" from <xref target="RFC9839"/>, as well as classes and profiles from
	  PRECIS<xref target="RFC8264"/>.
	</t>
      </section>

      <section title="Purposefully limited: for safety">
	<t>
	  For the same reason, BULK parsing and evaluation needed to be secure by default and
	  feature a security model that would be safe enough that it can be a secure foundation
	  almost everywhere.
	</t>
	<t>
	  This is why BULK syntax and the BULK core namespace can't directly express notions like
	  the inclusion of an outside BULK stream, or referencing a file or URI to access. This is
	  also why this specification limits the discoverability of bootstrapping namespaces and
	  strongly limits discoverabilty of non immutable namespaces and packages.
	</t>
	<t>
	  A BULK format that can express such a dangerous combination of actions as reading files
	  and making network connections SHOULD carefully consider the attack surface they present
	  and the threat models for the format's use cases, and explain those in detail in the
	  format's documentation. Dhall's Safety Guarantees<xref target="dhall-sec"/> are a prime
	  example.
	</t>
	<t>
	  The BULK core namespace can't directly express Turing complete functions, and not even
	  functions that can take different execution paths depending on their arguments. The
	  functions that can be expressed can only use their arguments in a static transformation,
	  by design. While this drastically limit BULK's expressivity, it also drastically limit the
	  attack surface on what arbitrary input can make the BULK parser or evaluator do, while
	  still offering a decent power of abstraction (e.g. one could write a substitution function
	  to unpack an array containing an IPv6 header into its fields, but not a substitution
	  function to unpack an IPv6 extension header according to its type).
	</t>
	<t>
	  The goal is that a BULK protocol or format designer should be able to trust that if they
	  use BULK in accordance with the safety recommendations of this specification, they don't
	  need to carefully weigh the benefits of evaluation vs. its dangers, like it has been the
	  case with a couple of previous formats.
	</t>
      </section>
    </section>

    <section title="Marking and extending media types with BULK" anchor="typing-medias">
      <t>
	One possible use of BULK is the ability to add metadata around a file. The most basic
	metadata is the media type and its parameters. Although many media types can be expressed
	with a file extension, this usually doesn't encode parameters like the charset used for a
	plain text file format. This is why most spreadsheet software present the user with a
	preview of a few rows while asking them for the encoding, when importing CSV data.
      </t>
      <t>
	A media type vocabulary would make it possible to encode parameters and provide names for
	some known parameter values.
      </t>
      <t>
	While a <tt>.md</tt> file extension only encodes the media type <tt>text/markdown</tt>, a
	short BULK header could encode the full media type <tt>text/markdown; charset=ISO-8859-15;
	variant=GFM</tt>:
      </t>
      <sourcecode>( version 1 0 )
( import 20 ( package ( shake128 #[8] 0xDABBED01 ) 2 ) )
( markdown ( iana-charset 111 ) "GFM" )</sourcecode>
      <t>
	While a <tt>.csv</tt> file extension only encodes the media type <tt>text/csv</tt>, a short
	BULK header could encode the full media type <tt>text/csv; charset=UTF-8; header=absent</tt>
      </t>
      <sourcecode>( version 1 0 )
( import 20 ( package ( shake128 #[8] 0xDABBED01 ) 2 ) )
( csv ( iana-charset 106 ) csv-header-absent )</sourcecode>
      <t>
	Media type metadata could be mixed with other metadata:
      </t>
      <sourcecode>( version 1 0 )
( import 20 ( package ( shake128 #[8] 0xDABBED02 ) 4 ) )
( description
  ( media-type ( csv ( iana-charset 106 ) csv-header-present ) )
  ( licence cc0 ) )</sourcecode>
    </section>
  </back>
</rfc>

