Grammar
The DXN text format is defined by a PEG grammar written in Aether, the grammar language of Ichor. dextrin’s parser is generated from it rather than written by hand.
Pragmas and tokens
This is the grammar dextrin’s parser is generated from, priv/grammar/dxn.aether as released with dextrin 0.1.3, with its comments removed. Upper-case names are tokens, matched longest-first with ties going to the earlier declaration; lower-case names are rules, matched over the tokens by ordered choice.
@grammar "dxn"
@root document
@skip TRIVIA
SPACE := [ \t\r\n,]+
COMMENT := "#" (!"\n" .)* "\n"?
TRIVIA := (SPACE | COMMENT)*
NIL := "nil"
TRUE := "true"
FALSE := "false"
EXPONENT := [eE] [+-]? DIGIT+
FLOAT := ("-"? DIGIT+ (("." DIGIT+ EXPONENT?) | EXPONENT)) | "NaN" | "Infinity" | "-Infinity"
DECIMAL := "-"? DIGIT+ ("." DIGIT+)? "M"
RATIONAL := "-"? DIGIT+ "/" DIGIT+
INTEGER := "-"? DIGIT+
IDENT_START := [_\u{41}-\u{5A}\u{61}-\u{7A}\u{AA} … \u{30000}-\u{3134A}\u{31350}-\u{33479}]
IDENT_CONT := IDENT_START | [\u{30}-\u{39}\u{41}-\u{5A}\u{5F}\u{61}-\u{7A} … \u{E0100}-\u{E01EF}\-?!]
IDENTIFIER := IDENT_START IDENT_CONT* ("/" IDENT_START IDENT_CONT*)?
MAP_KEY := IDENTIFIER ":"
KEYWORD := ":" (IDENTIFIER | STRING)
AT_DISCARD := "@_"
AT := "@"
ESCAPE := "\\" ("\"" | "\\" | "n" | "t" | "r" | "0" | "a" | "b" | "f" | "v" | ("x{" HEX+ "}"))
STRING := "\"" (ESCAPE | (![\"\\] .))* "\""
CHAR := "?" (ESCAPE | "s" | (![ \t\r\n] .))
DATE_SIGIL := "~D[" (!"]" .)* "]"
TIME_SIGIL := "~T[" (!"]" .)* "]"
INSTANT_SIGIL := "~U[" (!"]" .)* "]"
REGEX_SIGIL := "~r/" (("\\" .) | (!"/" .))* "/" [imsuxfr]*Declaration order does real work here. NIL, TRUE and FALSE come before IDENTIFIER so they win the tie. MAP_KEY fuses a name and its colon into one token, so name:value without a space can never be read as a name followed by the keyword :value. AT_DISCARD beats AT by length, so @_ is always a discard and never a custom tag named _.
Rules
document := header? value TRIVIA?
header := "@dxn" version:STRING
value := discard* value_body
discard := AT_DISCARD value
value_body := nil | boolean | number | string | char | keyword | symbol
| list | tuple | map_lit | struct_lit | set_lit
| sigil | tag_form
nil := NIL
boolean := TRUE | FALSE
number := DECIMAL | RATIONAL | FLOAT | INTEGER
string := STRING
char := CHAR
symbol := IDENTIFIER
keyword := KEYWORD
list := "[" value* "]"
tuple := "{" value* "}"
set_lit := AT "{" value* "}"
map_lit := "%" "{" map_entry* "}"
map_entry := (short_key:MAP_KEY short_val:value) | (arrow_key:value "=>" arrow_val:value)
struct_lit := "%" name:IDENTIFIER (keyed:struct_keyed | positional:struct_positional)
struct_keyed := "{" map_entry* "}"
struct_positional := "[" value* "]"
sigil := DATE_SIGIL | TIME_SIGIL | INSTANT_SIGIL | REGEX_SIGIL
tag_form := AT name:IDENTIFIER valueThe sigil tokens capture their bodies whole; dates, times and regular expressions are checked after parsing, by the host language’s own parsers. Because "@dxn" is a longer token than AT everywhere in a document, not only at its start, @dxn can never be a custom tag.
Where the implementation differs
In a few places dextrin 0.1.3 accepts more than the specification, or settles something the specification leaves open. These are the portable choices:
| Specification | dextrin 0.1.3 | Write |
|---|---|---|
~U[…] is UTC only, ending in Z | Accepts an offset and converts to UTC | Always Z; @datetime for offsets |
@datetime needs a non-UTC offset | Accepts Z and +00:00, and writes the value back as a timestamp | ~U[…] for UTC |
| The header names the format version | Any version string is accepted unchecked | @dxn "1.0" |
| Binary structs are positional, in schema order | Without a schema, a keyed struct is written in source order and loses its names: %Point{y: 2, x: 1} comes back as %Point[2 1] | Positional form, or encode with the schema loaded |
| Tag 29 refers to “the n-th marked item” | Numbers marked items as they finish decoding, innermost first; the registered CBOR extension numbers them as they open | Leave sharing off for files that other CBOR tools read |