Expressions
The body of every definition is an expression built from the same small set of operators. What may appear inside depends on whether the body belongs to a token or to a rule.
Operators
| Form | Meaning |
|---|---|
a b | sequence: a, then b |
a | b | ordered choice: a, or if that fails, b |
a* a+ a? | zero or more, one or more, zero or one — always greedy |
a{3} a{3,} a{3,7} | exactly 3, at least 3, from 3 to 7 |
&a !a | only if a follows, only if it does not — consuming nothing |
name:a | capture a under name |
~a | no skipped whitespace before a |
( a ) | grouping |
Choice binds loosest, then sequence; a capture, ~, & and ! apply to the single term after them, and a quantifier to the term before it. & and ! cannot carry a quantifier themselves.
Tokens and rules
| Element | In a token | In a rule |
|---|---|---|
"literal" | yes | yes — becomes an anonymous token |
| token name | yes | yes |
| rule name | no | yes |
[class] . /regex/ | yes | no |
name: ~ @indent @samecol | no | yes |
The split follows from the two stages: a rule only ever sees tokens, never characters. A pattern a rule needs gets a token name first.
a := [a-z]+
3:6: rules may not contain an inline character class -- give it a named TOKEN instead
A := a
3:6: a token body may only reference other tokens, not rule "a"
String literals
Strings are written in double quotes and may contain any character. The escapes are \n, \r, \t, \\, \", \[, \], \/, \xHH with exactly two hex digits, and \u{H…} with one or more; any other punctuation character escaped stands for itself, so \- and \* work too.
A suffix decides case sensitivity for one literal: "lit"i always ignores case, "lit"cs never does. Without a suffix, a literal follows @case_insensitive if the grammar has it. Character classes are not affected.
@grammar "literal" @root value @case_insensitive value := "true" | "false" | "null"cs
"TRUE" → ok "False" → ok "null" → ok "NULL" → 1:1: no token matches here
Character classes
[…] matches one character from a set of characters and a-z ranges; [^…] matches one character outside it, but never the end of the input. A - right before ] is a literal dash. Inside the brackets, [:alpha:], [:alnum:], [:digit:], [:space:] and [:hex:] stand for the predefined tokens. . matches any one character.
IDENT := [[:alpha:]_][[:alnum:]_]*
"_x1" → ok "1x" → 1:1: no token matches here
Repetition and lookahead
Repetition is greedy and never gives back what it took. Inside a token, choice is ordered too, which this colour token shows: #ff88 is neither six digits nor three, so HEX{3} takes ff8 and the last 8 is left over.
COLOR := "#" (HEX{6} | HEX{3})"#ff8800" → ok "#f80" → ok "#ff88" → 1:5: no token matches here
Lookahead checks without consuming. Here IF matches if only when no letter or digit follows, so iffy stays an identifier:
IF := "if" !ALNUM IDENT := ALPHA ALNUM* stmt := IF IDENT | IDENT
"if x" → ok "iffy" → ok
Regex literals
In a token body, /pattern/ is shorthand for the equivalent Aether expression. It supports characters, ., classes, |, groups, all the quantifiers, (?=…) and (?!…), and \d \w \s \h with their negations, which stand for DIGIT, ALNUM, SPACE and HEX. Anchors, back-references and capturing groups have no PEG meaning and are rejected outright:
NUMBER := /\d+(\.\d+)?/
"3.14" → ok "3." → 1:2: no token matches here
A := /^a+$/
3:6: anchors (^) are not supported in /pattern/