joetjen.net
EN DE
Grammar language

Expressions

The body of every definition is an expression built from the same small set of operators. What may appear inside depends on whether the body belongs to a token or to a rule.

Ichor 0.3

Operators

FormMeaning
a bsequence: a, then b
a | bordered choice: a, or if that fails, b
a* a+ a?zero or more, one or more, zero or one — always greedy
a{3} a{3,} a{3,7}exactly 3, at least 3, from 3 to 7
&a !aonly if a follows, only if it does not — consuming nothing
name:acapture a under name
~ano skipped whitespace before a
( a )grouping

Choice binds loosest, then sequence; a capture, ~, & and ! apply to the single term after them, and a quantifier to the term before it. & and ! cannot carry a quantifier themselves.

Tokens and rules

ElementIn a tokenIn a rule
"literal"yesyes — becomes an anonymous token
token nameyesyes
rule namenoyes
[class] . /regex/yesno
name: ~ @indent @samecolnoyes

The split follows from the two stages: a rule only ever sees tokens, never characters. A pattern a rule needs gets a token name first.

a := [a-z]+
3:6: rules may not contain an inline character class -- give it a named TOKEN instead
A := a
3:6: a token body may only reference other tokens, not rule "a"

String literals

Strings are written in double quotes and may contain any character. The escapes are \n, \r, \t, \\, \", \[, \], \/, \xHH with exactly two hex digits, and \u{H…} with one or more; any other punctuation character escaped stands for itself, so \- and \* work too.

A suffix decides case sensitivity for one literal: "lit"i always ignores case, "lit"cs never does. Without a suffix, a literal follows @case_insensitive if the grammar has it. Character classes are not affected.

@grammar "literal"
@root value
@case_insensitive
value := "true" | "false" | "null"cs
"TRUE"  → ok
"False" → ok
"null"  → ok
"NULL"  → 1:1: no token matches here

Character classes

[…] matches one character from a set of characters and a-z ranges; [^…] matches one character outside it, but never the end of the input. A - right before ] is a literal dash. Inside the brackets, [:alpha:], [:alnum:], [:digit:], [:space:] and [:hex:] stand for the predefined tokens. . matches any one character.

IDENT := [[:alpha:]_][[:alnum:]_]*
"_x1" → ok
"1x"  → 1:1: no token matches here

Repetition and lookahead

Repetition is greedy and never gives back what it took. Inside a token, choice is ordered too, which this colour token shows: #ff88 is neither six digits nor three, so HEX{3} takes ff8 and the last 8 is left over.

COLOR := "#" (HEX{6} | HEX{3})
"#ff8800" → ok
"#f80"    → ok
"#ff88"   → 1:5: no token matches here

Lookahead checks without consuming. Here IF matches if only when no letter or digit follows, so iffy stays an identifier:

IF    := "if" !ALNUM
IDENT := ALPHA ALNUM*
stmt  := IF IDENT | IDENT
"if x" → ok
"iffy" → ok

Regex literals

In a token body, /pattern/ is shorthand for the equivalent Aether expression. It supports characters, ., classes, |, groups, all the quantifiers, (?=…) and (?!…), and \d \w \s \h with their negations, which stand for DIGIT, ALNUM, SPACE and HEX. Anchors, back-references and capturing groups have no PEG meaning and are rejected outright:

NUMBER := /\d+(\.\d+)?/
"3.14" → ok
"3."   → 1:2: no token matches here
A := /^a+$/
3:6: anchors (^) are not supported in /pattern/