File structure
A grammar file is a header followed by definitions. The header names the grammar and its starting rule and sets grammar-wide options; each definition gives a token or a rule its body.
The header
Every file begins with @grammar and a name in quotes, then @root and the rule matching starts from. Optional pragmas follow, in any order, each at most once:
| Pragma | Effect |
|---|---|
@skip TOKEN | skip this token between the parts of a rule — see whitespace |
@noskip | skip nothing; whitespace is syntax |
@case_insensitive | quoted literals without a suffix ignore case |
@engine peg | lr | glr | the parsing algorithm — see engines |
@skip and @noskip exclude each other. Without either, the grammar behaves as if it said @skip SPACE. All pragmas belong to the header; one written after the first definition is an error:
@grammar "x" @root a a := "a" @skip SPACE
4:1: expected a token or rule definition (NAME := ...)
Comments
A ; starts a comment that runs to the end of the line, anywhere in the file. There are no block comments. Spaces, tabs and line breaks separate the parts of a grammar and mean nothing else — a definition may span as many lines as it likes.
Names
| Spelling | Declares | Examples |
|---|---|---|
[A-Z_][A-Z0-9_]* | a token | NUMBER DQ_STRING _WS |
[a-z][a-z0-9_-]* | a rule | expr map_entry key-path |
Any other spelling is an error, not a third kind of name:
Foo := "a"
3:1: invalid identifier "Foo" -- must be ALL_UPPERCASE (a token) or snake_case/kebab-case (a rule)
Definitions
A definition is a name, := and an expression. There is no terminator: a definition ends where the next NAME := begins. Definitions come in any order, and a body may use a name declared further down. Declaring a name twice, or using one that is never declared, is an error:
a := "a" a := "b"
4:1: rule a is already declared
a := b
3:6: reference to undefined token or rule :b
Order still matters among tokens, because it settles ties in the lexer — see longest match.
Predefined tokens
| Token | Default |
|---|---|
DIGIT | [0-9] |
ALPHA | [a-zA-Z] |
ALNUM | DIGIT | ALPHA |
SPACE | [ \t\r\n] |
HEX | [a-fA-F0-9] |
All five exist in every grammar without being declared. Each may be redeclared, but only before its first use — and @skip SPACE, written or implied, counts as a use. This grammar narrows SPACE to spaces and tabs, so a line break is no longer whitespace:
@grammar "words" @root line SPACE := [ \t]+ WORD := ALPHA+ line := WORD+
"two words" → ok "two\nlines" → 1:4: no token matches here
With @skip SPACE written above the redeclaration, the grammar is rejected instead:
@grammar "words" @root line @skip SPACE SPACE := [ \t]+
4:1: cannot override SPACE: already used at line 3, column 7