Grammar language
The Aether grammar language
Aether describes a language completely in one file: its tokens, its syntax, how whitespace is treated and which parts of a match are kept. Upper-case names are tokens, matched by a lexer; lower-case names are rules, matched by a parser over those tokens.
A first look
This grammar reads a small settings file — lines of name = value, with comments:
@grammar "settings" @root document @skip TRIVIA COMMENT := "#" (!"\n" .)* TRIVIA := (SPACE | COMMENT)* TRUE := "true" FALSE := "false" NAME := [a-z_] [a-z0-9_]* NUMBER := DIGIT+ STRING := "\"" (!"\"" .)* "\"" document := TRIVIA? entry* TRIVIA? entry := key:NAME "=" value:value value := TRUE | FALSE | NUMBER | STRING
@grammarnames the grammar and@rootthe rule matching starts from. Both are required, in that order.@skip TRIVIAlets whitespace and comments stand between any two parts of a rule.COMMENTtoSTRINGare tokens.TRUEcomes beforeNAMEon purpose: both matchtrue, and the earlier one wins.document,entryandvalueare rules. The quoted"="becomes a token of its own without being declared.key:andvalue:name the parts of an entry that matter to whatever processes the match.
# service port = 8080 debug = true name = "api"
ok
true = 1
1:1: unexpected "true" -- did not expect more input here 1 | true = 1 | ^
Design
- Lexer and parser in one file. Spelling decides which is which, so there is no separate token file to keep in step.
- Two stages, always. Tokens take the longest match, rules take the first alternative that fits — each stage keeps its own simple rule.
- Whitespace is a setting, not a chore. One pragma makes every rule tolerate it; one prefix forbids it where two parts must touch.
- Syntax only. A grammar says what is valid and names the parts; what they mean is decided elsewhere.
- Escape hatches, clearly marked. Where a language changes its own syntax mid-file,
@nativehands one token or rule to hand-written code.
File facts
| Property | Value |
|---|---|
| Extension | .aether |
| Encoding | UTF-8; names are ASCII, strings and classes may hold any character |
| Header | @grammar "name", then @root rule — mandatory, in that order |
| Comments | ; to the end of the line |
| Result | A lexer and a parser — PEG by default, LR or GLR on request |
| Implementation | Ichor, which compiles .aether files |
Written in Aether
On these pages
- File structure — header, pragmas, comments, names, definitions and predefined tokens.
- Expressions — operators, strings, character classes, repetition, lookahead and regex literals.
- Lexing and parsing — longest match, ordered choice, whitespace, captures, layout and engines.
- Escape hatches —
@keywords,@refineand@native. - Grammar — Aether written in Aether, and where the implementation differs from the reference.