joetjen.net
EN DE
Grammar language

The Aether grammar language

Aether describes a language completely in one file: its tokens, its syntax, how whitespace is treated and which parts of a match are kept. Upper-case names are tokens, matched by a lexer; lower-case names are rules, matched by a parser over those tokens.

Ichor 0.3

A first look

This grammar reads a small settings file — lines of name = value, with comments:

settings.aether
@grammar "settings"
@root document
@skip TRIVIA

COMMENT := "#" (!"\n" .)*
TRIVIA  := (SPACE | COMMENT)*

TRUE   := "true"
FALSE  := "false"
NAME   := [a-z_] [a-z0-9_]*
NUMBER := DIGIT+
STRING := "\"" (!"\"" .)* "\""

document := TRIVIA? entry* TRIVIA?
entry    := key:NAME "=" value:value
value    := TRUE | FALSE | NUMBER | STRING
  • @grammar names the grammar and @root the rule matching starts from. Both are required, in that order.
  • @skip TRIVIA lets whitespace and comments stand between any two parts of a rule.
  • COMMENT to STRING are tokens. TRUE comes before NAME on purpose: both match true, and the earlier one wins.
  • document, entry and value are rules. The quoted "=" becomes a token of its own without being declared.
  • key: and value: name the parts of an entry that matter to whatever processes the match.
# service
port = 8080
debug = true
name = "api"
ok
true = 1
1:1: unexpected "true" -- did not expect more input here
1 | true = 1
  | ^

Design

  • Lexer and parser in one file. Spelling decides which is which, so there is no separate token file to keep in step.
  • Two stages, always. Tokens take the longest match, rules take the first alternative that fits — each stage keeps its own simple rule.
  • Whitespace is a setting, not a chore. One pragma makes every rule tolerate it; one prefix forbids it where two parts must touch.
  • Syntax only. A grammar says what is valid and names the parts; what they mean is decided elsewhere.
  • Escape hatches, clearly marked. Where a language changes its own syntax mid-file, @native hands one token or rule to hand-written code.

File facts

PropertyValue
Extension.aether
EncodingUTF-8; names are ASCII, strings and classes may hold any character
Header@grammar "name", then @root rule — mandatory, in that order
Comments; to the end of the line
ResultA lexer and a parser — PEG by default, LR or GLR on request
ImplementationIchor, which compiles .aether files

Written in Aether

GrammarDefinesUsed by
casc.aetherthe CASC configuration format — grammarcooper
dxn.aetherthe DXN text format — grammardextrin
logos.aetherProlog sourcelogos
xeger.aetherregular expressionsxeger
abnf.aether …ABNF, BNF, ISO and W3C EBNF, and PEGIchor’s own importers

On these pages

  • File structure — header, pragmas, comments, names, definitions and predefined tokens.
  • Expressions — operators, strings, character classes, repetition, lookahead and regex literals.
  • Lexing and parsing — longest match, ordered choice, whitespace, captures, layout and engines.
  • Escape hatches — @keywords, @refine and @native.
  • Grammar — Aether written in Aether, and where the implementation differs from the reference.