joetjen.net
EN DE
Grammar language

Grammar

Ichor reads .aether files with a hand-written lexer and reader. The grammar below states the same syntax in Aether itself.

Ichor 0.3

Pragmas and tokens

This grammar was written for these pages from Aether.Lexer and Aether.Reader in Ichor 0.3.0. Compiled with Ichor, it accepts exactly the files Ichor’s own reader accepts among the 30 Aether grammars in the Ichor, Cooper, dextrin, logos and xeger repositories, and it rejects the same broken headers, names, pragmas and bounds.

aether.aether
@grammar "aether"
@root grammar_file
@skip TRIVIA

COMMENT := ";" (!"\n" .)*
TRIVIA  := (SPACE | COMMENT)*

UPPER_NAME := [A-Z_] [A-Z0-9_]*
LOWER_NAME := [a-z] [a-z0-9_\-]*
NUMBER     := DIGIT+

ESCAPE     := "\\" ("x" HEX HEX | "u{" HEX+ "}" | [nrt] | [^a-zA-Z0-9])
STRING     := "\"" (ESCAPE | !"\"" .)* "\"" (("cs" | "i") ![a-zA-Z0-9_])?
POSIX      := "[:" ("alpha" | "alnum" | "digit" | "space" | "hex") ":]"
CHAR_CLASS := "[" "^"? (POSIX | ESCAPE | !"]" .)* "]"
REGEX      := "/" ("\\" . | "[" ("\\" . | !"]" .)* "]" | ![/\[] .)* "/"

Header and definitions

grammar_file := TRIVIA? header definition* TRIVIA?

header := "@grammar" name:STRING "@root" root:LOWER_NAME pragma*
pragma := "@skip" UPPER_NAME | "@noskip" | "@case_insensitive" | "@engine" LOWER_NAME

definition   := token_def | rule_def | keywords_def
token_def    := UPPER_NAME ":=" choice refine?
rule_def     := LOWER_NAME ":=" choice
keywords_def := "@keywords" UPPER_NAME "{" keyword ("," keyword)* "}"
keyword      := STRING "->" UPPER_NAME
refine       := "@refine" "(" STRING "," STRING ("," UPPER_NAME)* ")"

sequence stops in front of NAME :=. That lookahead is the whole reason a definition needs no terminator.

Expressions

choice     := sequence ("|" sequence)*
sequence   := (!def_head item)+
def_head   := (UPPER_NAME | LOWER_NAME) ":="
item       := "~"? term
term       := (LOWER_NAME ":")? postfix
postfix    := ("&" | "!") primary | primary quantifier?
quantifier := "*" | "+" | "?" | "{" NUMBER ("," NUMBER?)? "}"
primary    := STRING | CHAR_CLASS | REGEX | "." | UPPER_NAME | LOWER_NAME
            | "(" choice ")" | layout | native

layout     := ("@indent" | "@samecol") ("(" choice ")" | postfix)
native     := "@native" "(" STRING "," STRING ("," LOWER_NAME)* ")" hint?
hint       := "@hint" "(" hint_entry ("," hint_entry)* ")"
hint_entry := LOWER_NAME ":" (LOWER_NAME | "(" (LOWER_NAME ("," LOWER_NAME)*)? ")")

Beyond syntax

Ichor’s reader checks more than this grammar can express. It also rejects:

  • a character class, . or /regex/ in a rule, and a rule name in a token;
  • ~ anywhere but directly before a bare name, in a token, or under @noskip;
  • a pragma given twice, @skip together with @noskip, and an @engine other than peg, lr or glr;
  • a name declared twice, and both @keywords and @refine on one token;
  • in @hint, keys other than nullable and leading, and a leading: name that is not a dependency.

Undefined references, redeclared predefined tokens and left recursion are checked once the whole grammar has been read.

Where the implementation differs

ReferenceIchor 0.3.0Write
Captures belong in rules onlyA capture in a token body is accepted when the grammar is read, then crashes the matcherNever name a part of a token
Both backends run every @engineThe interpreted backend runs only peg; lr and glr need their own runnersCompile lr and glr grammars ahead of time