Readable Regular Expressions for JavaScript/TypeScript, Inspired by Emacs' rx

Quick, what does this match?

/^(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)(?:-((?:0|[1-9]\d*|\d*[a-zA-Z-][0-9a-zA-Z-]*)(?:\.(?:0|[1-9]\d*|\d*[a-zA-Z-][0-9a-zA-Z-]*))*))?(?:\+([0-9a-zA-Z-]+(?:\.[0-9a-zA-Z-]+)*))?$/;

// Take

// your

// time...

//

// ...still decoding?

//

// OK, keep reading :)

That's the official regexp from semver.org. It validates version numbers like:

// matches

"1.2.3"

"0.10.0"

"2.0.0-rc.1"

"1.0.0-alpha.1+build.5"

"1.0.0+20260930"

// doesn't match

"01.2.3" // leading zero

"1.2" // missing patch

"v1.2.3" // no "v" prefix allowed

"1.0.0-01" // numeric pre-release with a leading zero

"1.2.3-" // empty pre-release

Don't get me wrong, I love regexps, but in practice you probably spend a bunch of time writing one, testing it against some cases, and moving on, proud of your achievement!

Some time passes and lucky future you (or unlucky someone else) has to change it. Dramatic pause here.

I bet you've been there. Now your options are probably: decode it again from the start, rewrite the whole thing, or, in the age of AI, ask (and hopefully not blindly accept) an LLM for a new recipe.

Emacs has had a nice answer for more readable regexps for a long time:

the rx macro. I started using it all the time in Emacs Lisp, as

reviewers always suggested it to me. Later, I started missing this DSL

in JavaScript and TypeScript, so I wrote a small version of it for my

projects.

So, what about reading that SemVer regexp like semver in the code

below?

const num = or("0", seq(anyOf("1-9"), zeroOrMore(digit)));

const idChar = anyOf(alnum, "-");

const preId = or(num, seq(zeroOrMore(digit), anyOf(alpha, "-"), zeroOrMore(idChar)));

const dotted = (x: Item) => seq(x, zeroOrMore(".", x));

const semver = RX(

start,

named("major", num), ".",

named("minor", num), ".",

named("patch", num),

optional("-", named("pre", dotted(preId))),

optional("+", named("build", dotted(oneOrMore(idChar)))),

end,

);

The same strings match, and you get named groups as a bonus. By the end of this post you'll know every piece of it.

TL;DR: jump straight to the cheat sheet, the side-by-side examples, the full source, or grab the gist to sneak a peek at the result.

NOTE: the

RXhere has nothing to do with RxJS, which is an amazing library for reactive programming with observables.

With rx you describe a regexp as a tree of named forms, and Emacs

turns it into the regexp string for you:

(rx bos (+ digit) eos)

;; => "\\`[[:digit:]]+\\'"

(rx bol "colo" (? "u") "r" eol)

;; => "^colou?r$"

(rx bos "(" (= 3 digit) ")" space (= 3 digit) "-" (= 4 digit) eos)

;; => "\\`([[:digit:]]\\{3\\})[[:space:]][[:digit:]]\\{3\\}-[[:digit:]]\\{4\\}\\'"

A few things to notice:

- Strings are literals. "("means a parenthesis. You don't need to escape anything by hand.

- Sequence is implicit. Every form takes a list of things and

matches them one after the other. You don't need to wrap them in a

seq, even thoughseqexists.

- Groups appear only when needed. (+ digit)becomes[[:digit:]]+, not\(?:[[:digit:]]\)+.

The proposed JavaScript/TypeScript version in this post reads like this:

const phone = RX(

start, "(", repeat(3, digit), ")", space,

repeat(3, digit), "-", repeat(4, digit), end,

);

// => /^\(\d{3}\)\s\d{3}-\d{4}$/

If you want strings to be literals, you can't represent a regexp piece

as a plain string, otherwise you can't tell "(" (a literal

parenthesis) apart from "(?:...)" (a group you built). So every

piece is a small object:

type Kind = "atom" | "seq" | "alt";

interface RxNode {

readonly src: string;

readonly kind: Kind;

readonly set?: string; // char sets only, see below

readonly neg?: boolean;

}

type Item = string | RxNode;

src is the regexp text. kind records how that text behaves when

you glue it to other things:

- atom: a single unit, like- a,- \d,- [a-z]or- (...). You can put a quantifier right after it.

- seq: safe to concatenate, but a quantifier needs- (?:...)around it.- abcis a- seq, and so is- a+, since- a+?would silently turn into a lazy quantifier.

- alt: has a- |at the top level, so it needs- (?:...)almost everywhere.

Plain strings go through literal, which escapes them:

const esc = (s: string) => s.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");

const literal = (s: string): RxNode => ({

src: esc(s),

kind: s.length === 1 ? "atom" : "seq",

});

const toNode = (x: Item): RxNode => (typeof x === "string" ? literal(x) : x);

With that in place, seq joins nodes and only brackets alternations:

const seq = (...xs: Item[]): RxNode => {

const nodes = xs.map(toNode).filter((n) => n.src !== "");

if (nodes.length === 0) return { src: "", kind: "seq" };

if (nodes.length === 1) return nodes[0];

let src = "";

for (const n of nodes) {

const part = n.kind === "alt" ? `(?:${n.src})` : n.src;

// `\1` followed by a literal `0` would read as `\10`

if (/\\\d+$/.test(src) && /^\d/.test(part)) src += "(?:)";

src += part;

}

return { src, kind: "seq" };

};

(That backreference check is one of those bugs you only find by

writing tests, or when it happens to you in prod. backref(1)

followed by the literal "0" gives you backreference number ten.)

Every quantifier is a seq of its arguments plus a suffix, bracketed

only when the body isn't an atom:

const quantifiable = (n: RxNode) =>

n.kind === "atom" ? n.src : `(?:${n.src})`;

const quantifier =

(suffix: string) =>

(...xs: Item[]): RxNode => ({

src: quantifiable(seq(...xs)) + suffix,

kind: "seq",

});

const zeroOrMore = quantifier("*");

const oneOrMore = quantifier("+");

const optional = quantifier("?");

Because each quantifier calls seq on its arguments, you get the

implicit sequence for free: optional("-", group(x)) becomes

(?:-(x))?.

And finally, the two entry points. As in Emacs, rx returns a

string. RX returns a RegExp you can use right away:

const rx = (...xs: Item[]): string => seq(...xs).src;

function RX(...xs: Item[]): RegExp {

return new RegExp(rx(...xs));

}

RX.flags = (flags: string, ...xs: Item[]): RegExp =>

new RegExp(rx(...xs), flags);

RX.flags exists because Emacs controls case folding through the

case-fold-search variable, and JavaScript puts it on the regexp

itself.

That's the whole engine! Now, let's build our vocabulary.

In Emacs you write (any "a-z" "_"). Inside those strings, a-z is a

range, and a - at either end is a plain dash. I kept the same rule:

const hexDigit = anyOf("0-9a-fA-F");

RX(start, "#", repeat(6, hexDigit), end);

// => /^#[0-9a-fA-F]{6}$/

RX(start, optional(anyOf("+-")), oneOrMore(digit), end);

// => /^[+\-]?\d+$/

The dash comes out escaped because sets can merge. If you combine

anyOf("+-") with anyOf("0-9"), an unescaped - would end up in

the middle and create a range from + to 0. Escaping it costs one

backslash.

And merging is the reason why RxNode has a set field. It is there

to hold the text that goes between [ and ], so anyOf can take

other sets as arguments:

const lower = anyOf("a-z");

const upper = anyOf("A-Z");

const alpha = anyOf(lower, upper); // [a-zA-Z]

const alnum = anyOf(alpha, "0-9"); // [a-zA-Z0-9]

not negates a set, and it knows the shorthand classes:

not(digit); // \D

not(anyOf(space, "@")); // [^\s@]

notChar(","); // [^,] (rx's not-char)

The simple email check, which most of us have written as

/^[^\s@]+@[^\s@]+\.[^\s@]+$/ at some point, becomes:

const part = oneOrMore(not(anyOf(space, "@")));

RX(start, part, "@", part, ".", part, end);

// => /^[^\s@]+@[^\s@]+\.[^\s@]+$/

The rest of the Emacs character classes are there too: digit,

hexDigit, space, blank, wordChar, notWordChar, alpha,

alnum, lower, upper, punct, control, graphic, printing,

ascii and nonascii. One difference: in Emacs they understand

Unicode, and mine are ASCII only. alpha won't match é.

Two more come from rx's symbol list, and people (me, many times) mix them up:

const notNewline: RxNode = { src: ".", kind: "atom" }; // rx: nonl

const anything = set("\\s\\S"); // rx: anything / anychar

In rx, anything really means anything, newlines included. Here is

where the difference shows up:

const code = "a = 1; /* first\n second */ b = 2;";

RX("/*", zeroOrMoreLazy(notNewline), "*/").exec(code);

// => null

RX("/*", zeroOrMoreLazy(anything), "*/").exec(code)?.[0];

// => "/* first\n second */"

or works as you'd expect, and gets bracketed when it lands inside a

sequence:

RX(start, or("cat", "dog", "bird"), end);

// => /^(?:bird|cat|dog)$/

Did you notice the order changed? I copied that behavior from

Emacs. When every branch of an or is a plain string, rx hands them

to regexp-opt, which builds a pattern that prefers the longest

match:

(rx (or "in" "int" "interface"))

;; => "\\(?:in\\(?:t\\(?:erface\\)?\\)?\\)"

JavaScript alternation takes the first branch that matches, going left to right. So the naive regexp for a list of keywords has a 'bug':

/in|int|interface/.exec("interface Foo")?.[0];

// => "in"

RX(or("in", "int", "interface")).exec("interface Foo")?.[0];

// => "interface"

I don't build a trie like regexp-opt does. Sorting the strings by

length, longest first, is enough to get the same behavior:

const or = (...xs: Item[]): RxNode => {

if (xs.length === 0) return unmatchable;

if (xs.length === 1) return toNode(xs[0]);

const branches = xs.every((x) => typeof x === "string")

? [...(xs as string[])].sort((a, b) => b.length - a.length)

: xs;

return { src: branches.map((x) => toNode(x).src).join("|"), kind: "alt" };

};

As in Emacs, or() with no branches returns unmatchable, which is

(?!) here. It's handy when you build the branch list at runtime and

it might come out empty.

Emacs has (= n ...), (>= n ...) and (** n m ...). Here they are

repeat, atLeast and between:

RX(start, between(2, 4, digit), end); // /^\d{2,4}$/

RX(atLeast(3, digit)); // /\d{3,}/

The lazy versions *?, +? and ?? are zeroOrMoreLazy,

oneOrMoreLazy and optionalLazy. The classic HTML tag example:

const html = "<b>bold</b> and <i>italic</i>";

RX("<", oneOrMore(notNewline), ">").exec(html)?.[0];

// => "<b>bold</b> and <i>italic</i>"

RX("<", oneOrMoreLazy(notNewline), ">").exec(html)?.[0];

// => "<b>"

group is a capturing group, and backref points back to it:

RX(start, group(oneOrMore(wordChar)), space, backref(1), end);

// => /^(\w+)\s\1$/ matches "hello hello", not "hello world"

Emacs also has (group-n N ...) to pick the group number. JavaScript

can't do that, but it has named groups, which serve the same purpose

and read better:

const date = RX(

named("y", repeat(4, digit)), "-",

named("m", repeat(2, digit)), "-",

named("d", repeat(2, digit)),

);

date.exec("2026-09-30")?.groups;

// => { y: '2026', m: '09', d: '30' }

backref accepts a name as well:

RX(

"<", named("tag", oneOrMore(wordChar)), ">",

zeroOrMoreLazy(notNewline),

"</", backref("tag"), ">",

);

// => /<(?<tag>\w+)>.*?<\/\k<tag>>/

rx distinguishes the start of the string (bos) from the start of a

line (bol). In JavaScript both are ^, and the m flag decides

which one you get. I kept both names so the intent shows in the code:

const text = "TODO: write post\nDONE: fix rx\nTODO: publish";

const todo = RX.flags(

"gm",

lineStart, "TODO: ", named("task", oneOrMore(notNewline)), lineEnd,

);

[...text.matchAll(todo)].map((m) => m.groups?.task);

// => [ 'write post', 'publish' ]

Why not add start and end automatically? Because you only want

them when validating a whole string. When searching inside a text, as

in split, replace or matchAll, a hidden ^ and $ would break

everything. Emacs agrees: bos and eos are explicit in rx too.

wordBoundary and notWordBoundary map straight to \b and \B.

Emacs also has bow and eow (\< and \>), start and end of a

word. JavaScript lacks those, so I combined \b with a lookaround:

const wordStart = zeroWidth("\\b(?=\\w)");

const wordEnd = zeroWidth("\\b(?<=\\w)");

Plain strings are already literals, but rx has an explicit literal

form for strings computed at runtime, and I kept it. It documents that

the value came from somewhere else:

const userInput = "1+1=2? (maybe)";

new RegExp(userInput).test(userInput); // false, oops

RX(literal(userInput)).test(userInput); // true

The opposite direction is rx's (regexp ...) form, the escape hatch.

Here it's raw, and it receives either a string or an existing

RegExp. It lets you adopt the DSL in a codebase full of old regexps

without rewriting all of them, like:

const legacyZip = /\d{5}(?:-\d{4})?/;

RX(start, repeat(2, upper), " ", raw(legacyZip), end);

// => /^[A-Z]{2} (?:\d{5}(?:-\d{4})?)$/

raw can't see inside the text it gets, so it adds brackets whenever

it's combined with something else. It's an extra (?:), and the

regexp still works.

Back to the regexp from the intro. In Emacs you would give names to

the pieces with rx-define or rx-let. In TypeScript those are just

consts:

const num = or("0", seq(anyOf("1-9"), zeroOrMore(digit)));

const idChar = anyOf(alnum, "-");

const preId = or(num, seq(zeroOrMore(digit), anyOf(alpha, "-"), zeroOrMore(idChar)));

const dotted = (x: Item) => seq(x, zeroOrMore(".", x));

const semver = RX(

start,

named("major", num), ".",

named("minor", num), ".",

named("patch", num),

optional("-", named("pre", dotted(preId))),

optional("+", named("build", dotted(oneOrMore(idChar)))),

end,

);

Now you can read the spec in the code. A numeric identifier is 0, or

a non-zero digit followed by any number of digits. A pre-release is a

dotted list of identifiers, and so is build metadata. dotted is a

plain function returning a node, which is as far as abstraction needs

to go here.

It matches the same strings as the official regexp, and the named groups give you a result like:

semver.exec("1.0.0-alpha.1+build.5")?.groups;

// => {

// major: '1',

// minor: '0',

// patch: '0',

// pre: 'alpha.1',

// build: 'build.5'

// }

Next time the spec changes, you can understand what the current regex does at a glance, instead of fighting an army of punctuation.

I tried to map every rx form, and a few have no JavaScript equivalent:

- point: JavaScript regexps don't know about a cursor.

- symbol-start,- symbol-end,- syntax,- category: these depend on Emacs syntax tables.

- intersection: possible with the- vflag, but I haven't needed it.

- minimal-match/- maximal-match: these flip the greediness of everything inside them. Doable, but it would need a separate pass, and the- *Lazyfunctions cover my use cases.

- eval: TypeScript already evaluates expressions everywhere, so you get it for free.

And one addition Emacs doesn't need: RX.flags.

Each example below shows the goal, the Emacs rx form in a comment,

the regexp you would write by hand, and the RX version. When RX

produces a different regexp text, the // => line shows it. The

results at the bottom come from running both against the same strings.

The whole string is digits.

// Emacs: (rx bos (+ digit) eos)

const regex = /^\d+$/;

const dsl = RX(start, oneOrMore(digit), end);

// "123" -> true

// "12a" -> false

The whole string is ASCII letters.

// Emacs: (rx bos (+ alpha) eos)

const regex = /^[a-zA-Z]+$/;

const dsl = RX(start, oneOrMore(alpha), end);

// "Hello" -> true

// "He11o" -> false

Both spellings, color and colour.

// Emacs: (rx bos "colo" (? "u") "r" eos)

const regex = /^colou?r$/;

const dsl = RX(start, "colo", optional("u"), "r", end);

// "color" -> true

// "colour" -> true

// "colouur" -> false

Two words separated by a space.

// Emacs: (rx bos (+ wordchar) space (+ wordchar) eos)

const regex = /^\w+\s\w+$/;

const word = oneOrMore(wordChar);

const dsl = RX(start, word, space, word, end);

// "hello world" -> true

// "hello" -> false

(123) 456-7890, parentheses and all.

// Emacs: (rx bos "(" (= 3 digit) ")" space (= 3 digit) "-" (= 4 digit) eos)

const regex = /^\(\d{3}\)\s\d{3}-\d{4}$/;

const dsl = RX(

start, "(", repeat(3, digit), ")", space,

repeat(3, digit), "-", repeat(4, digit), end,

);

// "(123) 456-7890" -> true

// "123-456-7890" -> false

#ff00aa-style colors.

// Emacs: (rx bos "#" (= 6 hex-digit) eos)

const regex = /^#[0-9a-fA-F]{6}$/;

const dsl = RX(start, "#", repeat(6, hexDigit), end);

// "#ff00aa" -> true

// "#ff00ag" -> false

An optional sign, then digits.

// Emacs: (rx bos (? (any "+-")) (+ digit) eos)

const regex = /^[+-]?\d+$/;

const dsl = RX(start, optional(anyOf("+-")), oneOrMore(digit), end);

// => /^[+\-]?\d+$/

// "-42" -> true

// "42" -> true

// "*42" -> false

Something@something.something, no spaces.

// Emacs: (rx-let ((part (+ (not (any space "@")))))

// (rx bos part "@" part "." part eos))

const regex = /^[^\s@]+@[^\s@]+\.[^\s@]+$/;

const part = oneOrMore(not(anyOf(space, "@")));

const dsl = RX(start, part, "@", part, ".", part, end);

// "a@b.com" -> true

// "a b@c.com" -> false

// "a@b" -> false

A string without any digit.

// Emacs: (rx bos (+ (not digit)) eos)

const regex = /^[^\d]+$/;

const dsl = RX(start, oneOrMore(not(digit)), end);

// => /^\D+$/

// "abc" -> true

// "a1b" -> false

Exactly three comma-separated fields.

// Emacs: (rx-let ((field (+ (not-char ","))))

// (rx bos field "," field "," field eos))

const regex = /^[^,]+,[^,]+,[^,]+$/;

const field = oneOrMore(notChar(","));

const dsl = RX(start, field, ",", field, ",", field, end);

// "a,b,c" -> true

// "a,b," -> false

A fixed list of words.

// Emacs: (rx bos (or "cat" "dog" "bird") eos)

const regex = /^(?:cat|dog|bird)$/;

const dsl = RX(start, or("cat", "dog", "bird"), end);

// => /^(?:bird|cat|dog)$/

// "cat" -> true

// "bird" -> true

// "cow" -> false

mr or ms, then a name, keeping the title.

// Emacs: (rx bos (group (or "mr" "ms")) space (+ wordchar) eos)

const regex = /^(mr|ms)\s\w+$/;

const dsl = RX(start, group(or("mr", "ms")), space, oneOrMore(wordChar), end);

// "mr john" -> true

// "dr john" -> false

Two to four digits.

// Emacs: (rx bos (** 2 4 digit) eos)

const regex = /^\d{2,4}$/;

const dsl = RX(start, between(2, 4, digit), end);

// "12" -> true

// "12345" -> false

Three or more digits, anywhere.

// Emacs: (rx (>= 3 digit))

const regex = /\d{3,}/;

const dsl = RX(atLeast(3, digit));

// "12" -> false

// "a123" -> true

The same word twice.

// Emacs: (rx bos (group (+ wordchar)) space (backref 1) eos)

const regex = /^(\w+)\s\1$/;

const dsl = RX(start, group(oneOrMore(wordChar)), space, backref(1), end);

// "hello hello" -> true

// "hello world" -> false

An open tag and its own closing tag.

// Emacs: (rx "<" (group-n 1 (+ wordchar)) ">" (*? nonl) "</" (backref 1) ">")

const regex = /<(?<tag>\w+)>.*?<\/\k<tag>>/;

const dsl = RX(

"<", named("tag", oneOrMore(wordChar)), ">",

zeroOrMoreLazy(notNewline),

"</", backref("tag"), ">",

);

// "<b>bold</b>" -> true

// "<b>oops</i>" -> false

cat as a word, not inside another one.

// Emacs: (rx word-boundary "cat" word-boundary)

const regex = /\bcat\b/;

const dsl = RX(wordBoundary, "cat", wordBoundary);

// "the cat sat" -> true

// "concatenate" -> false

hello, in any case.

// Emacs: (let ((case-fold-search t))

// (string-match-p (rx bos "hello" eos) "HeLLo"))

const regex = /^hello$/i;

const dsl = RX.flags("i", start, "hello", end);

// "HeLLo" -> true

// "help" -> false

It's a single file with no dependencies. Copy it into your project and start deleting the forms you don't need, or adding the ones you miss.

You can check the same code, plus all the examples from this post (and a few more), in this gist. If you'd rather not set anything up, paste it into the TypeScript Playground, hit "Run", and check the "Logs" tab.

/* =========================================================

* CORE

* ========================================================= */

// How a node behaves when combined with others:

// atom -> single unit, a quantifier can be glued right after it

// seq -> safe to concatenate, needs (?:) to be quantified

// alt -> has a top-level `|`, needs (?:) almost everywhere

type Kind = "atom" | "seq" | "alt";

interface RxNode {

readonly src: string;

readonly kind: Kind;

// char sets only: the text that goes inside [ ], so sets can merge

readonly set?: string;

readonly neg?: boolean;

}

type Item = string | RxNode;

const esc = (s: string) => s.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");

const escSet = (s: string) => s.replace(/[\]\\^-]/g, "\\$&");

// rx: (literal EXPR) — a string computed at runtime, matched as-is

const literal = (s: string): RxNode => ({

src: esc(s),

kind: s.length === 1 ? "atom" : "seq",

});

const toNode = (x: Item): RxNode => (typeof x === "string" ? literal(x) : x);

// wrap a node so a quantifier applies to all of it

const quantifiable = (n: RxNode) =>

n.kind === "atom" ? n.src : `(?:${n.src})`;

// rx: (regexp EXPR) — escape hatch, trust the regexp as-is

const raw = (re: string | RegExp): RxNode => ({

src: typeof re === "string" ? re : re.source,

kind: "alt",

});

/* =========================================================

* COMPOSITION

* ========================================================= */

const seq = (...xs: Item[]): RxNode => {

const nodes = xs.map(toNode).filter((n) => n.src !== "");

if (nodes.length === 0) return { src: "", kind: "seq" };

if (nodes.length === 1) return nodes[0];

let src = "";

for (const n of nodes) {

const part = n.kind === "alt" ? `(?:${n.src})` : n.src;

// `\1` followed by a literal `0` would read as `\10`

if (/\\\d+$/.test(src) && /^\d/.test(part)) src += "(?:)";

src += part;

}

return { src, kind: "seq" };

};

const unmatchable: RxNode = { src: "(?!)", kind: "atom" };

// Like rx: when every branch is a plain string, try the longest first,

// so or("in", "int") matches "int" instead of stopping at "in".

const or = (...xs: Item[]): RxNode => {

if (xs.length === 0) return unmatchable;

if (xs.length === 1) return toNode(xs[0]);

const branches = xs.every((x) => typeof x === "string")

? [...(xs as string[])].sort((a, b) => b.length - a.length)

: xs;

return { src: branches.map((x) => toNode(x).src).join("|"), kind: "alt" };

};

/* =========================================================

* CHARACTER SETS

* ========================================================= */

const set = (body: string, neg = false): RxNode => ({

src: neg ? `[^${body}]` : `[${body}]`,

kind: "atom",

set: body,

neg,

});

// class escapes are sets too, so they can go inside anyOf(...)

const classEscape = (e: string): RxNode => ({

src: e,

kind: "atom",

set: e,

neg: false,

});

// Same reading as rx: inside a string, "a-z" is a range, while a `-`

// at the start or end is just a dash ("+-" is plus or minus).

const intervals = (s: string): string => {

let body = "";

let i = 0;

while (i < s.length) {

if (i < s.length - 2 && s[i + 1] === "-") {

body += `${escSet(s[i])}-${escSet(s[i + 2])}`;

i += 3;

} else {

body += escSet(s[i]);

i += 1;

}

}

return body;

};

// rx: (any "a-z" "_" digit) — also known as `in` and `char`

const anyOf = (...xs: Item[]): RxNode => {

const body = xs

.map((x) => {

if (typeof x === "string") return intervals(x);

if (x.set === undefined || x.neg)

throw new Error(`anyOf: not a positive char set: ${x.src}`);

return x.set;

})

.join("");

return set(body);

};

// rx: (not charset) — not(digit) -> \D, not(anyOf(",;")) -> [^,;]

const not = (x: Item): RxNode => {

const n = typeof x === "string" ? anyOf(x) : x;

if (n.set === undefined) throw new Error(`not: not a char set: ${n.src}`);

if (n.neg) return set(n.set);

if (/^\\[dswDSW]$/.test(n.src)) {

const c = n.src[1];

const flipped = c === c.toLowerCase() ? c.toUpperCase() : c.toLowerCase();

return classEscape(`\\${flipped}`);

}

return set(n.set, true);

};

// rx: (not-char "a-z" ...) — shorthand for (not (any ...))

const notChar = (...xs: Item[]) => not(anyOf(...xs));

// rx char classes, `[[:name:]]` in Emacs

const digit = classEscape("\\d");

const space = classEscape("\\s");

const wordChar = classEscape("\\w");

const notWordChar = not(wordChar);

const lower = anyOf("a-z");

const upper = anyOf("A-Z");

const alpha = anyOf(lower, upper);

const alnum = anyOf(alpha, "0-9");

const hexDigit = anyOf("0-9a-fA-F");

const blank = set(" \\t");

const control = set("\\x00-\\x1f\\x7f");

const punct = anyOf("!-/:-@[-`{-~");

const graphic = anyOf("!-~");

const printing = anyOf(" -~");

const ascii = set("\\x00-\\x7f");

const nonascii = set("\\u0080-\\uffff");

// rx: `nonl` is any char but newline; `anything` really is anything

const notNewline: RxNode = { src: ".", kind: "atom" };

const anything = set("\\s\\S");

/* =========================================================

* ANCHORS (zero-width)

* ========================================================= */

const zeroWidth = (src: string): RxNode => ({ src, kind: "seq" });

// rx: bos / eos

const start = zeroWidth("^");

const end = zeroWidth("$");

// rx: bol / eol — same symbols, only per line with the "m" flag

const lineStart = start;

const lineEnd = end;

const wordBoundary = zeroWidth("\\b");

const notWordBoundary = zeroWidth("\\B");

// rx: bow / eow — JS has no \< \>, so a boundary plus a lookaround

const wordStart = zeroWidth("\\b(?=\\w)");

const wordEnd = zeroWidth("\\b(?<=\\w)");

/* =========================================================

* GROUPS & BACKREFERENCES

* ========================================================= */

const group = (...xs: Item[]): RxNode => ({

src: `(${seq(...xs).src})`,

kind: "atom",

});

// rx has (group-n N ...); JS can't pick group numbers, but it can name them

const named = (name: string, ...xs: Item[]): RxNode => ({

src: `(?<${name}>${seq(...xs).src})`,

kind: "atom",

});

const backref = (ref: number | string): RxNode => ({

src: typeof ref === "number" ? `\\${ref}` : `\\k<${ref}>`,

kind: "atom",

});

/* =========================================================

* QUANTIFIERS

* ========================================================= */

const quantifier =

(suffix: string) =>

(...xs: Item[]): RxNode => ({

src: quantifiable(seq(...xs)) + suffix,

kind: "seq",

});

// greedy — rx: * + ?

const zeroOrMore = quantifier("*");

const oneOrMore = quantifier("+");

const optional = quantifier("?");

// lazy — rx: *? +? ??

const zeroOrMoreLazy = quantifier("*?");

const oneOrMoreLazy = quantifier("+?");

const optionalLazy = quantifier("??");

// rx: (= n ...) (>= n ...) (** n m ...)

const repeat = (n: number, ...xs: Item[]) => quantifier(`{${n}}`)(...xs);

const atLeast = (n: number, ...xs: Item[]) => quantifier(`{${n},}`)(...xs);

const between = (n: number, m: number, ...xs: Item[]) =>

quantifier(`{${n},${m}}`)(...xs);

/* =========================================================

* ENTRY POINTS

* ========================================================= */

// rx(...) -> the regexp source string (like Emacs, rx returns a string)

// RX(...) -> a ready-to-use RegExp, no more `new RegExp(seq(...))`

const rx = (...xs: Item[]): string => seq(...xs).src;

function RX(...xs: Item[]): RegExp {

return new RegExp(rx(...xs));

}

// Emacs uses `case-fold-search` for this; JS puts it on the regexp

RX.flags = (flags: string, ...xs: Item[]): RegExp =>

new RegExp(rx(...xs), flags);

None of this is new. On the Emacs side, as I said before, rx has

shipped for decades, and the Elisp version is more complete than mine.

The idea of describing patterns with a small DSL instead of raw syntax isn't new either. Plenty of people have tried it, each in their own way. One project I like a lot in this space is Zod, which I wrote about in my Zod quick tutorial. It's not a regexp builder: you compose small schema pieces, and Zod gives you back a parser and a TypeScript type from the same construction. It follows the same spirit, though: build big things out of small named pieces you can read.

If you write Elisp and have never tried rx, open *scratch*, type

(rx (+ digit)), and C-x C-e it. If you write JavaScript or

TypeScript, the file above is yours. And if you port it to another

language, send me a link.