markz
On this page

Design

How markz is built: the tree it produces, how offsets map to the source, how the parser reads in one pass, and what html() writes and what it refuses. How it is tested is in Quality. Syntax is the language this reads, and the home page says what markz is for.

AST

The AST is a flat, indexed store, not mdast and not nested objects. A node is a number (its index). Its fields live in parallel typed arrays:

type NodeId = number; // -1 means "none"

// per node, in typed arrays
type_: NodeType; // small integer
start: number; // UTF-16 offset into the source, inclusive
end: number; // exclusive
parent: NodeId;
firstChild: NodeId;
nextSibling: NodeId;

Data specific to one node type (heading depth and id, link destination, code lang, element name, …) lives in a side table indexed by node. Nodes are indices, so parent and sibling links cost nothing, and every node has offsets.

Read-only, and transformations are folds

A Document can't be changed after parsing. Rendering to HTML, or to a framework's template tree, is a walk that builds a new structure. So markz provides a fast walk with enter/exit callbacks and typed accessors, and no mutation API. In-place mutation would mean the offsets could lie. A consumer that wants a rewritten document writes Markdown and parses it again.

Ergonomics

const doc = parse(source);

for (const child of doc.children(doc.root)) console.log(doc.type(child)); // "heading", "paragraph", …

walk(doc, {
  enter(node) {
    if (doc.type(node) === "link") console.log(doc.data(node, "link").destination);
  },
});

doc.data(doc.root, "heading"); // throws a TypeError: the root is not a heading

Iteration follows firstChild/nextSibling and allocates no arrays. Public type names are strings (NodeType), and the numeric codes stay internal. doc.data(node, type) reads a node's side-table entry and throws if the node is of another type, so a wrong guess fails loudly. A value that is decoded (a text node's) or stripped of container prefixes (a code block inside a blockquote loses its > ) is stored as a string. Everything else is a range into the source.

Node types

There is a node for each construct the syntax has and a consumer uses, and nothing else. A node is a block, an inline, or both for elements and math. Each type, and the data it carries, is in Reference.

Elements, blocks, images and links can carry attributes: an id, classes and key-value pairs, each with its source range. They're kept in a side table, so the common case (no attributes) costs nothing.

There are no html, definition or footnote nodes. HTML exists only in raw blocks, and reference links and footnotes would need the whole document read before a reference resolves.

Source locations

Offsets are canonical, and they are UTF-16 code units into the exact string passed to parse. Line and column are computed from them, not stored.

Rules:

  • A node's range covers its markers. A heading includes ##, a fence includes both fences, and a link includes [, ](…). The trailing line ending is excluded.
  • Content ranges are exposed as extra fields where consumers need them: a code block's body, a link's destination, the metadata block, and every attribute block.
  • Text nodes map to source, not just to their value. value is the rendered text: decoded (© → ©, \* → *). start/end cover the raw characters. A soft line break is a \n in the text before it, whose range covers the line ending, never the next line's container prefix, so one text node spans lines only where the source has nothing between them. A consumer scanning for syntax of its own reads source.slice(start, end), so a decoded escape can't shift its columns.
  • Containers with prefixed lines (blockquotes, list items) span from their first marker to the end of their last content. The > and indentation prefixes inside that span belong to no child.
  • Line endings and BOM: CRLF and a lone \r both end a line. A leading BOM is part of the source and falls before document.start.

position(source) builds a line-start table once and converts offsets to { line, column } by binary search. Lines are 1-based and columns are 0-based, as in source-map v3. A consumer that parses an extracted string (a comment body with its * gutters stripped) maps lines back to the file itself. markz only promises offsets into what it was given.

Parser foundation

markz has its own parser. It is written for this one dialect, runs in linear time with no backtracking, and emits straight into the flat AST:

source
  ↓  block pass: lines → containers and leaves
  ↓  inline pass: each leaf, as it closes
flat AST + warnings
  ↓
html()

Everything is built in. Elements, expressions, math, attributes, raw blocks, metadata and heading ids are cases in the same two scanners. They aren't plug-ins layered on a CommonMark core, because a fixed dialect needs no extension points. That also keeps precedence in one place: ${…} binding tighter than emphasis is just the order of the inline scanner's cases.

No backtracking. The language leaves out every construct whose meaning depends on text after it:

Cut What forced backtracking
setext headings a paragraph becomes a heading when the next line is read
reference links [x]: url might be a definition or a paragraph, and links resolve only at the end
raw HTML < might open a tag or be text
lazy continuation lines which block a line belongs to depends on context
indented code indentation means code, except inside lists
CommonMark emphasis rules delimiter runs, flanking by punctuation class, the rule of 3
bare-URL autolinks an email is known only at its @, and trailing punctuation is trimmed afterwards

What remains is small enough to write by hand, needs no named-entity table, lets html() be safe outside raw blocks, and never depends on later text, which keeps streaming simple. The rest is openers ([, ![, _, **, ~~, `, $, ${, and { after a ) or ]) that either close or turn out to be text.

  • Openers go on a stack. The inline pass keeps what it has read as a linked list of items. A closer wraps the items since its opener into one node; an opener still unmatched at the end of its block is text. The input is never read again.
  • A scan that fails settles the ones after it. An opener that never closes (${, [ in a label, <!--, a ` run, {) scans to the end of its range. That scan records what it learned, such as where each brace it passed closed, or that no closer is left, so a later opener reads the answer instead of scanning again. A paragraph full of unclosed openers stays linear.
  • Nesting costs nothing per line. A line is checked against the containers that consume a prefix from it (>, an item's indent). Elements and blank lines, which consume none, are settled for a whole run of containers at once, so a thousand unclosed {@div} don't make every line cost a thousand.
  • Block attributes are one line, so the block pass never looks ahead.

The grammar states the dialect, and the parser is its one reading. grammar.md states each construct in EBNF, with the side rules that settle what the productions leave ambiguous. The parser isn't generated from it. It is written by hand, to four rules.

  • Single pass: the block pass reads each line once, and the inline pass reads each leaf once, as it closes.
  • Deterministic: at every point the side rules allow exactly one reading. No alternative is tried and undone.
  • Grammar-directed: every case in the two scanners is a construct of the grammar or a Not supported form, and every construct is a case.
  • Bounded local lookahead: a scan ahead either stays within the line (a fence, an attribute line, a table's delimiter row) or records where it failed, so no character is scanned more than a constant number of times.

Heading ids are settled in the pass. CommonMark defines headings but not ids, so every renderer adds them its own way or not at all. markz uses GitHub's algorithm, so a link to a heading works on GitHub and on the site. A {#id} line sets one by hand. An id is settled as its heading is parsed, against the ids used so far, so it never depends on a later heading. The steps are in grammar.md.

micromark is the test oracle, not a dependency. markz must match GFM on the constructs they share, not all of GFM, and micromark checks exactly that (Quality).

Streaming

This is not built, because every use so far parses complete files. A live editor preview works with plain parse on every change. An unclosed ** shows as text until its closer is typed, as in every Markdown preview.

The use case it would serve is showing Markdown while it is still arriving, as in an LLM chat UI. The design is kept ready for it:

  • One place to heal. Openers wait on a stack until the end of their block (see Parser foundation). parse turns unmatched ones into text there. A future parsePartial(source) would close them instead, along with an open element, and flag those nodes partial. Offsets would stay within the source, with no synthetic text inserted.
  • Parse the whole prefix again, once per chunk. The language has no reference definitions, so later text never changes an earlier block. That makes reusing finished blocks a safe optimization, if it's ever needed.
  • It would be a new export. Adding it later changes nothing for code that calls parse.

Deciding whether a half-typed opener (hello *) shows or vanishes, and avoiding flicker when * turns out not to be emphasis, is left for when a consumer needs it.

Public API

The package exports parse, html, walk, textContent, headings and position, the read-only Document with NONE for "no node", and the types around them. Reference lists each one. Nothing takes an options object.

headings is the one helper that reads the tree for a consumer's outline, and it is not part of html(): depth, nesting and numbering are choices, and markz has none.

HTML output

html() is part of the package. It is a fold over the AST, and works the same in Node, Workers and the browser, because it builds a string and never touches the DOM.

  • Heading ids are written as id.
  • Attributes are written onto the element they belong to (see Security for the ones that are dropped).
  • ```=html raw blocks are written verbatim. Raw blocks in other formats are skipped. Everywhere else, &, <, > and " are escaped, and comments are dropped.
  • Math, expressions and elements are written in the shapes syntax.md gives for each. An expression is <code>, not <output>. markz never evaluates, so it has only the code, and <output> would present that code as a result. <output> is also a form control that screen readers announce. A host that computes the value writes it as it likes.

Framework output is not part of markz. An element's name is the element html() writes, custom elements included, and a framework's own fold maps names to its components (chart-view to ChartView). A custom element is inline until CSS says display: block.

Element names

{…} with @name makes an element, and its name is the element written: an HTML element on the allowlist for its kind, or a valid custom-element name. The allowlists (src/elements.ts) leave out three kinds of name:

  • What Markdown already writes: em, strong, code, a, img, headings, lists and the rest, so each element has one way in. A class on one of them goes on a span around it, as in [_word_]{.x}.
  • A second way: span, since [text]{.x} already writes one.
  • Anything active: script, style, iframe, object, embed, svg, math, template, form controls, dialog, audio and video.

The parser decides whether a name is an element and warns when it isn't, and the whole {…} stays text. Writing {@note} as <custom-note> is ruled out. The writer would style note and get custom-note. A name would also change meaning on the day HTML adds it as an element. html() still refuses active elements as a second line, because a document can be built without the parser. A leaf's [label] and a span's [text] are content, and a container has no label: details takes a […]{@summary /} line, so html() needs no rule for where a label goes. A closing line names what it closes, so a mismatch is a warning, not a silent wrong nesting.

Security

  • Raw blocks are trusted. A ```=html block is written out verbatim, so it can contain script. That's the point for your own content: embeds, SVG, a checkout script. For untrusted input (comments, LLM output), the host should do one of two things:
    • reject documents that contain raw nodes. They're easy to find in the AST.
    • pass the output through the platform Sanitizer (below).
  • Everything else is safe by construction. Outside raw blocks, unsafe values can reach the output only through URLs and attributes, and html() handles both:
    • It drops event-handler attributes (on*).
    • A URL with an unsafe scheme (javascript:, vbscript:, and data: other than a raster image) is written empty, and an attribute value with one is dropped.
    • style and other attributes pass through, since limited styling is the point of attributes.
  • The AST keeps everything verbatim and makes no safety promise. A consumer building its own output owns its policy.
  • markz ships no sanitizer. In the browser, a host that wants defence in depth passes html()'s output to the platform's HTML Sanitizer API (Element.setHTML()). Chrome 146 and Firefox 148 ship it and Safari doesn't yet, so feature-detect it and fall back to DOMPurify.

Package

The package is @amitkaps/markz (npm refuses the bare markz as too close to marked), and it belongs to no application. There is one package, with no markz-* companions. It is ESM only, built by vp pack into dist/index.js and its types, with one entry, src/index.ts: anything it doesn't re-export is private. It has no runtime dependencies and sideEffects: false, so a consumer tree-shakes what it doesn't call, and prepack builds, so a tarball never ships a stale dist/.

What ships is readable code without its prose. dist/index.js isn't minified, since a consumer's bundler minifies it for its own app, and anyone reading node_modules can follow it. A page with no bundler imports it from jsDelivr's /+esm, which minifies it (Usage). Its comments are stripped, but license comments and @__PURE__ annotations stay, and a /*! banner names the license so it survives a consumer's bundle. There are no sourcemaps. dist/index.d.ts keeps the comment above each export, @prose included, as that export's documentation on hover. Nothing strips it, so each export carries a comment written for that, and a section's @prose sits on the declaration it describes. exports is the only entry, with no main or types, which nothing on Node 26 reads. vp pack runs publint on the package, and fails the build on a problem.

A release is a merge that changes package.json's version. .github/workflows/release.yml, the same in every package, runs verify and packs it, stages that tarball on npm with provenance, then tags the commit vX.Y.Z and attaches the tarball to a GitHub release. npm trusts the workflow to stage, so no npm token is stored anywhere. Nothing reaches users until a maintainer approves the staged version with 2FA (npm stage approve), so a merge alone can't publish. A pre-release (vX.Y.Z-rc.N) goes to the next dist-tag and never becomes latest.

Performance and size

The budget is 20 KB gzip for the whole package, every export bundled and minified. pnpm size measures it and CI fails above it. One budget covers Node, Workers and the browser, and how it shaped the design is in Lessons.

Speed is a property of the design before it is a number: one pass over the source, linear time, no backtracking, and a flat tree of typed arrays with offsets into the source. test/complexity.test.ts holds linear time on adversarial patterns and a multi-megabyte document, and is the only timing CI gates on.

  • pnpm size prints the gzip size against the budget, and what holding the CommonMark spec's tree costs, as a multiple of its source.
  • pnpm speed is markz alone (test/speed.ts), the working tree against origin/main, in about 20 seconds: MB/s per document tier and per construct, and the change and its range across three processes. --against <ref> picks another commit.
  • pnpm hotspots shows where the time goes: self time by area (block pass, inline pass, html()) and by function, warm, over every tier or one tier or construct by name. It finds where to look, and pnpm speed says whether a change helped.
  • pnpm compare times markz beside markdown-exit, marked and micromark on the blocks they all read alike, each in a fresh process. It is for our own insight. The parsers make different trade-offs, so nothing from it is published.
  • The Quality report measures this commit when it is generated (pnpm quality): the gzip size against the budget, the memory held by one document's tree, and parse + HTML time on documents a reader can picture, each a range, with the machine named.

Open questions are kept with the ideas for later, in Plan.