On this page
Design
How markz is built: the tree it produces, how offsets map to the source, how the parser reads in
one pass, and what html() writes and what it refuses. How it is tested is in
Quality. Syntax is
the language this reads, and the home page says what markz is for.
AST
The AST is a flat, indexed store, not mdast and not nested objects. A node is a number (its index). Its fields live in parallel typed arrays:
type NodeId = number; // -1 means "none"
// per node, in typed arrays
type_: NodeType; // small integer
start: number; // UTF-16 offset into the source, inclusive
end: number; // exclusive
parent: NodeId;
firstChild: NodeId;
nextSibling: NodeId;
Data specific to one node type (heading depth and id, link destination, code lang, element name, …) lives in a side table indexed by node. Nodes are indices, so parent and sibling links cost nothing, and every node has offsets.
Read-only, and transformations are folds
A Document can't be changed after parsing. Rendering to HTML, or to a framework's template
tree, is a walk that builds a new structure. So markz provides a fast walk with enter/exit
callbacks and typed accessors, and no mutation API. In-place mutation would mean the offsets
could lie. A consumer that wants a rewritten document writes Markdown and parses it again.
Ergonomics
const doc = parse(source);
for (const child of doc.children(doc.root)) console.log(doc.type(child)); // "heading", "paragraph", …
walk(doc, {
enter(node) {
if (doc.type(node) === "link") console.log(doc.data(node, "link").destination);
},
});
doc.data(doc.root, "heading"); // throws a TypeError: the root is not a heading
Iteration follows firstChild/nextSibling and allocates no arrays. Public type names are
strings (NodeType), and the numeric codes stay internal. doc.data(node, type) reads a node's
side-table entry and throws if the node is of another type, so a wrong guess fails loudly. A
value that is decoded (a text node's) or stripped of container prefixes (a code block inside a
blockquote loses its > ) is stored as a string. Everything else is a range into the source.
Node types
There is a node for each construct the syntax has and a consumer uses, and nothing else. A node is a block, an inline, or both for elements and math. Each type, and the data it carries, is in Reference.
Elements, blocks, images and links can carry attributes: an id, classes and key-value pairs, each with its source range. They're kept in a side table, so the common case (no attributes) costs nothing.
There are no html, definition or footnote nodes. HTML exists only in raw blocks, and
reference links and footnotes would need the whole document read before a reference resolves.
Source locations
Offsets are canonical, and they are UTF-16 code units into the exact string passed to parse.
Line and column are computed from them, not stored.
Rules:
- A node's range covers its markers. A heading includes
##, a fence includes both fences, and a link includes[,](…). The trailing line ending is excluded. - Content ranges are exposed as extra fields where consumers need them: a code block's body, a link's destination, the metadata block, and every attribute block.
- Text nodes map to source, not just to their value.
valueis the rendered text: decoded (©→©,\*→*).start/endcover the raw characters. A soft line break is a\nin the text before it, whose range covers the line ending, never the next line's container prefix, so one text node spans lines only where the source has nothing between them. A consumer scanning for syntax of its own readssource.slice(start, end), so a decoded escape can't shift its columns. - Containers with prefixed lines (blockquotes, list items) span from their first
marker to the end of their last content. The
>and indentation prefixes inside that span belong to no child. - Line endings and BOM: CRLF and a lone
\rboth end a line. A leading BOM is part of the source and falls beforedocument.start.
position(source) builds a line-start table once and converts offsets to { line, column } by
binary search. Lines are 1-based and columns are 0-based, as in source-map v3. A consumer that
parses an extracted string (a comment body with its * gutters stripped) maps lines back to the
file itself. markz only promises offsets into what it was given.
Parser foundation
markz has its own parser. It is written for this one dialect, runs in linear time with no backtracking, and emits straight into the flat AST:
source
↓ block pass: lines → containers and leaves
↓ inline pass: each leaf, as it closes
flat AST + warnings
↓
html()
Everything is built in. Elements, expressions, math, attributes, raw blocks, metadata
and heading ids are cases in the same two scanners. They aren't plug-ins
layered on a CommonMark core, because a fixed dialect needs no extension points. That also keeps
precedence in one place: ${…} binding tighter than emphasis is just the order of the inline
scanner's cases.
No backtracking. The language leaves out every construct whose meaning depends on text after it:
| Cut | What forced backtracking |
|---|---|
| setext headings | a paragraph becomes a heading when the next line is read |
| reference links | [x]: url might be a definition or a paragraph, and links resolve only at the end |
| raw HTML | < might open a tag or be text |
| lazy continuation lines | which block a line belongs to depends on context |
| indented code | indentation means code, except inside lists |
| CommonMark emphasis rules | delimiter runs, flanking by punctuation class, the rule of 3 |
| bare-URL autolinks | an email is known only at its @, and trailing punctuation is trimmed afterwards |
What remains is small enough to write by hand, needs no named-entity table, lets html() be safe
outside raw blocks, and never depends on later text, which keeps streaming simple. The rest is
openers ([, ![, _, **, ~~, `, $, ${, and { after a ) or ]) that either
close or turn out to be text.
- Openers go on a stack. The inline pass keeps what it has read as a linked list of items. A closer wraps the items since its opener into one node; an opener still unmatched at the end of its block is text. The input is never read again.
- A scan that fails settles the ones after it. An opener that never closes (
${,[in a label,<!--, a`run,{) scans to the end of its range. That scan records what it learned, such as where each brace it passed closed, or that no closer is left, so a later opener reads the answer instead of scanning again. A paragraph full of unclosed openers stays linear. - Nesting costs nothing per line. A line is checked against the containers that consume a
prefix from it (
>, an item's indent). Elements and blank lines, which consume none, are settled for a whole run of containers at once, so a thousand unclosed{@div}don't make every line cost a thousand. - Block attributes are one line, so the block pass never looks ahead.
The grammar states the dialect, and the parser is its one reading. grammar.md
states each construct in EBNF, with the side rules that settle what the productions leave
ambiguous. The parser isn't generated from it. It is written by hand, to four rules.
- Single pass: the block pass reads each line once, and the inline pass reads each leaf once, as it closes.
- Deterministic: at every point the side rules allow exactly one reading. No alternative is tried and undone.
- Grammar-directed: every case in the two scanners is a construct of the grammar or a Not supported form, and every construct is a case.
- Bounded local lookahead: a scan ahead either stays within the line (a fence, an attribute line, a table's delimiter row) or records where it failed, so no character is scanned more than a constant number of times.
Heading ids are settled in the pass. CommonMark defines headings but not ids, so every
renderer adds them its own way or not at all. markz uses GitHub's algorithm, so a link to a
heading works on GitHub and on the site. A {#id} line sets one by hand. An id is settled as its
heading is parsed, against the ids used so far, so it never depends on a later heading. The steps
are in grammar.md.
micromark is the test oracle, not a dependency. markz must match GFM on the constructs they share, not all of GFM, and micromark checks exactly that (Quality).
Streaming
This is not built, because every use so far parses complete files. A live editor preview works
with plain parse on every change. An unclosed ** shows as text until its closer is typed, as
in every Markdown preview.
The use case it would serve is showing Markdown while it is still arriving, as in an LLM chat UI. The design is kept ready for it:
- One place to heal. Openers wait on a stack until the end of their block (see
Parser foundation).
parseturns unmatched ones into text there. A futureparsePartial(source)would close them instead, along with an open element, and flag those nodespartial. Offsets would stay within the source, with no synthetic text inserted. - Parse the whole prefix again, once per chunk. The language has no reference definitions, so later text never changes an earlier block. That makes reusing finished blocks a safe optimization, if it's ever needed.
- It would be a new export. Adding it later changes nothing for code that calls
parse.
Deciding whether a half-typed opener (hello *) shows or vanishes, and avoiding flicker when
* turns out not to be emphasis, is left for when a consumer needs it.
Public API
The package exports parse, html, walk, textContent, headings and position, the
read-only Document with NONE for "no node", and the types around them. Reference lists each one.
Nothing takes an options object.
headings is the one helper that reads the tree for a consumer's outline, and it is not part of
html(): depth, nesting and numbering are choices, and markz has none.
HTML output
html() is part of the package. It is a fold over the AST, and works the same in Node, Workers
and the browser, because it builds a string and never touches the DOM.
- Heading ids are written as
id. - Attributes are written onto the element they belong to (see Security for the ones that are dropped).
```=htmlraw blocks are written verbatim. Raw blocks in other formats are skipped. Everywhere else,&,<,>and"are escaped, and comments are dropped.- Math, expressions and elements are written in the shapes
syntax.mdgives for each. An expression is<code>, not<output>. markz never evaluates, so it has only the code, and<output>would present that code as a result.<output>is also a form control that screen readers announce. A host that computes the value writes it as it likes.
Framework output is not part of markz. An element's name is the element html() writes, custom
elements included, and a framework's own fold maps names to its components (chart-view to
ChartView). A custom element is inline until CSS says display: block.
Element names
{…} with @name makes an element, and its name is the element written: an HTML element on the
allowlist for its kind, or a valid custom-element name. The allowlists (src/elements.ts) leave
out three kinds of name:
- What Markdown already writes:
em,strong,code,a,img, headings, lists and the rest, so each element has one way in. A class on one of them goes on a span around it, as in[_word_]{.x}. - A second way:
span, since[text]{.x}already writes one. - Anything active:
script,style,iframe,object,embed,svg,math,template, form controls,dialog,audioandvideo.
The parser decides whether a name is an element and warns when it isn't, and the whole {…}
stays text. Writing {@note} as <custom-note> is ruled out. The writer would style note and
get custom-note. A name would also change meaning on the day HTML adds it as an element. html() still refuses active elements as a second line, because a document can be
built without the parser. A leaf's [label] and a span's [text] are content, and a container
has no label: details takes a […]{@summary /} line, so html() needs no rule for where a
label goes. A closing line names what it closes, so a mismatch is a warning, not a silent wrong
nesting.
Security
- Raw blocks are trusted. A
```=htmlblock is written out verbatim, so it can contain script. That's the point for your own content: embeds, SVG, a checkout script. For untrusted input (comments, LLM output), the host should do one of two things:- reject documents that contain
rawnodes. They're easy to find in the AST. - pass the output through the platform Sanitizer (below).
- reject documents that contain
- Everything else is safe by construction. Outside raw blocks, unsafe values can reach the
output only through URLs and attributes, and
html()handles both:- It drops event-handler attributes (
on*). - A URL with an unsafe scheme (
javascript:,vbscript:, anddata:other than a raster image) is written empty, and an attribute value with one is dropped. styleand other attributes pass through, since limited styling is the point of attributes.
- It drops event-handler attributes (
- The AST keeps everything verbatim and makes no safety promise. A consumer building its own output owns its policy.
- markz ships no sanitizer. In the browser, a host that wants defence in depth passes
html()'s output to the platform's HTML Sanitizer API (Element.setHTML()). Chrome 146 and Firefox 148 ship it and Safari doesn't yet, so feature-detect it and fall back to DOMPurify.
Package
The package is @amitkaps/markz (npm refuses the bare markz as too close to marked), and it
belongs to no application. There is one package, with no markz-* companions. It is ESM only, built by vp pack
into dist/index.js and its types, with one entry, src/index.ts: anything it doesn't
re-export is private. It has no runtime dependencies and sideEffects: false, so a consumer
tree-shakes what it doesn't call, and prepack builds, so a tarball never ships a stale dist/.
What ships is readable code without its prose. dist/index.js isn't minified, since a
consumer's bundler minifies it for its own app, and anyone reading node_modules can follow it.
A page with no bundler imports it from jsDelivr's /+esm, which minifies it
(Usage).
Its comments are stripped, but license comments and @__PURE__ annotations stay, and a /*!
banner names the license so it survives a consumer's bundle. There are no sourcemaps.
dist/index.d.ts keeps the comment above each export, @prose included, as that export's
documentation on hover. Nothing strips it, so each export carries a comment written for that,
and a section's @prose sits on the declaration it describes. exports is the only entry, with
no main or types, which nothing on Node 26 reads. vp pack runs publint on the package, and fails the build on a problem.
A release is a merge that changes package.json's version. .github/workflows/release.yml, the
same in every package, runs verify and packs it, stages that tarball on npm with provenance,
then tags the commit vX.Y.Z and attaches the tarball to a GitHub release. npm trusts the workflow to stage, so no npm token is stored anywhere. Nothing
reaches users until a maintainer approves the staged version with 2FA (npm stage approve), so a
merge alone can't publish. A pre-release (vX.Y.Z-rc.N) goes to the next dist-tag and
never becomes latest.
Performance and size
The budget is 20 KB gzip for the whole package, every export bundled and minified. pnpm size
measures it and CI fails above it. One budget covers Node,
Workers and the browser, and how it shaped the design is in
Lessons.
Speed is a property of the design before it is a number: one pass over the source, linear
time, no backtracking, and a flat tree of typed arrays with offsets into the source.
test/complexity.test.ts holds linear time on adversarial patterns and a multi-megabyte
document, and is the only timing CI gates on.
pnpm sizeprints the gzip size against the budget, and what holding the CommonMark spec's tree costs, as a multiple of its source.pnpm speedis markz alone (test/speed.ts), the working tree againstorigin/main, in about 20 seconds: MB/s per document tier and per construct, and the change and its range across three processes.--against <ref>picks another commit.pnpm hotspotsshows where the time goes: self time by area (block pass, inline pass,html()) and by function, warm, over every tier or one tier or construct by name. It finds where to look, andpnpm speedsays whether a change helped.pnpm comparetimes markz beside markdown-exit, marked and micromark on the blocks they all read alike, each in a fresh process. It is for our own insight. The parsers make different trade-offs, so nothing from it is published.- The Quality report measures this commit when it is generated (
pnpm quality): the gzip size against the budget, the memory held by one document's tree, and parse + HTML time on documents a reader can picture, each a range, with the machine named.
Open questions are kept with the ideas for later, in Plan.