Corpus
Real documents, and the variants built from them, which documents.test.ts holds markz to and
pnpm speed times. The variants are built on each run and never committed. Each tier of
documents answers its own question:
- markz: this repository's own writing (
README.md,AGENTS.mdanddocs/), read where it lives. It is the only real writing in the dialect, with its metadata, elements and math, and it changes with markz. - agents: the instructions open-source projects keep for coding agents (
AGENTS.md,CLAUDE.md). It is the Markdown agents read and write most, and none of it is markz's. Its warnings show how much of that writing the dialect cuts. - public: documentation written by people, in other styles: Node.js's API reference (dense
links and code), the Rust book (narrative with listings) and Vite's guide (VitePress, with
:::containers), so no one voice speaks for Markdown in general. - spec: the CommonMark spec's own text, which other parsers benchmark on. It is dense with edge cases, not a typical document.
A document comes in three variants. dialect is the document as written. common keeps only the top-level blocks every parser reads alike. formatted is the document after oxfmt, which is how the consumers store it. The scaling tier repeats the markz and public documents to a size. It leaves out the agents tier, whose files read as checklists more than pages.
The agents, public and spec documents are vendored in test/documents/, pinned and never edited. The
markz tier is never vendored. Frozen copies of markz's docs and of the repos that use it (base,
prose, visdown) were dropped: they duplicated the live docs and snapshotted repos that keep
changing.
Nothing here parses at run time: a variant that needs markz's reading is given the parsed
document, so the tests and the Quality page both pass src/.
import { execFileSync } from "node:child_process";
import { mkdtempSync, readdirSync, readFileSync, rmSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import type { Document, NodeId } from "../../src/index";
export type Tier = "markz" | "agents" | "public" | "spec";
export const TIERS: Tier[] = ["markz", "agents", "public", "spec"];
const root = join(import.meta.dirname, "../..");
const DOCUMENTS = join(root, "test/documents");
/** A tier's documents, named `<source>-<file>`, in a stable order. */
export function documents(tier: Tier): Map<string, string> {
if (tier === "markz") return own();
const dir = join(DOCUMENTS, tier);
const out = new Map<string, string>();
for (const entry of readdirSync(dir, { withFileTypes: true })) {
if (entry.isFile() && entry.name.endsWith(".md")) {
out.set(entry.name, readFileSync(join(dir, entry.name), "utf8"));
} else if (entry.isDirectory()) {
for (const file of readdirSync(join(dir, entry.name))) {
if (file.endsWith(".md")) {
out.set(`${entry.name}-${file}`, readFileSync(join(dir, entry.name, file), "utf8"));
}
}
}
}
return new Map([...out].sort(([a], [b]) => a.localeCompare(b)));
}
/** The markz tier: the root's documents, and `docs/` as `docs-<file>`. */
function own(): Map<string, string> {
const out = new Map<string, string>();
for (const name of ["README.md", "AGENTS.md"]) {
out.set(name, readFileSync(join(root, name), "utf8"));
}
for (const file of readdirSync(join(root, "docs"))) {
if (file.endsWith(".md"))
out.set(`docs-${file}`, readFileSync(join(root, "docs", file), "utf8"));
}
return new Map([...out].sort(([a], [b]) => a.localeCompare(b)));
}#The common variant
A top-level block is kept when nothing in it is markz's alone (metadata, comments, elements, math, raw blocks, expressions, attributes, an explicit heading id, a task item) and no warning touches it, since a warning marks a form markz cuts and the others read. The kept blocks are joined by blank lines, so each is still read as the block it was.
const OWN = new Set(["metadata", "comment", "element", "math", "raw", "expression"]);
export function common(doc: Document): string {
return commonBlocks(doc).join("\n\n") + "\n";
}
/** The blocks the common variant keeps, each on its own. */
export function commonBlocks(doc: Document): string[] {
const kept: string[] = [];
for (const block of doc.children(doc.root)) {
const start = doc.start(block);
const end = doc.end(block);
const warned = doc.warnings.some((w) => w.start < end && w.end > start);
if (!warned && shared(doc, block)) kept.push(doc.source.slice(start, end));
}
return kept;
}
function shared(doc: Document, node: NodeId): boolean {
const type = doc.type(node);
if (
OWN.has(type) ||
doc.attributes(node) ||
(type === "heading" && doc.data(node, "heading").idExplicit) ||
(type === "listItem" && doc.data(node, "listItem").checked !== null)
) {
return false;
}
return [...doc.children(node)].every((child) => shared(doc, child));
}
/** The documents after oxfmt with the repo's settings, formatted in a scratch copy. */
export function format(docs: Map<string, string>): Map<string, string> {
const dir = mkdtempSync(join(tmpdir(), "markz-corpus-"));
try {
for (const [name, text] of docs) writeFileSync(join(dir, name), text);
execFileSync(join(root, "node_modules/.bin/vp"), ["fmt", dir], { cwd: root, stdio: "ignore" });
return new Map([...docs.keys()].map((name) => [name, readFileSync(join(dir, name), "utf8")]));
} finally {
rmSync(dir, { recursive: true, force: true });
}
}#Repeating without repeating titles
A document repeated to a size is a benchmark's workload, and a real document doesn't hold the
same heading hundreds of times. Every copy after the first puts its number at the start of each
heading's text (## Setup becomes ## 2 Setup), so ids don't pile up as setup-1 to
setup-300 and time id numbering no author would ask for. The number goes straight after the
#s, so each heading keeps its level, its closing #s and its reading.
const ATX = /^( {0,3}#{1,6})(?=[ \t]|$)/gm;
/** `text` with `label()` put at the start of each ATX heading's text. */
export function titled(text: string, label: () => string): string {
return text.replace(ATX, (marker) => `${marker} ${label()}`);
}
/** `text` repeated to `bytes` characters, cut at the end of a line, each copy's titles its own. */
export function repeat(text: string, bytes: number): string {
let out = "";
for (let copy = 1; out.length < bytes; copy++) {
out += (copy === 1 ? text : titled(text, () => String(copy))) + "\n\n";
}
const cut = out.lastIndexOf("\n", bytes);
return out.slice(0, cut > 0 ? cut + 1 : bytes);
}
/** The scaling tier's text: the markz and public documents, one after another. */
export function mix(): string {
return [...documents("markz").values(), ...documents("public").values()].join("\n\n");
}