Migration and linting

September 25, 2026 · View on GitHub

Tools for moving Markdown or Djot documents to Carve, and for catching constructs that would silently mis-render under Carve's rules.

Migrating from Markdown

markdownToCarve(md) rewrites common Markdown into equivalent Carve. It is a source-to-source transform, not a parser, so it works on raw text and leaves fenced/inline code untouched.

import { markdownToCarve } from '@markup-carve/carve'

markdownToCarve('a *very* **bold** ~~old~~ idea')
// => 'a /very/ *bold* ~old~ idea'

It handles the inline constructs that differ between Markdown and Carve, plus Carve's blank-line-around-blocks rule:

MarkdownCarveNote
*x*, _x_/x/_x_ is underline in Carve, not emphasis
**x**, __x__*x*Carve strong is a single *
***x***, ___x___/*x*/Carve's canonical bold-italic
~~x~~~x~Carve strikethrough is a single ~
==x====x==literal by default - not CommonMark or GFM (dialect.highlight converts it to =x=)
^x^^x^literal by default - not CommonMark or GFM (dialect.superscript converts it to {^x^})
<mark>x</mark>=x=highlight tag → bare marker (brace-forced {=x=} intraword)
<sub>x</sub>,x,subscript tag → bare marker (H<sub>2</sub>O → H{,2,}O intraword)
<sup>x</sup>^x^superscript tag → bare marker (brace-forced {^x^} intraword)
$x$$x$literal by default - not CommonMark or GFM (dialect.math converts it to $`x`, leaving $5 as currency)
<em>/<strong>/<del>/…Carve formother inline HTML tags map to their Carve markers

The default dialect is CommonMark plus GFM, plus footnote references, so a source construct none of those defines stays as written rather than becoming Carve markup. Seven flavor extensions are opt-in through the second argument:

markdownToCarve('a ==hi== ^up^ $x$', { highlight: true, superscript: true, math: true })
// => 'a =hi= {^up^} $`x`'
SourceFlavorFlagOff (the default)On
==x==Obsidian, Quartohighlightliteral=x=
^x^Pandocsuperscriptliteral{^x^}
$x$Pandoc, GitHubmathliteral$`x`
^[body]PandocinlineFootnotes\^[body]an inline footnote
*[HTML]: HyperTextPHP Markdown Extraabbreviations\*[HTML]: HyperTextan abbreviation
::: notePandoc, QuartofencedDivs\::: notea div
[t]{.c}, {.c}Pandoc, kramdownattributes[t]\{.c}, \{.c}attributes

attributes covers an attribute list wherever it would attach, not just the bare span: [t](u){.c}, ![alt](u){.c}, `x`{.c}, <https://e.com/>{.c} and the emphasis family are all escaped with it off. A braced delimiter pair is not an attribute list and is never escaped as one - {,x,} is a subscript in Carve wherever it stands, and it is what <sub>x</sub> converts to.

The last four differ from the first three in how they got here: Carve spells them the way the source does, so nothing had to be rewritten for a CommonMark document to grow markup its author never saw. Escaping is what keeps them literal, which is why the "off" column shows a backslash.

A handful of Carve constructs have no Markdown spelling in any flavor and are always escaped, with no flag: a $`x` and a $$`x` (math spans), a !`x` (a literal span), a :term[x] (an extension call), and a leading ^ on a paragraph (a caption, which binds to the block above it).

One exception to the contract: [^1] footnote references convert by default. They are in neither CommonMark nor GFM, so the rule above would put them behind a flag. The rule exists to stop a migrated document from rendering differently than its author saw it, and here it would cause exactly that: github.com renders footnotes, so an author who wrote one saw one, and leaving it literal would take it away. Where the letter of the contract and the reason for it disagree, the reason governs. [^1] with its [^1]: … definition therefore migrates to a Carve footnote, and this is the only construct the contract makes room for.

The HTML tags below are unaffected: <mark>, <sub> and <sup> mean one thing in every dialect, so they always convert.

On the command line the same transform is carve migrate --from markdown (--from md works too), which reads a named file or stdin and writes Carve to stdout. The dialect flags have no CLI spelling yet, so the command is CommonMark plus GFM only.

carve migrate --from markdown README.md > README.crv
cat post.txt | carve migrate --from bbcode

Note

Carve's highlight and subscript markers are single characters (=x=, ,x,); the doubled forms ==x== and ,,x,, are literal text in Carve (see the corpus pair 79-two-char-delimiter-runs). A bare ,x, / ^x^ / =x= only renders at a word boundary, so the <mark>/<sub>/<sup> tags map to the bare markers when they sit between non-alphanumeric neighbors (the common, whitespace-separated case) and to the forced brace forms {=x=} / {,x,} / {^x^} only when intraword (e.g. H<sub>2</sub>O → H{,2,}O), where the brace form renders in every position (corpus 40-superscript-and-subscript).

It also rewrites GFM tables to Carve's native form: a header row followed by a | --- | delimiter row becomes |=-prefixed header cells, and the delimiter row is dropped (Carve needs no separator). Column alignment from the delimiter (:--, --:, :--:) is glued onto the header marker as |=<, |=>, |=~:

| L | C | R |
| :-- | :--: | --: |
| a | b | c |

becomes

|=< L |=~ C |=> R |
| a | b | c |

Body rows are rewritten in the same spelling carve fmt writes, and each one is fitted to the header's column count the way GFM reads it.

A row whose every cell is blank is dropped. Carve has no spelling for one - | | | is a paragraph, not a row, so writing it would split the table in two - and the importer drops the row rather than inventing a cell the author never typed. migrateMarkdown reports each dropped row as a structure-unspellable diagnostic, so nothing goes missing silently.

A pipe row without a delimiter row stays text. Carve reads any line that begins and ends with | as a table row, with no delimiter row anywhere, so a row GFM shows as a paragraph would otherwise become a table on migration:

| a | b |
| c | d |

GFM renders that as one paragraph, and so does the migrated document - the opening pipe of each line is escaped, which keeps the row literal and keeps it in the paragraph it belongs to:

\| a | b |
\| c | d |

The same applies to every partly-formed table: a delimiter row with no header above it, a header and delimiter whose column counts disagree, and a stray pipe row before or after a real table. A table GFM does read - inside a block quote or a list item as much as at the top level - is untouched.

Note

A --- kept as literal text still renders as an em dash, because Carve applies smart typography to prose. That is true of any --- a migrated document carries, not only one inside a pipe row.

A task checkbox GFM does not read stays text. cmark-gfm's task-list extension takes a box off a line carrying one list marker, bullet or ordered, and off the states [ ], [x] and [X]. Carve spells a task item only behind a bullet, while its parser accepts four more task states. A quoted item, two markers on one line, or one of those extra states would otherwise grow a box the source never spelled. The pair is escaped there:

> - [ ] foo
- - [x] bar
- [-] baz

migrates to

> - \[ ] foo
- - \[x] bar
- \[-] baz

A top-level - [ ] foo still imports a checkbox, and so does a sublist that opens on a line of its own: indentation never puts a list out of the extension's reach, a second marker on the line does. An ordered item takes no escape, Carve spelling a task marker behind a bullet only, so 1. [x] done is already the text it renders.

For an ordered task item, the migration report records the missing checkbox as structure-unspellable beside fidelity-unverified.

Migrating from Djot

djotToCarve(source) converts a complete Djot document. Unlike applyMigrationFixes, it also escapes source text that is inert in Djot but would become Carve markup, including tags, mentions, bare / and =, and the %% comment opener. Code spans, fenced code and link destinations remain opaque.

import { djotToCarve } from '@markup-carve/carve'

djotToCarve('_em_ and ~sub~; literal #tag and /path/')
// => '/em/ and {,sub,}; literal \\#tag and \\/path/'

The CLI spelling reads a file or stdin and writes Carve to stdout:

carve migrate --from djot document.djot > document.crv

To only flag delimiter collisions without importing the document, use djotMigrationWarnings; to rewrite only those collisions, use applyMigrationFixes (or the carve fix CLI below):

import { applyMigrationFixes } from '@markup-carve/carve'

const { output, applied, skipped } = applyMigrationFixes('use _emphasis_ here')
// output  -> 'use /emphasis/ here'
// applied -> the warnings that were spliced in (nested ones compose, so
//            **_x_** fixes to a single-star bold wrapping a slash emphasis)
// skipped -> crossing collisions (e.g. **_x**_) left for manual review

Command line: carve fix

Installing the package provides a carve binary. Its carve fix subcommand wraps applyMigrationFixes to rewrite Djot/Markdown delimiter collisions to their Carve equivalents.

carve fix < in.crv > out.crv     # stdin -> stdout (default)
carve fix --write doc.crv …      # rewrite files in place
carve fix --check doc.crv …      # report only; exit 1 if any would change (CI)
carve fix --stdout doc.crv       # print the fix for one file, don't modify it

With no files it reads stdin and writes the fixed result to stdout. Nested collisions compose (**_x_** fixes in one pass); only crossing collisions that are ambiguous (e.g. **_x**_) are reported on stderr for manual review. --check is a gate: it exits non-zero when a file would change or has manual-review collisions, so it drops into a pre-commit hook or CI step.

Linting

djotMigrationWarnings catches source-level delimiter collisions; lintCarve catches silent-failure problems - markup that parses without error but renders as the wrong thing, so nothing throws. Every rule here is about Carve alone; for "does this document also mean the same thing in Djot" see Portability, which measures the answer rather than linting for it.

import { lintCarve } from '@markup-carve/carve'

lintCarve('# Setup\n\n## Setup\n\nSee </#ghost>.')
// [
//   { rule: 'duplicate-heading-id', line: 3, ... },  // second "Setup" -> id setup-2
//   { rule: 'broken-crossref',      line: 5, ... },  // </#ghost> has no heading
// ]
RuleCatches
duplicate-heading-idtwo headings producing the same id (slug collision or repeated explicit {#id}); ambiguous references resolve to the first
broken-crossrefa </#id> cross-reference with no matching heading or numbered caption id; it renders as literal text
unresolved-reference-linka [text][label] or [text][] reference link with no matching link definition or implicit heading target; it renders as literal text
unresolved-footnotea [^label] footnote reference with no matching [^label]: ... definition; it renders as literal text
duplicate-footnote-definitiona repeated [^label]: ... definition; the parser keeps the first definition and ignores the later one
unused-footnote-definitiona footnote definition that is never referenced; it is omitted from rendered output
heading-trailing-attributea trailing {#id} / {.class} on a heading line; under heading-strict this is literal text, so the attributes never attach (put them on a {…} line above the heading)
raw-block-syntaxa legacy ```raw FORMAT fence; the Carve raw block is ```=FORMAT, and the wrong form fails to open and desyncs the rest of the document's fences
block-marker-as-texta line that opens like a block (:::, {#, {.) but parsed as a paragraph because the block never opened
fence-opener-fallbacka line that starts with a backtick or tilde fence but whose info string is not a language, an optional "title" and an optional [label] (for example ```js {.diff} or ```js title="x"); the fence does not open and the block renders as inline code or plain text. A trailing {…} belongs on its own line above the fence, and the message says so
fence-delimiter-indentationa fenced-code delimiter (``` / ~~~) indented where no list-item authored base applies; at top level and below a container's minimum column the run does not open a code block and instead degrades to inline code/plain text
list-item-block-overindenteda recognized block group authored past a list item's canonical content column; current readers parse it structurally and carve fmt dedents it, while older readers may have treated the marker literally (dedent for explicit structural intent, or escape the opener for literal intent)
list-item-body-detacheda block-shaped line that does not reach the preceding list item's minimum content column and therefore parses outside the item; indent it to the reported column to attach it, or escape the opener to keep literal text
blockquote-marker-without-spacea > blockquote marker with no space after it. Carve requires the separator space, so the marker does not open a quote
empty-include-patha {{ … }} run shaped like an include directive but with no path (empty braces, or only a #section / @option); an empty path is not a directive, so it renders as literal text - add a path or remove the braces
footnotes-placement-in-containera ::: footnotes marker inside a block-level container - a block quote, a list item, a div or directive body, a definition description, a footnote definition. Only a top-level marker places (PART 9 §16, CARVE-P9-073), so the marker renders the <div class="footnotes"> floor where it is written and the endnotes section goes where it would without the marker. Move the marker to document level, or delete it to accept the default position
references-placement-in-containera ::: references marker inside a container while the citations extension is enabled. It renders as a div there, while the reference list keeps its document position. Move the marker to document level to place the list

Pass { extensions: [citations()] } to lintCarve when rendering with citations, or run carve lint --extension citations doc.crv. The references rule is off for a core render.

The carve lint CLI reports collision warnings and lint findings as file:line:col rule - message, and exits non-zero if anything is found:

carve lint doc.crv …   # report; exit 1 if any finding (CI / pre-commit)
carve lint < doc.crv   # read stdin

Platform rules (opt-in, default OFF)

Two further rules answer a different question: not "is this document right in Carve" but "does a HOST mangle it after publication". No render-time construct prevents a host from re-linkifying published output, so a bare #123 becomes a link to an unrelated issue and a bare at-word becomes a mention that notifies an uninvolved person. The source is the only place the author's intent still exists.

They are off by default and enabled per platform, because unlike every rule above they are target-specific - an over-eager rule people disable wholesale would be worse than none.

lintCarve(src)                            // never emits a platform rule
lintCarve(src, { platforms: ['github'] }) // opts in
carve lint --platform github doc.crv   # repeatable; an unknown name is an error
RuleCatches
platform-mention-tokenan at-prefixed word (@minutely, @param, @property, @types/node) outside a fenced block; the host turns it into a mention that notifies whoever owns that handle
platform-issue-referencea hash-number (#1, #123) outside a fenced block; the host turns it into a link to an unrelated issue, and posts a backlink there

Two ids rather than one, because the two token shapes have different false-positive profiles and an author will want to silence one without the other.

They look in prose and in inline code spans - those are not reliably safe, since some host surfaces (a pull-request list, a commit log view) still linkify inside them. They do not look in fenced code blocks, which are reliably safe, nor in raw blocks or comments, nor in text that is never published: frontmatter, link and abbreviation definitions, an unreferenced footnote definition, and an inline link's destination. A token inside a URL is part of that URL, so nothing in a bare URL's path, query or fragment is flagged either. A captioned listing's caption and a referenced footnote's body are published, so both are checked. The suggested fix in each message is to move the example into a fenced block, strip the sigil and rephrase, or rewrite an enumerated reference as "item 1" / "point 1".

Portability

Linting answers "is this document right in Carve". A different question comes up when a file has to survive both readers - a README rendered by Carve here and by a Djot processor somewhere else: does it mean the same thing in Djot?

That one is not a lint. It was tried as one (carve-js#546): a rule reasoned about when a block opener would be absorbed into a paragraph by Djot but not by Carve. The divergence it described is real, but the rule tested a property of the Carve tree while the divergence is a statement about Djot's block model, and the two came apart on documents where Djot absorbs the paragraph into something before it. Measured false positives ran from 11.5% to 36.5% depending on the generator, and its advice - "add a blank line" - changed the Carve document in the cases it got wrong.

So carve portability does not reason about it. It renders the document with both engines and reports the first place they disagree:

carve portability doc.crv     # exit 0 portable, 1 diverges
carve portability --json *.crv
doc.crv:1: diverges from Djot
  carve: </p><blockquote><p>A quote.</p></blockquote>
  djot:   &gt; A quote.</p>

It needs djot.js, which Carve does not depend on - install it alongside:

npm install @djot/djot

Two things to expect from the output:

  • Carve's deliberate departures are divergences. /italic/, =mark= and a quoted link title mean something else in Djot, so a document using them is reported. That is the correct answer to the question being asked, not noise - but it does mean a Carve-flavored document is rarely portable, and the check is most useful on prose you intend to keep neutral.
  • Only the first divergence is reported. Once the engines disagree about a block boundary everything after it is displaced, so the rest of the report would restate one difference as many.

Differences in how the two renderers write the same document are not divergences: attribute order, a boolean attribute spelled disabled or disabled="", a self-closing slash, and whitespace at a block boundary are all normalized away first. Whitespace between inline siblings and inside <pre> is content and is compared as-is.

Programmatically the engine is injected, so importing @markup-carve/carve never pulls in a Djot parser:

import { checkPortability, carveToHtml } from '@markup-carve/carve'
import { parse, renderHTML } from '@djot/djot'

const report = checkPortability(
  source,
  { parse, renderHTML },
  (src) => carveToHtml(src, { sourceLine: true }),
)
// { portable: false, divergence: { line: 1, carve: '…', djot: '…' } }

The sourceLine render option is what lets the report name a line: the check reads Carve's own data-source-line output and drops it before comparing, so the line comes from the parser rather than from a guess about the source.