blogs/docs/PARSER.md
Neeldhara Misra d598004d79
Some checks failed
Build / build (push) Has been cancelled
Sort data table resource numbers
2026-06-29 03:41:39 +05:30

22 KiB

Semantic Markdown Parser

This document describes the reusable parser contract for the blog-family Markdown pipeline. It is intentionally project-independent: the same ideas should be portable to future site families such as neeldhara.courses and books.neeldhara.com.

The design goal is a single Markdown source that renders predictably to:

  • Astro HTML, with React interactives where needed.
  • PDF, through Pandoc and LuaLaTeX.
  • Obsidian preview, with syntax that remains legible even when Obsidian does not understand a custom semantic feature.

Principles

The source format is mostly ordinary Markdown. Custom syntax is introduced only when the output needs semantic information that plain Markdown cannot carry.

Use these rules when adding new syntax:

  • Prefer fenced attributes for semantic blocks. Pandoc understands them, and they are easy to parse in Astro.
  • Keep labels explicit. A label such as thm:hall should map to a LaTeX \label{thm:hall} and an HTML id="thm:hall".
  • Use normal Markdown links for references. [Hall's theorem](#thm:hall) becomes an HTML anchor in Astro and \cref{thm:hall} in PDF.
  • Make HTML and PDF fallbacks explicit for interactive material.
  • Keep Obsidian-readable text in the source. A reader opening the vault should understand the note even without the final renderers.

Current Implementation

The parser is split across an HTML path and a PDF path.

HTML path:

Astro Markdown/MDX
  -> remark-gfm
  -> remarkInlineFootnotes
  -> remarkCodeFenceLanguages
  -> remarkSemanticBlocks
  -> remark-math
  -> remarkTypography
  -> remarkCallouts
  -> remarkSymbols
  -> rehype-katex
  -> Astro pages

PDF path:

Markdown
  -> Pandoc markdown reader
  -> scripts/pandoc-filters/semantic-blocks.lua
  -> LuaLaTeX
  -> scripts/pandoc-templates/semantic-preamble.tex

The HTML parser lives per site at:

sites/<site>/src/remark/

The current files are duplicated across the seven blog sites. When reusing this system in another project, prefer extracting these files into a small shared package or copying them together as a unit.

The PDF parser lives at:

scripts/pandoc-filters/semantic-blocks.lua
scripts/pandoc-templates/semantic-preamble.tex
scripts/export-pdf.mjs

Dependencies

Astro-side dependencies:

@astrojs/mdx
@astrojs/react
rehype-katex
remark-gfm
remark-math
mdast-util-from-markdown

PDF-side dependencies:

pandoc
lualatex
amsmath
amssymb
amsthm
cleveref
tcolorbox
emoji
chessfss
xcolor

For TeX Live installations, make sure the emoji, chessfss, tcolorbox, and cleveref packages are installed.

Base Markdown

The expected source is normal Markdown with YAML frontmatter:

---
title: "Example"
description: "A short summary."
pubDate: "2026-01-31"
---

Body text.

Supported common Markdown features:

  • Headings.
  • Lists.
  • Tables via remark-gfm in HTML and Pandoc's Markdown reader in PDF.
  • Footnotes.
  • Inline math with $...$.
  • Display math with $$...$$.
  • Code fences.
  • Raw HTML for web-only fragments.

Use colocated relative assets:

![A small diagram](diagram.png)

Avoid raw LaTeX in source unless the post is intended only for PDF.

Typography

Write three hyphens for an em dash:

This is one thought --- and this is another.

HTML converts text-node --- into through remarkTypography.

PDF leaves --- in the generated LaTeX, where LuaLaTeX typesets it as an em dash. This means the source stays ASCII-friendly while both outputs display the proper dash.

Callouts

Callouts are portable blockquotes. The callout type controls styling but does not create a visible title by itself.

Supported types:

note
tip
warning
caution
aside

Canonical titleless form:

> **Note**
>
> This is a note without a visible heading.

Obsidian-style titleless form:

> [!tip]
> This is a titleless tip.

Explicit title form:

> [!warning: Check this assumption]
>
> This warning has a visible title.

The pipe spelling is equivalent:

> [!warning|Check this assumption]
>
> This warning has a visible title.

Do not use this form for a title:

> [!tip] Useful observation

That is intentionally interpreted as a titleless callout whose body begins with "Useful observation". This rule keeps parsing unambiguous across Obsidian, Astro, and Pandoc.

Quick marker forms are supported:

> 📝 This becomes a note.
> 💡 This becomes a tip.
> ⚡ This becomes a tip.
> ⚠️ This becomes a warning.

HTML output:

<aside class="callout callout-warning">
  <p class="callout-title">Check this assumption</p>
  <p>...</p>
</aside>

If no explicit title is present, no .callout-title paragraph is emitted.

PDF output:

\begin{blogwarningbox}{Check this assumption}
...
\end{blogwarningbox}

If no explicit title is present, the second environment argument is empty and the tcolorbox title is omitted.

Semantic Environments

Use semantic environments for theorem-like content, definitions, examples, and proofs.

Canonical form:

```{.env .theorem #thm:hall title="Hall's theorem" html="callout"}
Every bipartite graph satisfying Hall's condition has a matching that covers
the left side.
```

Supported environment types:

theorem
lemma
proposition
corollary
definition
example
remark
proof

Attributes:

Attribute Required Meaning
.env Recommended Marks this as a semantic environment.
.<type> Yes One of the supported environment types.
#label Recommended Cross-reference label and HTML id.
title="..." Optional Environment title.
html="callout" Optional Render as a styled block in HTML.
html="plain" Optional Render as a quieter left-rule block in HTML.

Default HTML mode:

  • Theorem-like blocks default to html="callout".
  • proof defaults to html="plain".

HTML output:

<section
  id="thm:hall"
  class="semantic-env semantic-env-theorem semantic-env-callout callout"
  data-env-type="theorem"
  data-env-title="Hall's theorem"
  data-env-render="callout"
>
  <p class="semantic-env-heading"><strong>Theorem (Hall's theorem)</strong></p>
  ...
</section>

PDF output:

\begin{theorem}[Hall's theorem]
\label{thm:hall}
...
\end{theorem}

Short HTML-only form:

```theorem id="thm:hall" title="Hall's theorem"
Every bipartite graph satisfying Hall's condition has a matching...
```

Use the canonical Pandoc-style form for any material that should export to PDF.

Data Tables

Use data tables when a post should keep rows in a colocated YAML or JSON file but render as a searchable, sortable table in Astro.

Canonical form:

```{.data-table #tbl:resources src="resources.yml" columns="type,title,summary" headers="type:Type,title:Resource,summary:Notes" link="title" number="slno" caption="Resources" search-placeholder="Search resources"}
The HTML version renders this as a searchable, sortable table. In Obsidian,
edit `resources.yml` to update the rows.
```

Attributes:

Attribute Required Meaning
.data-table Yes Marks this as a data-backed table.
#tbl:<id> Optional HTML id and future cross-reference label.
src="resources.yml" Yes YAML or JSON file path, resolved relative to the Markdown.
columns="type,title,summary" Optional Visible columns and their order.
headers="key:Label,..." Optional Human-readable labels for selected columns.
link="title" Optional Column whose value links through the row's URL field.
url="link" Optional URL field used by the linked column; defaults to link.
number="slno" Optional Prefix and sort the linked column by a row number field.
caption="Resources" Optional Visible toolbar title and accessible table caption.
search="false" Optional Disable local table search; enabled by default.
sort="false" Optional Disable sortable headers; enabled by default.
search-placeholder="Search" Optional Placeholder text for the search input.
pdf="Fallback text" Optional PDF-specific fallback text.
fallback="Fallback text" Optional General fallback text when pdf is absent.

The data file should contain either a top-level array:

- slno: 01
  title: "Cognitive Productivity"
  summary: "Book by Luc P. Beaudoin."
  link: https://leanpub.com/cognitiveproductivity/
  type: book

or an object with an items array:

items:
  - title: "Cognitive Productivity"
    summary: "Book by Luc P. Beaudoin."
    link: https://leanpub.com/cognitiveproductivity/
    type: book

HTML output:

<section class="semantic-data-table" id="tbl:resources" data-data-table-root>
  <div class="data-table-toolbar">
    <h2>Resources</h2>
    <div class="data-table-search">...</div>
  </div>
  <div class="data-table-scroll">
    <table>
      ...
    </table>
  </div>
</section>
<script>
  ...
</script>

The generated table is static HTML. The inline script only adds local search and sort behavior; it does not fetch data at runtime.

PDF output:

\begin{blogasidebox}{Resources}
The HTML version renders a searchable, sortable data table from `resources.yml`.
\end{blogasidebox}

If the code fence has body text, or pdf="...", that text is used as the PDF fallback. Full YAML-to-LaTeX table rendering is deliberately not implemented yet because long prose fields usually typeset better as curated PDF content.

Obsidian behavior:

  • The fence stays readable in the note.
  • The YAML file remains a normal sidecar file you can edit directly.
  • Obsidian does not render the enhanced table unless a future vault plugin adds a preview renderer for .data-table.

Interactive Blocks

Interactive blocks reserve a place for a React component in HTML and provide a text fallback for PDF and non-JavaScript output.

Canonical form:

```{.interactive #int:tape-layout caption="Tape layout explorer"}
Drag files along a tape and compare the access cost of each arrangement.
```

Attributes:

Attribute Required Meaning
.interactive Yes Marks this as an interactive block.
#int:<id> Yes Semantic label. The int: prefix is for references.
caption="..." Optional Figure caption in HTML and PDF.
title="..." Optional Fallback caption when caption is absent.
component="..." Optional Override the component id.
pdf="..." Optional PDF-specific replacement text.
fallback="..." Optional General fallback text.

Component resolution:

#int:tape-layout
  -> component id: tape-layout
  -> component path: sites/<site>/src/interactives/tape-layout/index.tsx

HTML output:

<figure class="semantic-interactive" id="int:tape-layout">
  <div
    class="interactive-mount"
    data-interactive-id="tape-layout"
    data-interactive-props="..."
  >
    <p class="interactive-fallback">...</p>
  </div>
  <figcaption>Tape layout explorer</figcaption>
</figure>

The runtime loaded by InteractiveRuntime.astro imports:

sites/<site>/src/interactives/runtime.tsx

The runtime uses import.meta.glob("./*/index.tsx"), so each interactive component should live in its own directory with an index.tsx entrypoint.

Component props:

type InteractiveProps = {
  id: string;
  label: string;
  caption: string;
  description: string;
  fallback: string;
  attrs: Record<string, string>;
};

PDF output:

\begin{figure}[htbp]
\centering
\fbox{\begin{minipage}{0.86\linewidth}
Fallback text...
\end{minipage}}
\caption{Tape layout explorer}
\label{int:tape-layout}
\end{figure}

Short HTML-only form:

```interactive id="tape-layout" caption="Tape layout explorer"
Drag files along a tape and compare the access cost.
```

Use the canonical Pandoc-style form for any material that should export to PDF.

Cross-References

Use ordinary Markdown links to ids:

The proof of [Hall's theorem](#thm:hall) is constructive.
See the [tape explorer](#int:tape-layout) for the interactive version.

HTML keeps the link:

<a href="#thm:hall">Hall's theorem</a>

PDF converts the link to:

\cref{thm:hall}

Recommended label prefixes:

thm: theorem
lem: lemma
prop: proposition
cor: corollary
def: definition
ex: example
rem: remark
prf: proof
int: interactive
fig: figure
tbl: table
eq: equation

Do not reuse labels within one document.

Inline Footnotes

The HTML parser supports inline footnotes:

This is a sentence.^[This is the footnote.]

remarkInlineFootnotes rewrites these into normal Markdown footnote definitions before rendering.

The parser also supports multi-node inline footnotes when formatting spans the inside of the footnote.

For maximum PDF portability, ordinary footnotes are safest:

This is a sentence.[^note]

[^note]: This is the footnote.

Math

Use standard dollar math:

Inline math: $a^2 + b^2 = c^2$.

$$
\sum_{i=1}^n i = \frac{n(n+1)}{2}
$$

HTML uses remark-math plus rehype-katex.

PDF passes math through Pandoc to LuaLaTeX.

Avoid math syntax that only one renderer understands.

Symbols, Emoji, Chess, And Cards

Write Unicode directly in Markdown:

The position after ♘f3 is pleasant.
Hearts and diamonds are red: ♥ ♦.
Spades and clubs are black: ♠ ♣.
This note is important 🎯.

HTML wraps these glyphs in semantic spans:

<span
  class="symbol symbol-chess"
  data-symbol="white-knight"
  aria-label="white knight"
  ></span
>
<span class="symbol symbol-card" data-symbol="heart" aria-label="heart"
  ></span
>
<span class="symbol symbol-emoji">🎯</span>

PDF maps them to LaTeX macros:

emoji and emoji sequences -> \BlogEmojiText{...}
♔ ♕ ♖ ♗ ♘ ♙       -> chessfss white-piece macros
♚ ♛ ♜ ♝ ♞ ♟       -> chessfss black-piece macros
♥ ♡ ♦ ♢           -> red amssymb suit macros
♠ ♤ ♣ ♧           -> black amssymb suit macros
playing-card emoji -> \BlogEmojiText{...}

Do not write LaTeX-only suit or chess macros in source unless the document is PDF-only.

Code Fences

The code-fence language parser normalizes a few imported Quarto forms:

```{ojs}
viewof x = Inputs.range([0, 10])
```

becomes a JavaScript-highlighted fence:

```js
viewof x = Inputs.range([0, 10])
```

This pass is intentionally conservative. Add aliases only when imported source material needs them.

Parser Ordering

Parser order matters.

Use this order for Astro:

const remarkPlugins = [
  remarkGfm,
  remarkInlineFootnotes,
  remarkCodeFenceLanguages,
  remarkSemanticBlocks,
  remarkMath,
  remarkTypography,
  remarkCallouts,
  remarkSymbols,
];
const rehypePlugins = [rehypeKatex];

Reasons:

  • Inline footnotes should be normalized before other structure changes.
  • Code-fence languages should be normalized before semantic fences are read.
  • Semantic blocks should run before code fences become highlighted code.
  • Math should be parsed before typography and symbol wrapping can interfere.
  • Typography should run before symbol wrapping.
  • Callouts should run before symbol wrapping so marker detection sees plain text.
  • Symbols should run late because it emits raw HTML spans.

Astro Integration Checklist

Copy these files into the target Astro project:

src/remark/callouts.mjs
src/remark/code-fence-languages.mjs
src/remark/inline-footnotes.mjs
src/remark/semantic-blocks.mjs
src/remark/symbols.mjs
src/remark/typography.mjs
src/interactives/runtime.tsx
src/components/InteractiveRuntime.astro
src/styles/article.css

Wire the plugins in astro.config.mjs:

import mdx from "@astrojs/mdx";
import rehypeKatex from "rehype-katex";
import remarkGfm from "remark-gfm";
import remarkMath from "remark-math";
import remarkCallouts from "./src/remark/callouts.mjs";
import remarkCodeFenceLanguages from "./src/remark/code-fence-languages.mjs";
import remarkInlineFootnotes from "./src/remark/inline-footnotes.mjs";
import remarkSemanticBlocks from "./src/remark/semantic-blocks.mjs";
import remarkSymbols from "./src/remark/symbols.mjs";
import remarkTypography from "./src/remark/typography.mjs";

const remarkPlugins = [
  remarkGfm,
  remarkInlineFootnotes,
  remarkCodeFenceLanguages,
  remarkSemanticBlocks,
  remarkMath,
  remarkTypography,
  remarkCallouts,
  remarkSymbols,
];
const rehypePlugins = [rehypeKatex];

export default defineConfig({
  integrations: [mdx({ remarkPlugins, rehypePlugins })],
  markdown: { remarkPlugins, rehypePlugins },
});

Add the runtime to post pages:

---
import InteractiveRuntime from "@/components/InteractiveRuntime.astro";
---

<article class="article-content">
  <Content />
</article>
<InteractiveRuntime />

Make sure the article body uses the CSS class that article.css targets.

PDF Integration Checklist

Copy these files:

scripts/export-pdf.mjs
scripts/pandoc-filters/semantic-blocks.lua
scripts/pandoc-templates/semantic-preamble.tex

Add a package script:

{
  "scripts": {
    "pdf:post": "node scripts/export-pdf.mjs"
  }
}

Export a post:

npm run pdf:post -- path/to/post/index.md --out /tmp/post.pdf

The export command currently uses:

markdown+fenced_code_attributes+tex_math_dollars+tex_math_single_backslash+footnotes+raw_html

and includes the shared preamble with:

--include-in-header scripts/pandoc-templates/semantic-preamble.tex

If LuaLaTeX cannot write its font cache in a sandboxed environment, set a writable TeX cache:

TEXMFVAR=/tmp/texmf-var TEXMFCONFIG=/tmp/texmf-config \
  npm run pdf:post -- path/to/post/index.md --out /tmp/post.pdf

Obsidian Behavior

The syntax is chosen to remain readable in Obsidian:

  • Plain Markdown stays plain Markdown.
  • > [!tip] and > [!tip: Title] render as Obsidian callouts.
  • Fenced semantic blocks remain visible as code blocks when Obsidian does not know how to render them.
  • Interactive blocks show their fallback description in the source.

Obsidian will not execute Astro interactives or render LaTeX environments as final theorem boxes. That is acceptable: Obsidian is the editing surface, not the final renderer.

Adding A New Semantic Feature

When adding a feature, update all three layers:

  1. Source syntax in this document.
  2. Astro parser and CSS.
  3. Pandoc Lua filter and LaTeX preamble.

Use this checklist:

  • Define the Markdown syntax and at least one canonical example.
  • Decide whether it needs a label and cross-reference behavior.
  • Decide the HTML output shape and CSS classes.
  • Decide the PDF output shape and required LaTeX packages.
  • Add an Obsidian-readable fallback.
  • Add a smoke-test Markdown file and render both HTML and PDF.
  • Document migration rules for old content, if any.

Known Tradeoffs

The current system is robust, but it has deliberate limits:

  • It is a convention layer over Markdown, not a full new markup language.
  • HTML and PDF are separate renderers. The contract is shared, but exact visual parity is not the goal.
  • Interactives are HTML-first. PDF receives a figure-like fallback.
  • Raw HTML is allowed for web-only material but should not be used for semantic features that need PDF output.
  • The blog repo currently duplicates parser files across seven sites. A future shared parser package would reduce drift.

Minimal Smoke Test

Use this sample when porting the parser to another project:

---
title: "Parser Smoke Test"
description: "Checks callouts, environments, links, symbols, math, and interactives."
pubDate: "2026-01-31"
---

This sentence has an em dash --- and inline math $a^2+b^2=c^2$.

> [!tip: Explicit title]
>
> This callout has a title, a chess piece ♘, a card suit ♥, and an emoji 🎯.

```{.env .definition #def:test title="Test object" html="callout"}
A test object is something we can refer to later.
```

See [the definition](#def:test).

```{.interactive #int:test-widget caption="Test widget" pdf="The HTML version contains a small interactive test widget."}
This fallback appears when the interactive component is unavailable.
```

Expected HTML:

  • One em dash.
  • KaTeX-rendered math.
  • A titled tip callout.
  • A semantic definition block with id="def:test".
  • A normal anchor link to #def:test.
  • A framed interactive fallback if no component exists.
  • Wrapped symbol spans for , , and 🎯.

Expected PDF:

  • A LuaLaTeX em dash.
  • Typeset math.
  • A titled tcolorbox tip.
  • A real definition environment with \label{def:test}.
  • A \cref{def:test} reference.
  • A figure fallback for int:test-widget.