Skip to content

feat: assemble the open-source pages from lancedb/lancedb - #355

Open
jackye1995 wants to merge 16 commits into
mainfrom
jack/a3-move
Open

jackye1995 wants to merge 16 commits into
mainfrom
jack/a3-move

Conversation

@jackye1995

Copy link
Copy Markdown
Contributor

A3 of the documentation migration. The open-source pages move to lancedb/lancedb#4122 and are assembled from there.

Merge order: #352 → this → and lancedb/lancedb#4122 must land first, since CI assembles from that repository's main.

What moves and what stays

102 pages, the static assets, the generated snippets and the test suite they come from all move to lancedb/lancedb/docs/web. This repository keeps what it still owns: the Geneva pages until Stage G replaces them, the generated dataset cards, the OpenAPI spec, and the Enterprise pages until A5 moves them to sophon.

Navigation is the interesting part

It is authored once, in the root that owns the pages, so that root can be served on its own — a contributor runs mint dev in docs/web with no second checkout. Everything this repository still holds is contributed as a fragment saying where each entry belongs: the chain of groups above it and the sibling it follows.

Appending was not good enough. Sidebar order is what a reader navigates by, and appending dropped Geneva below Support and Datasets past Use Cases — a silent reordering of the whole site.

scripts/split_nav.py generates that fragment by walking the original navigation rather than by hand. That is why it contains six entries rather than ninety-seven: a container that moves wholesale is carried across intact instead of rebuilt from its children. A5 reuses it when the Enterprise pages move.

Three things that had to match exactly

Each was caught by the harness, not by review:

  • The openapi block names a spec file this repository owns, and Mintlify refuses to build at all when it is missing. It is lifted out of the base navigation and restored by the fragment.
  • Key order in docs.json is load-bearing. Mintlify hashes its CSS and JS bundles from those bytes, so moving navigation renamed every asset on every page — 179 differences from one reordered key.
  • ensure_ascii likewise. The banner text carries an em-dash; escaping it differently changed the same bytes.

CI

Now checks out lancedb/lancedb@main and assembles twice, comparing the two. The site no longer exists as a single tree anywhere, so comparing the assembly against docs/ is no longer available — determinism is what the comparison can still establish, and a stale stored baseline is exactly what this harness was built to avoid.

Verified

  • assembled tree byte-identical to the one it replaces — diff -rq clean against the pre-move source
  • two consecutive assemblies identical
  • full mint export comparison: VERDICT: EQUIVALENT, zero real differences
  • link check clean
  • the standalone site in lancedb/lancedb builds and renders all 102 pages

The first published build dropped docs/.cursor/ silently: actions/upload-artifact
skips hidden files by default, so the assembled tree had 268 files and the
`assembled` branch had 267.

Editor configuration is not site content and should not be published, but it
should not vanish by accident either. The assembler now excludes dot-prefixed
paths deliberately and names what it dropped, so the assembled tree and the
published tree are the same thing — which is the guarantee the whole pipeline
rests on.
Two problems found while verifying the first publish to `assembled`.

The published branch had 267 files where the assembler produced 268:
actions/upload-artifact skips hidden files by default and silently dropped
docs/.cursor/. The first attempt excluded dotfiles from the assembly instead,
which was wrong — mint export carries that file today, so removing it would have
quietly deleted a published file at cutover. The harness caught that, which is
the strongest evidence so far that it works. The upload now includes hidden
files, and the assembler publishes what it assembled.

The harness itself was intermittently failing on identical trees. Mintlify
renders the OpenAPI reference non-deterministically: response code blocks come
out syntax-highlighted on one run and plain on the next, ~2 KB across ~78
fragments. One comparison passed, the next failed, on byte-identical input. A
gate that fails at random is one people learn to route around.

That subtree is now compared for presence but not for bytes, and only that
subtree. It gives up nothing about the assembler, which passes openapi.yml
through byte-identically and cannot influence one render differently from the
other; a page appearing or disappearing is still caught. Verified: five
consecutive comparisons pass, an authored-page change is caught, a reference
page being removed is caught, and a reference page's contents are knowingly not.
* feat: anchor the tables pages

First area of the anchor pass. Every section heading in docs/tables now carries
an explicit `{#anchor}` — 114 of them — which is the identity Enterprise
overlays attach to from A5 and which survives the heading being reworded.

Anchors are taken from the ids the site already renders, not derived from the
heading text. That distinction turned out to matter: Mintlify applies smart
quotes before slugifying, so `## What's next?` renders as `what’s-next` with a
curly apostrophe, and a derived slug would have silently changed the id and
broken every existing deep link to it. Reading the ids back from `mint export`
makes the pass correct by construction rather than by reimplementing rules we
would have to keep in sync.

Headings that begin with a number need their period escaped. `### 1. Setup`
renders with its number until an explicit anchor is added, at which point
Mintlify re-parses the text, treats the number as an ordered list marker, and
drops it from both the heading and the table of contents. 117 headings across 17
pages start this way, so the pass would have quietly renumbered a good deal of
the site.

Verified: the exported site is unchanged. All six existing deep links into
tables still resolve.

* feat: anchor the remaining pages

Completes the anchor pass. 805 of 810 section headings across the authored pages
now carry an explicit `{#anchor}`, the identity Enterprise overlays attach to
from A5 and the one thing that survives a heading being reworded.

Three more cases the tables pilot had not reached, each of which would have
corrupted anchors silently:

Setext headings. Eleven h2s in the reranking pages are written as text over a
rule rather than with hashes. Mintlify gives them ids like any other heading, so
a parser that only saw hashes consumed the rendered ids out of order and handed
every later heading on the page the wrong anchor — plausible names attached to
the wrong sections. They are rewritten to hashes, which renders identically, and
a strict count check now refuses to write anything when source headings and
rendered ids disagree.

Headings below h4. Mintlify emits no id for h5, so there is nothing to read back
and nothing to preserve. The eight in the corpus, all API method names on one
page, keep their generated markup and stay unanchored.

Ampersands. Mintlify keeps `&` in a generated id but strips it from an explicit
anchor, so `observability-&-performance` cannot be written down: any anchor set
on those headings changes the id and breaks links to it. Five headings join two
words this way; they keep their generated id. If A5 needs to attach to one,
rewording the heading is the honest fix rather than silently moving it.

Verified: the exported site is unchanged.

* refactor: retitle headings whose anchors carried punctuation

Eighteen headings across twelve pages produced anchors containing `&`, `/`, or
curly quotes. Ampersands were the pressing case — Mintlify keeps `&` in a
generated id but strips it from an explicit anchor, so those five headings could
not be anchored at all — but slashes made anchors read like paths and curly
quotes made them non-ASCII, and neither belongs in a key that Enterprise
overlays will be written against.

The punctuation is incidental in every case, so the headings say the same thing
without it: "Observability & performance" becomes "and", "S3 / GCS / Azure Blob"
becomes a comma list, "What's next?" becomes "Next steps". Nothing linked to any
of the old anchors, so nothing breaks.

Underscores and the plus in `analyze_plan`, `explain_plan`, `max_pooling`,
`approx_mode`, `torch_col` and `100B+ row scale` are left alone. Those
characters are part of an API name or a quantity rather than punctuation, they
are safe in a URL fragment, and renaming them would misname the thing the
heading documents.

All 810 headings now carry an anchor. The twelve pages whose rendered output
changed are exactly the twelve retitled here.

* refactor: retitle the last headings whose anchors held identifiers

Six headings named an API identifier or a quantity directly — `analyze_plan`,
`explain_plan`, `approx_mode`, `torch_col`, `max_pooling`, `100B+ row scale` —
so their anchors carried an underscore or a plus.

Renaming the identifier would have misnamed what the section documents, so the
headings now describe what the section does and the identifier stays in the
prose, where it was already being used: `analyze_plan` appears 14 times in that
page, `explain_plan` 11, `max_pooling` 9. Nothing about the API is lost, and
the heading reads better for it.

Every anchor in the corpus is now plain `[a-z0-9-]`.
The pages describing the open-source client now live in that repository, beside
the code, and are assembled from there. This repository keeps what it still
owns — the Geneva pages until Stage G replaces them, the generated dataset
cards, the OpenAPI spec, and the Enterprise pages until they move to sophon.

Navigation is the interesting part. It is authored once, in the root that owns
the pages, so that root can also be served on its own — a contributor previews
the open-source documentation with `mint dev` and no second checkout. Everything
this repository still holds is contributed as a fragment that says where each
entry belongs: the chain of groups above it and the sibling it follows.
Appending was not good enough, because sidebar order is what a reader navigates
by, and appending dropped Geneva below Support and Datasets past Use Cases.

`scripts/split_nav.py` produced that fragment by walking the original navigation
rather than by hand, which is how the six entries stayed six rather than
ninety-seven. A5 reuses it when the Enterprise pages move.

Three things had to match exactly, and each was found by the harness:

The `openapi` block names a spec file this repository owns, and Mintlify refuses
to build at all when it is missing — so it is lifted out of the base navigation
and restored by the fragment.

Key order in `docs.json` is load-bearing. Mintlify hashes its CSS and JS bundles
from those bytes, so moving `navigation` renamed every asset on every page.

`ensure_ascii` likewise: the banner text carries an em-dash, and escaping it
differently changed the same bytes.

CI now checks out lancedb/lancedb at main and assembles twice, comparing the two.
The site no longer exists as a single tree anywhere, so a direct comparison
against `docs/` is not available; determinism is what the comparison can still
establish.

Verified: the assembled tree is byte-identical to the one it replaces.
#350 was merged into deploy-freeze while the freeze was in effect, so the fix is
live but absent from main and would have been lost at cutover: the redirect for
/hybrid-search and the removal of a redundant layers diagram from the Enterprise
architecture page.

Redirects are now owned by this repository's navigation fragment, so that is
where the redirect goes. The diagram removal is applied as a patch rather than
by taking the whole file from deploy-freeze, which would have reverted that
page's anchors along with it.

Anything else that lands on deploy-freeze during the freeze needs the same
treatment — it is the branch production serves, and it does not flow back.
The eleven Enterprise pages move to sophon, and this repository assembles them
from there. Its navigation fragment keeps only what it still owns: the Geneva
group and the Datasets tab.

Adding a third root exposed two assumptions in the assembler that only held
while there were two. Both were caught by the collision guard rather than by
review, which is the guard doing its job.

Every root now ships a complete `docs.json` so it can be previewed on its own —
that is how a contributor reads the open-source pages without this repository,
and how the Enterprise pages are reviewed in isolation. Only the first reference
root's is the published navigation; the others are ignored rather than treated
as a conflict.

Shared assets legitimately appear in more than one root, because each root needs
them to render alone. Identical bytes are no longer an error. Differing bytes
still are: the site would otherwise depend on the order roots are declared in.
Retarget the Geneva group at the new Feature engineering section.

Two assembler fixes the restructure surfaced. Redirects now accumulate across
roots instead of the last fragment replacing them, so a repo can redirect its
own moved pages; a source claimed twice with different destinations is an error.
A root's README is no longer copied, which was publishing it at a URL nothing
links to.

The restructure also dropped the set directive supplying the REST reference its
spec, which assembled cleanly and rendered a group with no operations in it.
Validation now fails when the navigation never names the spec.
Removes the 34 pages, their nav group, and the redirect whose destination they
were. Feature engineering is now documented in lancedb and sophon.
The REST pages moved out of their own tab into Basics.
The REST group is gone from the sidebar, so nothing references the spec from
navigation and the old guard would fail every build. Check the synced file is
present instead, which is the invariant that still holds.
Robotics moves above the image groups and gains the LeRobotDataset guide, which
is how you read the two datasets already in it.
Same wasted level as the other tabs: a group named Overview holding one page.
An overlay page replaces the reference page at the same path. The reference root
owns every path and the navigation, so an overlay never adds or moves a page --
it only says more about one that exists. sophon switches to that role.
Carries #364 from deploy-freeze.
@jackye1995
jackye1995 changed the base branch from jack/exclude-dotfiles to main September 25, 2026 23:56
Nineteen modify/delete conflicts, all resolved as delete: main edited pages that
this branch moves to lancedb/lancedb, and every one of those edits has already
been ported there and verified line by line. Two of the nineteen are the
snippet-generating tests, which move to docs/web-tests/ in the same repository.

Three files merged cleanly that had to be removed as well -- training/data-loading.mdx
and two images added by #344 and #354. A clean merge is not the same as a correct
one when the branch's whole purpose is to delete the directory they land in.

Keeps main's pin of mint@4.2.888, which exists to stop newer releases rendering
nondeterministic OpenAPI examples -- the same byte-comparability the harness in
this branch checks.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant