feat(docs): serve Markdown to clients that ask for it - #12596
Conversation
|
This pull request is part of a Mergify stack:
|
Merge Protections🟢 All 7 merge protections satisfied — ready to merge. Show 7 satisfied protections🟢 ⛓️ Depends-On RequirementsRequirement based on the presence of
🟢 🤖 Continuous Integration
🟢 👀 Review Requirements
🟢 Enforce conventional commitMake sure that we follow https://www.conventionalcommits.org/en/v1.0.0/
🟢 🔎 Reviews
🟢 📕 PR description
🟢 🚦 Auto-queueWhen all merge protections are satisfied, this pull request will be queued automatically. |
1171c35 to
73f5d6b
Compare
Revision history
|
There was a problem hiding this comment.
Pull request overview
Adds Cloudflare Pages middleware-based content negotiation so clients explicitly requesting text/markdown receive the existing per-page *.md “twin”, while keeping CDN caching correct via Vary: Accept and avoiding Worker invocations for most static assets.
Changes:
- Introduces
prefersMarkdown()+isNegotiablePage()utilities (with Vitest coverage) to drive safe Accept-header negotiation. - Adds Pages middleware to rewrite negotiable page requests to their
*.mdtwin, return Markdown-formatted 404s when appropriate, and mergeAcceptintoVaryfor negotiated media types. - Adds
<link rel="alternate" type="text/markdown">in page head, and introduces_routes.jsonplus lint/format ignores for local Wrangler artifacts.
Reviewed changes
Copilot reviewed 7 out of 8 changed files in this pull request and generated 1 comment.
Show a summary per file
| File | Description |
|---|---|
| src/util/acceptMarkdown.ts | Implements Accept parsing + policy for when to serve Markdown. |
| src/util/acceptMarkdown.test.ts | Unit tests covering key negotiation scenarios and asset/page detection. |
| src/components/HeadSEO.astro | Publishes the Markdown twin via an alternate link tag. |
| public/_routes.json | Limits Pages Function routing to page requests by excluding major static asset paths. |
| functions/_middleware.ts | Rewrites requests to *.md when appropriate and sets Vary: Accept on negotiated responses. |
| eslint.config.js | Ignores local .wrangler/ artifacts for linting. |
| biome.json | Excludes .wrangler/ from Biome formatting/linting. |
| .gitignore | Ignores local Wrangler build artifacts. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
73f5d6b to
80ec774
Compare
`getOgImageUrl` strips the leading and trailing slashes off the pathname to build the image filename. For the homepage that pathname is `/`, so stripping left an empty string and the lookup missed — every docs page had an OpenGraph image and the homepage shipped `<meta property="og:image">` with no content. The homepage's collection id is `index`, which is what `getStaticPaths` names its image, so fall back to that when the slug comes out empty. Covered by a regression test that fails against the old expression. The generated-image set comes from the content collection and needs the Astro build pipeline, so the test stubs it and exercises the derivation, which is the half that was wrong. Change-Id: Ifb9a23ea2caa20d28489a4f21363d85ed5e3342c
Docset cards rendered their title as `h4`. Almost every grid sits directly under an `##`, so the outline jumped h2 to h4 — including on the homepage, whose whole body is the "Products" grid. Screen readers and anything parsing the document outline read that as a missing level. Default the card heading to `h3` and make it a prop, because one grid does belong at h4: the "Components" grid in `ci-insights.mdx` is nested under an `### Components`, where h3 would make the cards siblings of their own section heading instead of children. Change-Id: Ie01f4c2b03f1f2b7f798ae89b056135a5b00800e
The Mergify OpenAPI 3.1 document is already deployed — it is what the API Reference pages are generated from — but only at `/api-schemas.json`, a filename that exists nowhere outside this repository. Every OpenAPI client, SDK generator and crawler probes `/openapi.json`, so nothing finds it and the docs read as a site with no API at all. Serve the same bytes at `/openapi.json`. The route does not transform the document: the spec is synced from the engine repository, and two spellings of it that could disagree would be worse than one obscure path. Alongside it, three other entry points that were undiscoverable: - `<link rel="service-desc">` in every page head, the IANA relation for "the description of this site's API" (RFC 8631), plus one for `llms.txt`. - `robots.txt` had no `Sitemap:` line, so crawlers had to guess `sitemap-index.xml` rather than be told. - `/developers` is the path people and tools guess for a developer portal and was a 404; it now redirects to the API reference. Change-Id: I1b625f960363c8427d5282c052fee74111bf07fa
Every page already ships its Markdown source at `<path>.md`, which is what the "View as Markdown" button and `llms.txt` link to. But the convention agents reach for first is to request the page's own URL with `Accept: text/markdown` (https://acceptmarkdown.com), and that returned HTML. The site is a static build, so a Cloudflare Pages middleware is the only layer that sees request headers. It rewrites to the `.md` twin when Markdown is asked for, answers 404s in the format that was requested, and merges `Accept` into `Vary` on both variants — without it a CDN would hand whichever variant it cached first to everyone, which is the failure this is most likely to produce in production and the least likely to be noticed. The `Vary` goes on `304` responses too, since those are exactly the ones a cache is about to act on. Markdown is served only when `text/markdown` is named explicitly and not outranked by `text/html`. Browsers send `text/html,...,*/*;q=0.8`, so honouring wildcards would serve raw Markdown to every human visitor; the tests pin the real Chrome and Firefox Accept headers against that. HTML is served as the fallback only on a `404` from the `.md` route, not on anything that is merely not `ok`: `304 Not Modified` is the normal answer to a client revalidating Markdown it already holds, and treating it as a missing page answered it with HTML. Three things about adding a root `_middleware` needed care, all three verified against the real Pages runtime with `wrangler pages dev dist`: Cloudflare documents that `_redirects` are not applied to requests served by Functions, and this repository has 99 of them. `next()` does still route through the asset server, so every redirect, the `_headers` rules, the `.md` routes and the static 404 behave exactly as before. `next()` is called with an explicit request everywhere. A bare `next()` is documented as forwarding the original request, but once the middleware has asked for the Markdown twin the runtime forwards *that* request again — so `/api` and `/cli`, which are built from `src/pages/` and have no `.md`, were answered with the Markdown 404 instead of their own HTML. The test double refuses a call with no request so this cannot come back. A root middleware otherwise turns every request into a Worker invocation, assets included — roughly thirty per page view, none of which can be negotiated. `_routes.json` excludes the bundle, the search index, the OpenGraph images and the static files by name, so only page requests reach the Function. Also ignore `.wrangler/` — running the Pages runtime locally drops generated bundles there, and eslint linted them. Change-Id: I52f9d2a1c70248b1406412129a633cc3bee219de
Merge Queue Status
This pull request spent 3 minutes 11 seconds in the queue, including 2 minutes 41 seconds running CI. Required conditions to merge
|
Every page already ships its Markdown source at
<path>.md, which is what the"View as Markdown" button and
llms.txtlink to. But the convention agentsreach for first is to request the page's own URL with
Accept: text/markdown(https://acceptmarkdown.com), and that returned HTML.
The site is a static build, so a Cloudflare Pages middleware is the only layer
that sees request headers. It rewrites to the
.mdtwin when Markdown is askedfor, answers 404s in the format that was requested, and merges
AcceptintoVaryon both variants — without it a CDN would hand whichever variant itcached first to everyone, which is the failure this is most likely to produce
in production and the least likely to be noticed. The
Varygoes on304responses too, since those are exactly the ones a cache is about to act on.
Markdown is served only when
text/markdownis named explicitly and notoutranked by
text/html. Browsers sendtext/html,...,*/*;q=0.8, so honouringwildcards would serve raw Markdown to every human visitor; the tests pin the
real Chrome and Firefox Accept headers against that.
HTML is served as the fallback only on a
404from the.mdroute, not onanything that is merely not
ok:304 Not Modifiedis the normal answer to aclient revalidating Markdown it already holds, and treating it as a missing
page answered it with HTML.
Three things about adding a root
_middlewareneeded care, all three verifiedagainst the real Pages runtime with
wrangler pages dev dist:Cloudflare documents that
_redirectsare not applied to requests served byFunctions, and this repository has 99 of them.
next()does still routethrough the asset server, so every redirect, the
_headersrules, the.mdroutes and the static 404 behave exactly as before.
next()is called with an explicit request everywhere. A barenext()isdocumented as forwarding the original request, but once the middleware has
asked for the Markdown twin the runtime forwards that request again — so
/apiand/cli, which are built fromsrc/pages/and have no.md, wereanswered with the Markdown 404 instead of their own HTML. The test double
refuses a call with no request so this cannot come back.
A root middleware otherwise turns every request into a Worker invocation,
assets included — roughly thirty per page view, none of which can be
negotiated.
_routes.jsonexcludes the bundle, the search index, theOpenGraph images and the static files by name, so only page requests reach the
Function.
Also ignore
.wrangler/— running the Pages runtime locally drops generatedbundles there, and eslint linted them.
Depends-On: #12595