diff --git a/README.md b/README.md index 005d789..a8e827b 100644 --- a/README.md +++ b/README.md @@ -78,8 +78,9 @@ npm run check For a live spot-check that the public `/check`, `/demo`, `/methodology`, `/packages`, `/support`, `/terms`, and `/privacy` pages on the deployed site still show the anonymous proof check, proof loop, stated limits, package ladder, and no-ranking promise the README -makes, and that `/llms.txt`, `/sitemap.xml`, `/robots.txt`, `/api/health`, -`/api/deep-health`, and the `POST /api/public-check` route are still served: +makes, that `/llms.txt`, `/sitemap.xml`, `/robots.txt`, `/api/health`, +`/api/deep-health`, and the `POST /api/public-check` route are still served, and that +`www.seofixkit.com` still 301-redirects onto the apex host: ```bash npm run audit:live-promise @@ -202,7 +203,7 @@ Access-link, payment-success, repair-started, delivery-ready, and daily ops dige ## Custom domain -`seofixkit.com` is the production domain. The Worker config attaches both the apex and `www` hostnames: +`seofixkit.com` is the canonical host. The Worker config attaches both the apex and `www` hostnames: ```jsonc "routes": [ @@ -211,6 +212,12 @@ Access-link, payment-success, repair-started, delivery-ready, and daily ops dige ] ``` +The `www` hostname stays attached only so its requests reach the Worker, which +permanently 301-redirects every `www.seofixkit.com` request onto the apex host +with the path and query intact. Every URL the Worker serves (page canonicals, +social tags, `robots.txt`, `sitemap.xml`, `llms.txt`, and fixture URLs) is +generated from the apex origin, so canonicals and the sitemap are apex-only. + ## Product boundary This MVP can beat weak SEO audit tools on accuracy and fix quality. Competitor benchmarks are public homepage proof snapshots only. Self-serve rendered repair crawl currently supports up to 1,000 pages inside a normal report, while sitemap inventory can discover up to 50,000 URLs and separate large-crawl jobs can store 50,000-page frontier, batch, retry, proof, and incremental-crawl metadata. Large crawls are early access and must not be sold as completed 50,000-page rendered validation until every large-crawl batch has page-level proof and merge readiness is clear. Crawl intelligence is based on the rendered crawl and the sitemap inventory sample; orphan URLs and cannibalization are repair heuristics, not full-site rank or index data. Backlink data starts with supplied/imported rows and link-edge history; it does not provide proprietary backlink discovery. Keyword/rank data starts with supplied/imported Search Console or rank-tracker rows and observation history; it does not provide live keyword volume providers, traffic estimates, or continuous rank tracking yet. Local SEO audit uses supplied business details and citation URLs; it does not scrape private Google Business Profile data or discover every citation automatically. Platform SEO audit uses rendered public proof only; it does not log into WordPress, Shopify, WooCommerce, Magento, Google Business Profile, or private plugin/admin settings. AI Answer Readiness is proof-derived from rendered content, schema, canonical/link clarity, sitemap context, and optional `/llms.txt`; it does not sample answer engines or monitor citations. Growth opportunities are draft-only briefs from verified keyword, competitor, AI-readiness, or crawl gaps; they do not auto-publish, create CMS drafts, open pull requests, or promise rankings, traffic, citations, or revenue. Repair agent actions, implementation packs, and repair proof receipts are reviewable records, drafts, handoff documents, and proof artifacts only; they do not publish CMS changes, open GitHub pull requests, merge code, or call provider admin APIs. Paid Growth Add-On billing/integrations remain roadmap. Live AI-engine visibility tracking and AI citation monitoring are not live. diff --git a/scripts/live-promise-spot-check.mjs b/scripts/live-promise-spot-check.mjs index e79d672..b0cc43a 100644 --- a/scripts/live-promise-spot-check.mjs +++ b/scripts/live-promise-spot-check.mjs @@ -3,13 +3,14 @@ import { pathToFileURL } from "node:url"; // Live spot-check for the public promise surfaces the README "What is live in // this repo" section relies on: /demo (proof loop), /methodology (limits), // /packages (package ladder), /check (anonymous one-page check), the -// no-ranking /support, /terms, and /privacy pages, and the machine surfaces +// no-ranking /support, /terms, and /privacy pages, the machine surfaces // the README says stay served by the Worker (/llms.txt, /sitemap.xml, // /robots.txt, /api/health, /api/deep-health, and the POST /api/public-check -// route). Each must be served by the deployed Worker with the copy that backs -// the claims. This is the repeatable "spot-check" half of the lane-2 promise -// audit; the offline regression lock lives in shared/promise-audit.test.mjs -// and worker/routes/pages.test.mjs. +// route), and the canonical-host promise that every www.seofixkit.com request +// 301-redirects onto the apex host. Each must be served by the deployed Worker +// with the copy that backs the claims. This is the repeatable "spot-check" +// half of the lane-2 promise audit; the offline regression lock lives in +// shared/promise-audit.test.mjs and worker/routes/pages.test.mjs. // // Opt-in script, not part of `npm run check` (CI stays offline-only; the // offline regression lock for the same claims is in the check pipeline): @@ -67,13 +68,17 @@ export async function spotCheckPublicPages({ fetcher = fetch, timeoutMs = DEFAULT_TIMEOUT_MS }) { - const checks = [...publicPageSpotChecks(baseUrl), ...publicSurfaceSpotChecks(baseUrl)]; + const checks = [ + ...publicPageSpotChecks(baseUrl), + ...publicSurfaceSpotChecks(baseUrl), + ...canonicalHostSpotChecks(baseUrl) + ]; const results = []; for (const check of checks) { const { method = "GET", body } = check; const response = await fetchWithTimeout( - `${baseUrl}${check.path}`, - { method, body }, + check.url || `${baseUrl}${check.path}`, + { method, body, ...(check.redirectManual ? { redirect: "manual" } : {}) }, timeoutMs, fetcher ); @@ -89,7 +94,13 @@ export async function spotCheckPublicPages({ if (check.contentType && !new RegExp(check.contentType, "i").test(contentType)) { failures.push(`served as ${contentType || "no content-type"} instead of ${check.contentType}`); } - for (const { reason, match } of check.expectations) { + for (const { name, value, reason } of check.expectedHeaders || []) { + const header = response.headers.get(name) || ""; + if (!hasContent(header, value)) { + failures.push(reason || `missing ${name} header matching ${typeof value === "string" ? value : value.toString()}`); + } + } + for (const { reason, match } of check.expectations || []) { if (!hasContent(text, match)) { failures.push(reason); } @@ -277,6 +288,42 @@ export function publicSurfaceSpotChecks(baseUrl) { ]; } +// Canonical-host promise: the README "Custom domain" section claims every +// www.seofixkit.com request 301-redirects onto the apex host with its path and +// query intact, so canonicals, robots.txt, and sitemap.xml stay apex-only. +// Redirects are checked with `redirect: "manual"` so the 301 itself is +// observable instead of being silently followed. +export function canonicalHostSpotChecks(baseUrl) { + const apex = new URL(baseUrl); + const wwwOrigin = `https://www.${apex.hostname}`; + return [ + { + path: "www.seofixkit.com/", + name: "www.seofixkit.com 301-redirects onto the apex host", + url: `${wwwOrigin}/`, + redirectManual: true, + acceptStatuses: [301], + expectedHeaders: [ + { name: "location", value: `${baseUrl}/`, reason: "redirects to the apex root" } + ] + }, + { + path: "www.seofixkit.com/check", + name: "www.seofixkit.com deep paths redirect with path and query intact", + url: `${wwwOrigin}/check?utm_source=spot-check`, + redirectManual: true, + acceptStatuses: [301], + expectedHeaders: [ + { + name: "location", + value: `${baseUrl}/check?utm_source=spot-check`, + reason: "redirect preserves the path and query" + } + ] + } + ]; +} + function hasContent(text, match) { return typeof match === "string" ? text.includes(match) : match.test(text); } diff --git a/scripts/live-promise-spot-check.test.mjs b/scripts/live-promise-spot-check.test.mjs index 15df129..6da242d 100644 --- a/scripts/live-promise-spot-check.test.mjs +++ b/scripts/live-promise-spot-check.test.mjs @@ -3,7 +3,7 @@ import test from "node:test"; import { rootSitemap } from "../shared/audit-engine.js"; import { demoHtml, llmsText, methodologyHtml, packagesHtml, privacyHtml, supportHtml, termsHtml } from "../worker/routes/pages.js"; import { checkHtml } from "../worker/routes/public-check.js"; -import { publicPageSpotChecks, publicSurfaceSpotChecks, spotCheckPublicPages } from "./live-promise-spot-check.mjs"; +import { canonicalHostSpotChecks, publicPageSpotChecks, publicSurfaceSpotChecks, spotCheckPublicPages } from "./live-promise-spot-check.mjs"; const origin = "https://seofixkit.com"; const pages = { @@ -45,6 +45,14 @@ const surfaces = { function pageFetcher(overrides = {}) { return async (rawUrl, options = {}) => { const url = new URL(rawUrl); + if (url.hostname.startsWith("www.")) { + const apexUrl = new URL(rawUrl); + apexUrl.hostname = apexUrl.hostname.replace(/^www\./, ""); + return new Response(null, { + status: 301, + headers: { location: apexUrl.toString() } + }); + } if (url.pathname in surfaces) { return surfaces[url.pathname](); } @@ -80,9 +88,20 @@ test("live spot-check covers the public machine surfaces", () => { ); }); +test("live spot-check covers the www-to-apex canonical redirect", () => { + assert.deepEqual( + canonicalHostSpotChecks(origin).map((check) => check.path), + ["www.seofixkit.com/", "www.seofixkit.com/check"] + ); + for (const check of canonicalHostSpotChecks(origin)) { + assert.equal(check.redirectManual, true, "redirect checks must observe the 301 itself"); + assert.equal(check.acceptStatuses[0], 301); + } +}); + test("live spot-check passes against the shipped public page copy", async () => { const results = await spotCheckPublicPages({ baseUrl: origin, fetcher: pageFetcher() }); - assert.equal(results.length, 13); + assert.equal(results.length, 15); for (const result of results) { assert.deepEqual(result.failures, [], `${result.path} must pass: ${result.name}`); } @@ -136,6 +155,42 @@ test("live spot-check flags a stale Worker serving the SPA fallback", async () = assert.deepEqual(demo.failures, [], "worker-rendered pages must not be flagged as stale"); }); +test("live spot-check flags a www host that stops redirecting to the apex", async () => { + const noRedirectFetcher = async (rawUrl, options = {}) => { + const url = new URL(rawUrl); + if (url.hostname.startsWith("www.")) { + return htmlResponse("
"); + } + return pageFetcher()(rawUrl, options); + }; + const results = await spotCheckPublicPages({ baseUrl: origin, fetcher: noRedirectFetcher }); + const redirect = results.find((result) => result.name.includes("301-redirects onto the apex host")); + assert.ok( + redirect.failures.some((failure) => failure.includes("HTTP 200")), + "a www host serving 200 instead of 301 must be reported" + ); + assert.ok( + redirect.failures.some((failure) => failure.includes("redirects to the apex root")), + "a redirect without the apex Location header must be reported" + ); +}); + +test("live spot-check flags a www redirect that drops the path or query", async () => { + const rootOnlyFetcher = async (rawUrl, options = {}) => { + const url = new URL(rawUrl); + if (url.hostname.startsWith("www.")) { + return new Response(null, { status: 301, headers: { location: `${origin}/` } }); + } + return pageFetcher()(rawUrl, options); + }; + const results = await spotCheckPublicPages({ baseUrl: origin, fetcher: rootOnlyFetcher }); + const redirect = results.find((result) => result.name.includes("path and query intact")); + assert.ok( + redirect.failures.some((failure) => failure.includes("redirect preserves the path and query")), + "a redirect that drops the path and query must be reported" + ); +}); + test("live spot-check flags llms.txt that no longer lists the anonymous check", async () => { const overrides = { "/llms.txt": () => textResponse(llmsText(origin).replaceAll(`${origin}/check`, `${origin}/gone`)) diff --git a/worker/index.js b/worker/index.js index 466c7ad..b9b9a20 100644 --- a/worker/index.js +++ b/worker/index.js @@ -137,6 +137,22 @@ import { saveLargeRenderedCrawlBatchProof } from "./routes/large-crawls.js"; +// Canonical host: `www.seofixkit.com` is a serving alias that 301-redirects +// onto the apex host, and every URL the Worker emits (page canonicals, social +// tags, robots.txt, sitemap.xml, llms.txt, fixture URLs) is generated from the +// apex origin. This keeps canonicals, robots, and sitemap apex-only no matter +// which hostname carried the request, while the redirect converges crawlers +// and visitors on one host. +const CANONICAL_HOST = "seofixkit.com"; +const CANONICAL_ORIGIN = `https://${CANONICAL_HOST}`; + +function canonicalOrigin(url) { + const hostname = url.hostname.toLowerCase(); + return hostname === CANONICAL_HOST || hostname === `www.${CANONICAL_HOST}` + ? CANONICAL_ORIGIN + : url.origin; +} + export default { async scheduled(_event, env, ctx) { if (env.WAITLIST_DB) { @@ -168,6 +184,18 @@ export default { async fetch(request, env, ctx) { const url = new URL(request.url); + // Canonical host: permanently redirect every www.seofixkit.com request + // onto the apex host with its path and query intact before any route + // logic runs, so no content or API response is ever served from www. + if (url.hostname.toLowerCase() === `www.${CANONICAL_HOST}`) { + return new Response(null, { + status: 301, + headers: secureHeaders({ Location: `${CANONICAL_ORIGIN}${url.pathname}${url.search}` }) + }); + } + + const origin = canonicalOrigin(url); + try { if (url.pathname === "/api/health") { return json({ @@ -554,7 +582,7 @@ export default { } if (url.pathname === "/fixture/rendered-page") { - return new Response(renderedFixture(url.origin), { + return new Response(renderedFixture(origin), { headers: secureHeaders({ "content-type": "text/html; charset=utf-8", "x-robots-tag": "noindex, nofollow" @@ -563,74 +591,74 @@ export default { } if (url.pathname === "/fixture/robots.txt") { - return new Response(`User-agent: *\nAllow: /\n\nSitemap: ${url.origin}/fixture/sitemap.xml\n`, { + return new Response(`User-agent: *\nAllow: /\n\nSitemap: ${origin}/fixture/sitemap.xml\n`, { headers: secureHeaders({ "content-type": "text/plain; charset=utf-8" }) }); } if (url.pathname === "/fixture/sitemap.xml") { return new Response( - `\n${url.origin}/fixture/rendered-page`, + `\n${origin}/fixture/rendered-page`, { headers: secureHeaders({ "content-type": "application/xml; charset=utf-8" }) } ); } if (url.pathname === "/robots.txt") { - return new Response(`User-agent: *\nAllow: /\n\nSitemap: ${url.origin}/sitemap.xml\n`, { + return new Response(`User-agent: *\nAllow: /\n\nSitemap: ${origin}/sitemap.xml\n`, { headers: secureHeaders({ "content-type": "text/plain; charset=utf-8" }) }); } if (url.pathname === "/sitemap.xml") { - return new Response(rootSitemap(url.origin), { + return new Response(rootSitemap(origin), { headers: secureHeaders({ "content-type": "application/xml; charset=utf-8" }) }); } if (url.pathname === "/llms.txt") { - return new Response(llmsText(url.origin), { + return new Response(llmsText(origin), { headers: secureHeaders({ "content-type": "text/plain; charset=utf-8" }) }); } if (url.pathname === "/privacy") { - return new Response(privacyHtml(url.origin), { + return new Response(privacyHtml(origin), { headers: secureHeaders({ "content-type": "text/html; charset=utf-8" }) }); } if (url.pathname === "/support") { - return new Response(supportHtml(url.origin), { + return new Response(supportHtml(origin), { headers: secureHeaders({ "content-type": "text/html; charset=utf-8" }) }); } if (url.pathname === "/terms") { - return new Response(termsHtml(url.origin), { + return new Response(termsHtml(origin), { headers: secureHeaders({ "content-type": "text/html; charset=utf-8" }) }); } if (url.pathname === "/demo") { - return new Response(demoHtml(url.origin), { + return new Response(demoHtml(origin), { headers: secureHeaders({ "content-type": "text/html; charset=utf-8" }) }); } if (url.pathname === "/check") { - return new Response(checkHtml(url.origin), { + return new Response(checkHtml(origin), { headers: secureHeaders({ "content-type": "text/html; charset=utf-8" }) }); } if (url.pathname === "/methodology") { - return new Response(methodologyHtml(url.origin), { + return new Response(methodologyHtml(origin), { headers: secureHeaders({ "content-type": "text/html; charset=utf-8" }) }); } if (url.pathname === "/packages") { - return new Response(packagesHtml(url.origin), { + return new Response(packagesHtml(origin), { headers: secureHeaders({ "content-type": "text/html; charset=utf-8" }) }); } @@ -645,7 +673,7 @@ export default { url.pathname === "/" && (request.headers.get("accept") || "").includes("text/markdown") ) { - return new Response(homeMarkdown(url.origin), { + return new Response(homeMarkdown(origin), { headers: secureHeaders({ "content-type": "text/markdown; charset=utf-8" }) }); } diff --git a/worker/index.test.mjs b/worker/index.test.mjs index 1896ace..db0913f 100644 --- a/worker/index.test.mjs +++ b/worker/index.test.mjs @@ -170,6 +170,50 @@ test("Worker dispatch routes public pages and repair APIs", async () => { assert.match(await apiProof.text(), /# SEOFixKit Repair Proof Receipt/); }); +test("Worker dispatch 301-redirects www.seofixkit.com onto the apex host", async () => { + const env = await fakeWorkerEnv(); + + const root = await worker.fetch(new Request("https://www.seofixkit.com/"), env, fakeCtx()); + assert.equal(root.status, 301); + assert.equal(root.headers.get("location"), "https://seofixkit.com/"); + assert.match(root.headers.get("strict-transport-security") || "", /max-age=31536000/); + + const deep = await worker.fetch(new Request("https://www.seofixkit.com/packages?utm_source=scout"), env, fakeCtx()); + assert.equal(deep.status, 301); + assert.equal(deep.headers.get("location"), "https://seofixkit.com/packages?utm_source=scout"); + + const api = await worker.fetch(new Request("https://www.seofixkit.com/api/health"), env, fakeCtx()); + assert.equal(api.status, 301); + assert.equal(api.headers.get("location"), "https://seofixkit.com/api/health"); + + const sitemap = await worker.fetch(new Request("https://www.seofixkit.com/sitemap.xml"), env, fakeCtx()); + assert.equal(sitemap.status, 301); + assert.equal(sitemap.headers.get("location"), "https://seofixkit.com/sitemap.xml"); +}); + +test("Worker dispatch serves apex-only canonicals, robots, sitemap, and llms.txt", async () => { + const env = await fakeWorkerEnv(); + + const sitemap = await worker.fetch(new Request("https://seofixkit.com/sitemap.xml"), env, fakeCtx()); + assert.equal(sitemap.status, 200); + const sitemapBody = await sitemap.text(); + assert.match(sitemapBody, /https:\/\/seofixkit\.com\/<\/loc>/); + assert.doesNotMatch(sitemapBody, /www\.seofixkit\.com/); + + const robots = await worker.fetch(new Request("https://seofixkit.com/robots.txt"), env, fakeCtx()); + assert.equal(robots.status, 200); + assert.match(await robots.text(), /Sitemap: https:\/\/seofixkit\.com\/sitemap\.xml/); + + const methodology = await worker.fetch(new Request("https://seofixkit.com/methodology"), env, fakeCtx()); + assert.equal(methodology.status, 200); + const methodologyBody = await methodology.text(); + assert.match(methodologyBody, /rel="canonical" href="https:\/\/seofixkit\.com\/methodology"/); + assert.doesNotMatch(methodologyBody, /www\.seofixkit\.com/); + + const llms = await worker.fetch(new Request("https://seofixkit.com/llms.txt"), env, fakeCtx()); + assert.match(await llms.text(), /https:\/\/seofixkit\.com\/check/); +}); + test("Worker dispatch exposes public-safe deep health readiness", async () => { const env = await fakeWorkerEnv(); const response = await worker.fetch(new Request("https://seofixkit.test/api/deep-health"), env, fakeCtx());