Drive Sitemap Desk from your own code
A plan, not a change. Nothing on your site is touched: the reply is a structure
plan and a redirect map for you to review and apply yourself. Your URL list travels in
inventory (up to 300 paths are sent; larger sites are sampled so every section is
represented), so send only paths you are allowed to share.
Everything the web page does is available over HTTP. Send the page inventory of an existing site
(root-relative paths, as they appear in its sitemap.xml, a crawl export or a URL list) with a little
business context, and get the same architecture plan back: a status (sound,
tidy_up, restructure), a headline and summary, findings with fixes, the
proposed page hierarchy (with /* collection nodes for large sections), the 301
redirect map from old URLs to new ones (wildcard rules included), the header, CTA, footer and
breadcrumb navigation, the internal-linking plan, next steps, and one answer per browser flag. The
natural use is a migration or redesign routine: export the sitemap, run it through here, and review
the redirect map before anything moves.
The model does not do the reading alone. Hosts and protocols that disagree, the
same page under several spellings, a mixed trailing-slash policy, URL-convention problems, depth
against the site type, folders without a hub page, synonym top-level sections, orphans and missing
key pages are all worked out first by the page's free prescan in sitekit.js and sent
as facts, a JSON string. See the facts string.
One lane: the task field
Every request names its lane in task. Sitemap Desk has exactly one.
| task | what it does | what it needs |
|---|---|---|
architect | The structure audit and architecture plan (site-architecture): site, status, headline, summary, findings, hierarchy, redirects, navigation, linking, next_steps and one prescan_responses entry per browser flag. | site_name, inventory (the URL paths) and facts (the prescan, as a JSON string). |
Any other value, or a missing task, is still answered as architect and
the reply's lane is "architect". Always send
"task": "architect"; the page's own guard refuses a body without a task.
Input fields
The body is one flat JSON object, built in the page by SiteKit.buildInput. Every value
is a string. Required: task, site_name, inventory,
facts.
| field | type | required | what it holds |
|---|---|---|---|
task | string | yes | Always "architect". |
site_name | string | yes | The site's name; the page falls back to the host, then "Untitled site". Cut at 160 characters. |
site_type | string | no | saas, content, ecommerce, docs, hybrid or local (the page defaults to saas). Sets the typical maximum depth: saas 3, content 3, local 2, ecommerce 4, docs 4, hybrid 4. |
goals | string | no | What the site is for (sign-ups, calls, sales). Cut at 2,000 characters. |
audiences | string | no | Who it serves. Cut at 2,000 characters. |
key_pages | string | no | The paths that matter most, newline-separated ("/pricing\n/signup"); up to 20. |
current_nav | string | no | The current header and footer items, free text. Cut at 2,000 characters. |
inventory | string | yes | The site's URLs as root-relative paths, one per line. A path may be followed by a TAB and inlinks=N (internal links pointing at it, from a crawl) and/or a TAB and status=NNN (only when not 200): "/pricing\tinlinks=41\n/old-page\tinlinks=0\tstatus=404". At most 300 lines; for a larger site send a sample and say so in facts.clipped. |
facts | string | yes | A JSON string (the output of JSON.stringify), never an object. In the page it is the browser's free prescan; its keys are listed below. |
question | string | no | Your own question, answered inside summary. Cut at 4,000 characters; the page sends "" when empty. |
retry_note | string | no | Only on a reformat retry, after a reply that could not be parsed: say what was wrong. Never on a first run. |
The facts string
In the web page, facts is computed by the browser before you pay for anything: the
free prescan reads your sitemap, crawl export or URL list, runs every structural check, raises the
flags you see on the page, and serializes the result. An API caller has two options: build the same
object yourself with the keys below (easiest by loading sitekit.js, which runs
unchanged in Node, see building the body), or send a minimal one. The model
parses it either way; a minimal facts simply means there are no browser flags to
confirm or dismiss, so prescan_responses comes back empty and the structural checks
rest on the model's own reading of inventory.
| key | what it holds |
|---|---|
host, hosts | The main host ("" for a list of bare paths), and every host seen, most frequent first. |
format | How the inventory was read: urlset (sitemap.xml), index (a sitemap index), csv (crawl export), list or empty. |
url_count, sent_count, clipped | URLs found; URLs listed in inventory; and "", or a sentence saying the inventory is a sample of how many distinct URLs. |
slash_policy | trailing or none: the majority trailing-slash style. |
max_depth, depth_counts, typical_max_depth | The deepest path; pages per depth ({"0":1,"1":7,"2":6}); the usual maximum for site_type. |
sections | Up to 40 top-level folders: path, pages, hub (whether the folder has its own page), children. |
folders_without_hub | Folders that hold pages but have no page of their own (up to 20). |
has_inlinks, orphans | Whether the export carried an Inlinks column, and up to 40 pages with 0 inlinks. |
key_pages | Each key page: path, found, depth and, when known, inlinks. |
expected_missing | Pages a site of this type usually has that were not found (/pricing, /contact...). |
utility_urls | Present only when found: up to 40 search, cart, account or login URLs, which are not content. |
browser_status | The prescan's status hint: restructure (any high flag), tidy_up (any medium), else sound. |
flags | Every flag: id (F1..), severity (high, medium, low), category (hosts, sitemap, status, duplicates, conventions, coverage, depth, hierarchy, linking, key_pages), message and up to five examples. |
A minimal facts, before JSON.stringify, for a caller that runs no prescan:
{"host": "example.com", "format": "list", "url_count": 14, "sent_count": 14, "clipped": "",
"slash_policy": "none", "typical_max_depth": 2, "browser_status": "sound", "flags": []}
Building the body
The surest way to match the page is to run the page's own engine. Load sitekit.js
(it runs unchanged in Node via require()) and give analyze the same fields
the page's form has; buildInput then samples the inventory, clips the text fields and
serializes the facts:
| set field | what it holds |
|---|---|
site | The site's name. |
type | saas, content, ecommerce, docs, hybrid or local. |
goals, audiences | Free text. |
key_pages | Paths, one per line (commas also work). |
nav | The current header and footer, free text. |
inventory | The raw export: a sitemap.xml, a crawl .csv (Address, Status Code and Inlinks columns are recognised) or one URL per line. |
// make-body.js - build the run body with the SAME engine the web page uses.
// Save sitekit.js from https://sitemap-desk.skillsafe.ai/sitekit.js next to this file.
// Usage: node make-body.js sitemap.xml
const fs = require("fs");
const K = require("./sitekit.js");
const A = K.analyze({
site: "Brindlecott Plumbing",
type: "local", // saas, content, ecommerce, docs, hybrid or local
goals: "Phone calls and booking requests from homeowners in the service area.",
audiences: "Homeowners in Eugene and Springfield with an urgent or planned plumbing job.",
key_pages: "/services\n/book",
nav: "Header: Services, Service areas, About, Book a visit. Footer: Contact, Privacy.",
inventory: fs.readFileSync(process.argv[2] || "sitemap.xml", "utf8") // sitemap.xml, crawl CSV or URL list
});
const body = K.buildInput(A, { question: "" }); // a non-empty question is answered in summary
fs.writeFileSync("body.json", JSON.stringify(K.mustBeObject(body)));
console.log(A.flags.length, "flags; browser status", A.hint, "; sent", A.sent.paths.length, "of", A.sent.total);
console.log("Idempotency-Key: sitemap-desk:architect:" + K.hashInput(body) + ":a1");
Run on the Brindlecott sitemap from the page's examples, this produces the worked request
below (its hash is 212ua9axq9u). mustBeObject is the page's own guard:
it throws unless the body is a plain object. From another language, send the same field names and
build facts with the keys above.
Base URL and the envelope
Every endpoint lives under https://api.skillsafe.ai/v1/app-api and every response uses
the same envelope, so one helper covers the whole API:
{"ok": true, "data": {"job_id": "job_...", "status": "queued"}}
{"ok": false, "error": {"code": "payment_required", "message": "..."}}
The token is minted for this app (the guest endpoint takes {"slug":"sitemap-desk"} in
its body), so no slug header is needed afterwards. Send it as Authorization: Bearer ….
The input object IS the request body. There is no {"input": …}
wrapper. A wrapped body is answered with an unknown field 'input' warning, and the
model never sees your inventory.
Error codes
| status | code | what to do |
|---|---|---|
| 400 | validation_error | A field is missing or the wrong type. Every field is a string: facts must be a JSON-encoded string, not an object. |
| 401 | unauthorized | The token is missing, malformed or expired. Get a new one from the token page. |
| 402 | payment_required | The balance is below min_credits. Call /estimate first and top up. |
| 403 | forbidden | The token is valid but not for this app, or a guest token tried a metered run. A guest cannot run; sign in for a personal token. |
| 404 | not_found | Unknown job id, or the app slug does not exist. |
| 409 | conflict | The same Idempotency-Key was replayed with a different body. Change the key or send the original input. |
| 429 | rate_limited | Too many requests. Back off and retry; do not tight-loop. |
| 5xx | internal | A server-side failure. Retry with the SAME Idempotency-Key so you are not billed twice. |
1. A tiny client
One helper that sends the token, unwraps data and raises on ok: false.
The token comes from the token page (Copy token or
Copy shell export); step 2 covers the kinds of token and minting one from code.
# Every call is the same three things: the base URL, your bearer token,
# and a JSON body. Keep the token in a shell variable.
BASE="https://api.skillsafe.ai/v1/app-api"
SLUG="sitemap-desk"
TOKEN="$SKILLSAFE_TOKEN" # from https://sitemap-desk.skillsafe.ai/tokens.html
call() { # call <path> [json-body]
if [ -n "$2" ]; then
curl -sS -X POST "$BASE/$1" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d "$2"
else
curl -sS "$BASE/$1" -H "Authorization: Bearer $TOKEN"
fi
}
import json, os, urllib.error, urllib.request
BASE = "https://api.skillsafe.ai/v1/app-api"
SLUG = "sitemap-desk"
TOKEN = os.environ.get("SKILLSAFE_TOKEN", "YOUR_TOKEN") # from https://sitemap-desk.skillsafe.ai/tokens.html
def call(path, body=None):
"""Returns the unwrapped `data`, or raises with the API error code."""
data = json.dumps(body).encode() if body is not None else None
req = urllib.request.Request(f"{BASE}/{path}", data=data, method="POST" if body is not None else "GET")
req.add_header("Authorization", f"Bearer {TOKEN}")
if body is not None:
req.add_header("Content-Type", "application/json")
try:
with urllib.request.urlopen(req) as r:
payload = json.load(r)
except urllib.error.HTTPError as e:
payload = json.load(e)
if not payload.get("ok"):
err = payload.get("error", {})
raise RuntimeError(f"{err.get('code')}: {err.get('message')}")
return payload["data"]
import { readFileSync } from "node:fs";
const BASE = "https://api.skillsafe.ai/v1/app-api";
const SLUG = "sitemap-desk";
// Paste the token from https://sitemap-desk.skillsafe.ai/tokens.html into a file named "token",
// or replace the fallback with it.
let TOKEN = "YOUR_TOKEN";
try { TOKEN = readFileSync("token", "utf8").trim(); } catch {}
async function call(path, body) {
const res = await fetch(`${BASE}/${path}`, {
method: body ? "POST" : "GET",
headers: {
Authorization: `Bearer ${TOKEN}`,
...(body ? { "Content-Type": "application/json" } : {}),
},
body: body ? JSON.stringify(body) : undefined,
});
const payload = await res.json();
if (!payload.ok) throw new Error(`${payload.error.code}: ${payload.error.message}`);
return payload.data;
}
package main
import (
"bufio"
"bytes"
"crypto/sha256"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strings"
"time"
)
const (
base = "https://api.skillsafe.ai/v1/app-api"
slug = "sitemap-desk"
)
var token = os.Getenv("SKILLSAFE_TOKEN") // from https://sitemap-desk.skillsafe.ai/tokens.html
type envelope struct {
OK bool `json:"ok"`
Data json.RawMessage `json:"data"`
Error struct {
Code string `json:"code"`
Message string `json:"message"`
} `json:"error"`
}
func call(path string, body any) (json.RawMessage, error) {
method := http.MethodGet
var rdr io.Reader
if body != nil {
method = http.MethodPost
b, _ := json.Marshal(body)
rdr = bytes.NewReader(b)
}
req, _ := http.NewRequest(method, base+"/"+path, rdr)
req.Header.Set("Authorization", "Bearer "+token)
if body != nil {
req.Header.Set("Content-Type", "application/json")
}
res, err := http.DefaultClient.Do(req)
if err != nil {
return nil, err
}
defer res.Body.Close()
var env envelope
if err := json.NewDecoder(res.Body).Decode(&env); err != nil {
return nil, err
}
if !env.OK {
return nil, fmt.Errorf("%s: %s", env.Error.Code, env.Error.Message)
}
return env.Data, nil
}
import java.net.URI;
import java.net.http.*;
public class SitemapDesk {
static final String BASE = "https://api.skillsafe.ai/v1/app-api";
static final String SLUG = "sitemap-desk";
static final String TOKEN = System.getenv().getOrDefault("SKILLSAFE_TOKEN", "YOUR_TOKEN");
static final HttpClient HTTP = HttpClient.newHttpClient();
static String call(String path, String jsonBody) throws Exception {
HttpRequest.Builder b = HttpRequest.newBuilder(URI.create(BASE + "/" + path))
.header("Authorization", "Bearer " + TOKEN);
if (jsonBody != null) {
b.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(jsonBody));
} else {
b.GET();
}
HttpResponse<String> res = HTTP.send(b.build(), HttpResponse.BodyHandlers.ofString());
// The envelope is always {"ok":true,"data":...} or {"ok":false,"error":...}.
return res.body();
}
}
require "json"
require "net/http"
require "uri"
BASE = "https://api.skillsafe.ai/v1/app-api"
SLUG = "sitemap-desk"
TOKEN = ENV.fetch("SKILLSAFE_TOKEN", "YOUR_TOKEN") # from https://sitemap-desk.skillsafe.ai/tokens.html
def call(path, body = nil)
uri = URI("#{BASE}/#{path}")
req = body ? Net::HTTP::Post.new(uri) : Net::HTTP::Get.new(uri)
req["Authorization"] = "Bearer #{TOKEN}"
if body
req["Content-Type"] = "application/json"
req.body = JSON.generate(body)
end
res = Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) { |h| h.request(req) }
payload = JSON.parse(res.body)
raise "#{payload['error']['code']}: #{payload['error']['message']}" unless payload["ok"]
payload["data"]
end
<?php
const BASE = "https://api.skillsafe.ai/v1/app-api";
const SLUG = "sitemap-desk";
define("TOKEN", getenv("SKILLSAFE_TOKEN") ?: "YOUR_TOKEN"); // from /tokens.html
function call(string $path, ?array $body = null) {
$ch = curl_init(BASE . "/" . $path);
$headers = ["Authorization: Bearer " . TOKEN];
if ($body !== null) {
$headers[] = "Content-Type: application/json";
curl_setopt($ch, CURLOPT_POST, true);
curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode($body));
}
curl_setopt($ch, CURLOPT_HTTPHEADER, $headers);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
$payload = json_decode(curl_exec($ch), true);
curl_close($ch);
if (empty($payload["ok"])) {
throw new RuntimeException($payload["error"]["code"] . ": " . $payload["error"]["message"]);
}
return $payload["data"];
}
using System.Net.Http.Json;
using System.Text.Json;
static class SitemapDesk
{
const string Base = "https://api.skillsafe.ai/v1/app-api";
const string Slug = "sitemap-desk";
static readonly string Token =
Environment.GetEnvironmentVariable("SKILLSAFE_TOKEN") ?? "YOUR_TOKEN";
static readonly HttpClient Http = new();
public static async Task<JsonElement> Call(string path, object? body = null)
{
var req = new HttpRequestMessage(body is null ? HttpMethod.Get : HttpMethod.Post, $"{Base}/{path}");
req.Headers.Add("Authorization", $"Bearer {Token}");
if (body is not null) req.Content = JsonContent.Create(body);
var res = await Http.SendAsync(req);
var payload = await res.Content.ReadFromJsonAsync<JsonElement>();
if (!payload.GetProperty("ok").GetBoolean())
{
var e = payload.GetProperty("error");
throw new Exception($"{e.GetProperty("code")}: {e.GetProperty("message")}");
}
return payload.GetProperty("data");
}
}
2. Get a token
The easiest route is the token page: it shows the token this browser
already holds, with Copy token and Copy shell export buttons, and
a sign-in button for a personal token. A guest token, minted with
POST /guest and {"slug":"sitemap-desk"}, can call /me and
/estimate; the run is metered, so /run and /run-stream need
a personal token.
# The token page is the shortest path. It shows the token this browser holds and
# hands you a ready-made shell export:
#
# https://sitemap-desk.skillsafe.ai/tokens.html
# export SKILLSAFE_TOKEN="..."
#
# To mint a guest token from the command line instead. A guest token is enough
# for /me and /estimate; a run needs a personal token from signing in.
curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/guest" \
-H "Content-Type: application/json" -d '{"slug":"sitemap-desk"}'
# {"ok":true,"data":{"token":"…","subject_type":"guest"}}
# Open https://sitemap-desk.skillsafe.ai/tokens.html and press "Copy token",
# or mint a guest token here. A guest token can call /me and /estimate but
# cannot start a metered run.
import json, urllib.request
req = urllib.request.Request(
"https://api.skillsafe.ai/v1/app-api/guest", data=b'{"slug": "sitemap-desk"}', method="POST")
req.add_header("Content-Type", "application/json")
with urllib.request.urlopen(req) as r:
TOKEN = json.load(r)["data"]["token"]
// Open https://sitemap-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot start a metered run.
const res = await fetch("https://api.skillsafe.ai/v1/app-api/guest", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ slug: "sitemap-desk" }),
});
const TOKEN = (await res.json()).data.token;
// Open https://sitemap-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot start a metered run.
guestReq, _ := http.NewRequest(http.MethodPost,
"https://api.skillsafe.ai/v1/app-api/guest", bytes.NewReader([]byte(`{"slug":"sitemap-desk"}`)))
guestReq.Header.Set("Content-Type", "application/json")
guestRes, err := http.DefaultClient.Do(guestReq)
if err != nil {
panic(err)
}
defer guestRes.Body.Close()
var guest struct {
Data struct {
Token string `json:"token"`
} `json:"data"`
}
_ = json.NewDecoder(guestRes.Body).Decode(&guest)
fmt.Println(guest.Data.Token)
// Open https://sitemap-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot start a metered run.
var http = HttpClient.newHttpClient();
var guestReq = HttpRequest.newBuilder(URI.create("https://api.skillsafe.ai/v1/app-api/guest"))
.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString("{\"slug\":\"sitemap-desk\"}"))
.build();
HttpResponse<String> guest = http.send(guestReq, HttpResponse.BodyHandlers.ofString());
System.out.println(guest.body()); // {"ok":true,"data":{"token":"…","subject_type":"guest"}}
# Open https://sitemap-desk.skillsafe.ai/tokens.html and press "Copy token",
# or mint a guest token here. A guest token can call /me and /estimate but
# cannot start a metered run.
require "json"
require "net/http"
require "uri"
uri = URI("https://api.skillsafe.ai/v1/app-api/guest")
req = Net::HTTP::Post.new(uri)
req["Content-Type"] = "application/json"
req.body = JSON.generate({ slug: "sitemap-desk" })
res = Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) { |h| h.request(req) }
TOKEN = JSON.parse(res.body)["data"]["token"]
<?php
// Open https://sitemap-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot start a metered run.
$ch = curl_init("https://api.skillsafe.ai/v1/app-api/guest");
curl_setopt($ch, CURLOPT_POST, true);
curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode(["slug" => "sitemap-desk"]));
curl_setopt($ch, CURLOPT_HTTPHEADER, ["Content-Type: application/json"]);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
$guest = json_decode(curl_exec($ch), true);
curl_close($ch);
echo $guest["data"]["token"];
// Open https://sitemap-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot start a metered run.
using var http = new HttpClient();
var guestReq = new HttpRequestMessage(HttpMethod.Post, "https://api.skillsafe.ai/v1/app-api/guest");
guestReq.Content = new StringContent("{\"slug\":\"sitemap-desk\"}", Encoding.UTF8, "application/json");
var guestRes = await http.SendAsync(guestReq);
var guest = await guestRes.Content.ReadFromJsonAsync<JsonElement>();
Console.WriteLine(guest.GetProperty("data").GetProperty("token").GetString());
3. Check the session and the balance
call me
# {"ok":true,"data":{"subject_type":"user","username":"you","credits":51234}}
me = call("me")
print(me["subject_type"], me.get("credits"))
const me = await call("me");
console.log(me.subject_type, me.credits);
raw, err := call("me", nil)
if err != nil {
panic(err)
}
var me struct {
SubjectType string `json:"subject_type"`
Credits int `json:"credits"`
}
_ = json.Unmarshal(raw, &me)
fmt.Println(me.SubjectType, me.Credits)
System.out.println(call("me", null));
// {"ok":true,"data":{"subject_type":"user","username":"you","credits":51234}}
me = call("me")
puts "#{me['subject_type']} #{me['credits']}"
<?php
$me = call("me");
echo $me["subject_type"], " ", $me["credits"], PHP_EOL;
var me = await SitemapDesk.Call("me");
Console.WriteLine(me.GetProperty("subject_type").GetString());
4. Price the run (free)
/estimate returns the model binding and the credits a run would reserve. It creates no
job and charges nothing. Expect model_alias gpt-terra and
markup_bps 1000 (a 10% markup). hold_credits is what the
run reserves, not the price; min_credits is the least balance that
can start it; the charged_credits reported after the run is usually far lower than the
hold. The body is the input object itself, with no {"input": …} wrapper.
/estimate does not validate the body, so check the shape yourself: an object whose
every value is a string, task equal to architect,
site_name, inventory and facts non-empty, and
facts a JSON string that parses to an object.
# body.json is the input object itself - no {"input": ...} wrapper. Build it with
# make-body.js above, or by hand. estimate does not validate it, so check the shape first:
python3 -c 'import json;b=json.load(open("body.json"));assert isinstance(b,dict) and b.get("task")=="architect" and all(isinstance(v,str) for v in b.values()) and all(b.get(k,"").strip() for k in ("site_name","inventory","facts")) and isinstance(json.loads(b["facts"]),dict)'
INPUT=$(cat body.json)
call estimate "$INPUT"
# {"ok":true,"data":{"model":"...","model_alias":"gpt-terra",
# "markup_bps":1000,"hold_credits":...,"min_credits":...,"sponsor_enabled":false,
# "warnings":[]}}
#
# estimate creates no job and charges nothing. hold_credits is RESERVED, not the
# price; charged_credits after the run is usually far lower.
INPUT = json.load(open("body.json")) # built by make-body.js above, or by hand
assert isinstance(INPUT, dict) and INPUT.get("task") == "architect"
assert all(isinstance(v, str) for v in INPUT.values())
assert all(INPUT.get(k, "").strip() for k in ("site_name", "inventory", "facts"))
assert isinstance(json.loads(INPUT["facts"]), dict) # facts is a JSON STRING
est = call("estimate", INPUT)
print(est["model_alias"], est["markup_bps"], est["hold_credits"], est.get("warnings"))
me = call("me")
if me.get("credits", 0) < est["min_credits"]:
raise SystemExit("top up first: balance is below min_credits")
const INPUT = JSON.parse(readFileSync("body.json", "utf8")); // built by make-body.js above
if (!INPUT || typeof INPUT !== "object" || INPUT.task !== "architect") throw new Error("task must be architect");
for (const [k, v] of Object.entries(INPUT)) if (typeof v !== "string") throw new Error(k + " must be a string");
for (const k of ["site_name", "inventory", "facts"]) if (!INPUT[k]) throw new Error(k + " is required");
JSON.parse(INPUT.facts); // throws unless facts is a JSON string
const est = await call("estimate", INPUT);
console.log(est.model_alias, est.markup_bps, est.hold_credits, est.warnings);
const me = await call("me");
if ((me.credits ?? 0) < est.min_credits) throw new Error("top up first");
raw, _ := os.ReadFile("body.json") // built by make-body.js above
var input map[string]string // every field is a string, facts included
if err := json.Unmarshal(raw, &input); err != nil {
panic("body.json must be an object of strings: " + err.Error())
}
if input["task"] != "architect" {
panic("task must be architect")
}
for _, k := range []string{"site_name", "inventory", "facts"} {
if strings.TrimSpace(input[k]) == "" {
panic(k + " is required")
}
}
var facts map[string]any
if err := json.Unmarshal([]byte(input["facts"]), &facts); err != nil {
panic("facts must be a JSON string holding an object")
}
est, err := call("estimate", input)
if err != nil {
panic(err)
}
fmt.Println(string(est)) // model_alias gpt-terra, markup_bps 1000, hold_credits, min_credits
String input = Files.readString(Path.of("body.json")); // built by make-body.js above
if (!input.matches("(?s)\\s*\\{.*\"task\"\\s*:\\s*\"architect\".*\\}\\s*"))
throw new IllegalStateException("body.json must be an object with task architect");
for (String k : new String[] {"site_name", "inventory", "facts"})
if (!input.contains("\"" + k + "\"")) throw new IllegalStateException(k + " is required");
String est = call("estimate", input);
System.out.println(est); // model_alias gpt-terra, markup_bps 1000, hold_credits, min_credits
INPUT = JSON.parse(File.read("body.json")) # built by make-body.js above
raise "task must be architect" unless INPUT["task"] == "architect"
INPUT.each { |k, v| raise "#{k} must be a string" unless v.is_a?(String) }
%w[site_name inventory facts].each { |k| raise "#{k} is required" if INPUT[k].to_s.strip.empty? }
raise "facts must hold an object" unless JSON.parse(INPUT["facts"]).is_a?(Hash)
est = call("estimate", INPUT)
puts est["model_alias"], est["markup_bps"], est["hold_credits"]
<?php
$input = json_decode(file_get_contents("body.json"), true); // built by make-body.js above
if (!is_array($input) || ($input["task"] ?? "") !== "architect") { throw new Exception("task must be architect"); }
foreach ($input as $k => $v) { if (!is_string($v)) { throw new Exception("$k must be a string"); } }
foreach (["site_name", "inventory", "facts"] as $k) { if (trim($input[$k] ?? "") === "") { throw new Exception("$k is required"); } }
if (!is_array(json_decode($input["facts"], true))) { throw new Exception("facts must be a JSON string"); }
$est = call("estimate", $input);
echo $est["model_alias"], " ", $est["markup_bps"], " ", $est["hold_credits"], PHP_EOL;
var input = File.ReadAllText("body.json"); // built by make-body.js above
using var doc = JsonDocument.Parse(input);
var root = doc.RootElement;
if (root.GetProperty("task").GetString() != "architect") throw new Exception("task must be architect");
foreach (var p in root.EnumerateObject())
if (p.Value.ValueKind != JsonValueKind.String) throw new Exception($"{p.Name} must be a string");
JsonDocument.Parse(root.GetProperty("facts").GetString()!); // facts is a JSON string
var est = await SitemapDesk.Call("estimate", root);
Console.WriteLine(est); // model_alias gpt-terra, markup_bps 1000, hold_credits, min_credits
5. Run it, then poll
POST /run returns a job_id; poll GET /jobs/{id} until it is
terminal. The reply is a string at data.output.output: JSON.parse
it (step 7). Send an Idempotency-Key built from the lane, a hash of the input and the
attempt number, sitemap-desk:architect:<hash>:a<attempt>, so a retried
request returns the same job instead of billing a second run. Use one key per distinct input: a
changed inventory or changed facts is a new hash, and replaying an old key with a different body is
a 409. The page uses SiteKit.hashInput(body) for the hash
(make-body.js prints that key); any stable digest of the body works from other
languages. Leave retry_note out of the hash and bump the attempt instead.
# Always send an Idempotency-Key derived from the input. A retried request with
# the same key returns the SAME job instead of billing a second run.
KEY="sitemap-desk:architect:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"
JOB=$(curl -sS -X POST "$BASE/run" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $KEY" \
-d "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')
while :; do
OUT=$(call "jobs/$JOB")
STATUS=$(printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
[ "$STATUS" = "succeeded" ] && break
[ "$STATUS" = "failed" ] && echo "$OUT" && exit 1
sleep 2
done
# {"ok":true,"data":{"job_id":"job_...","status":"succeeded",
# "output":{"output":"{\"lane\":\"architect\",\"site\":{...},\"status\":\"sound\", ...}"},
# "charged_credits":...,"truncated":false}}
printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["output"]["output"])' > reply.json
import hashlib, time
digest = hashlib.sha256(json.dumps(INPUT, sort_keys=True).encode()).hexdigest()[:16]
key = f"sitemap-desk:architect:{digest}:a1"
req = urllib.request.Request(f"{BASE}/run", data=json.dumps(INPUT).encode(), method="POST")
req.add_header("Authorization", f"Bearer {TOKEN}")
req.add_header("Content-Type", "application/json")
req.add_header("Idempotency-Key", key)
with urllib.request.urlopen(req) as r:
job_id = json.load(r)["data"]["job_id"]
while True:
job = call(f"jobs/{job_id}")
if job["status"] in ("succeeded", "failed"):
break
time.sleep(2)
if job["status"] == "failed":
raise RuntimeError(job.get("error"))
text = job["output"]["output"] # the reply, as a string
print("charged", job.get("charged_credits"), "truncated", job.get("truncated"))
import { createHash } from "node:crypto";
const digest = createHash("sha256").update(JSON.stringify(INPUT)).digest("hex").slice(0, 16);
const key = `sitemap-desk:architect:${digest}:a1`;
const started = await fetch(`${BASE}/run`, {
method: "POST",
headers: { Authorization: `Bearer ${TOKEN}`, "Content-Type": "application/json", "Idempotency-Key": key },
body: JSON.stringify(INPUT),
}).then((r) => r.json());
if (!started.ok) throw new Error(`${started.error.code}: ${started.error.message}`);
let job = started.data;
while (job.status !== "succeeded" && job.status !== "failed") {
await new Promise((r) => setTimeout(r, 2000));
job = await call(`jobs/${job.job_id}`);
}
if (job.status === "failed") throw new Error(JSON.stringify(job.error));
const text = job.output.output; // the reply, as a string
console.log(job.charged_credits, job.truncated);
body, _ := json.Marshal(input)
sum := sha256.Sum256(body)
key := fmt.Sprintf("sitemap-desk:architect:%x:a1", sum[:8])
req, _ := http.NewRequest(http.MethodPost, base+"/run", bytes.NewReader(body))
req.Header.Set("Authorization", "Bearer "+token)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", key)
res, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
var started struct {
Data struct {
JobID string `json:"job_id"`
} `json:"data"`
}
_ = json.NewDecoder(res.Body).Decode(&started)
res.Body.Close()
var jobOutput string
for {
raw, err := call("jobs/"+started.Data.JobID, nil)
if err != nil {
panic(err)
}
var job struct {
Status string `json:"status"`
Output struct {
Output string `json:"output"`
} `json:"output"`
Charged int `json:"charged_credits"`
Truncated bool `json:"truncated"`
}
_ = json.Unmarshal(raw, &job)
if job.Status == "succeeded" {
jobOutput = job.Output.Output
fmt.Println(job.Charged, job.Truncated)
break
}
if job.Status == "failed" {
panic(string(raw))
}
time.Sleep(2 * time.Second)
}
String key = "sitemap-desk:architect:" + sha256Hex(input).substring(0, 16) + ":a1";
HttpRequest run = HttpRequest.newBuilder(URI.create(BASE + "/run"))
.header("Authorization", "Bearer " + TOKEN)
.header("Content-Type", "application/json")
.header("Idempotency-Key", key)
.POST(HttpRequest.BodyPublishers.ofString(input)).build();
String started = HTTP.send(run, HttpResponse.BodyHandlers.ofString()).body();
String jobId = started.replaceAll(".*\"job_id\":\"([^\"]+)\".*", "$1");
while (true) {
String job = call("jobs/" + jobId, null);
if (job.contains("\"status\":\"succeeded\"")) { System.out.println(job); break; }
if (job.contains("\"status\":\"failed\"")) throw new RuntimeException(job);
Thread.sleep(2000);
}
// Parse data.output.output (a string holding the reply JSON) with your JSON library.
// sha256Hex: HexFormat.of().formatHex(MessageDigest.getInstance("SHA-256").digest(input.getBytes(UTF_8)))
require "digest"
key = "sitemap-desk:architect:#{Digest::SHA256.hexdigest(JSON.generate(INPUT))[0, 16]}:a1"
uri = URI("#{BASE}/run")
req = Net::HTTP::Post.new(uri)
req["Authorization"] = "Bearer #{TOKEN}"
req["Content-Type"] = "application/json"
req["Idempotency-Key"] = key
req.body = JSON.generate(INPUT)
job = JSON.parse(Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) { |h| h.request(req) }.body)["data"]
until %w[succeeded failed].include?(job["status"])
sleep 2
job = call("jobs/#{job['job_id']}")
end
raise job.inspect if job["status"] == "failed"
text = job["output"]["output"] # the reply, as a string
puts job["charged_credits"], job["truncated"]
<?php
$key = "sitemap-desk:architect:" . substr(hash("sha256", json_encode($input)), 0, 16) . ":a1";
$ch = curl_init(BASE . "/run");
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_POSTFIELDS => json_encode($input),
CURLOPT_HTTPHEADER => ["Authorization: Bearer " . TOKEN, "Content-Type: application/json", "Idempotency-Key: " . $key],
CURLOPT_RETURNTRANSFER => true,
]);
$job = json_decode(curl_exec($ch), true)["data"];
curl_close($ch);
while (!in_array($job["status"], ["succeeded", "failed"], true)) {
sleep(2);
$job = call("jobs/" . $job["job_id"]);
}
if ($job["status"] === "failed") { throw new RuntimeException(json_encode($job)); }
$text = $job["output"]["output"]; // the reply, as a string
echo $job["charged_credits"], PHP_EOL;
using System.Security.Cryptography;
var json = input; // the body.json text from step 4
var key = "sitemap-desk:architect:" + Convert.ToHexString(SHA256.HashData(System.Text.Encoding.UTF8.GetBytes(json)))[..16].ToLower() + ":a1";
var req = new HttpRequestMessage(HttpMethod.Post, "https://api.skillsafe.ai/v1/app-api/run");
req.Headers.Add("Authorization", $"Bearer {Environment.GetEnvironmentVariable("SKILLSAFE_TOKEN") ?? "YOUR_TOKEN"}");
req.Headers.Add("Idempotency-Key", key);
req.Content = new StringContent(json, System.Text.Encoding.UTF8, "application/json");
var started = await (await new HttpClient().SendAsync(req)).Content.ReadFromJsonAsync<JsonElement>();
var jobId = started.GetProperty("data").GetProperty("job_id").GetString();
JsonElement job;
while (true)
{
job = await SitemapDesk.Call($"jobs/{jobId}");
var status = job.GetProperty("status").GetString();
if (status == "succeeded") break;
if (status == "failed") throw new Exception(job.ToString());
await Task.Delay(2000);
}
var output = job.GetProperty("output").GetProperty("output").GetString()!; // the reply, as a string
6. Or stream it
POST /run-stream takes the same body and headers and answers with server-sent events:
job (the job id), delta (chunks of the reply) and done (the
status, charged_credits, truncated and, when present, the full
output). A browser page may receive only tick heartbeats and then
done, never a delta, so take the reply from done.output.output
when it is there, fall back to the concatenated deltas, and fall back again to
GET /jobs/{id}.
# Server-sent events. `delta` events carry chunks of the reply; `done` carries the
# status, charged_credits and the truncated flag. Ignore `tick` heartbeats.
curl -N -X POST "$BASE/run-stream" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $KEY" \
-H "Accept: text/event-stream" \
-d "$INPUT"
# event: job {"job_id":"job_..."}
# event: delta {"text":"{\"lane\":\"architect\",\"site\":{\"name\":\"Brindlecott Plumbing\""}
# event: done {"status":"succeeded","charged_credits":...,"truncated":false}
req = urllib.request.Request(f"{BASE}/run-stream", data=json.dumps(INPUT).encode(), method="POST")
for h, v in (("Authorization", f"Bearer {TOKEN}"), ("Content-Type", "application/json"),
("Idempotency-Key", key), ("Accept", "text/event-stream")):
req.add_header(h, v)
raw, done, event = "", {}, None
with urllib.request.urlopen(req) as stream:
for line in stream:
line = line.decode().rstrip("\n")
if line.startswith("event: "):
event = line[7:]
elif line.startswith("data: ") and event == "delta":
raw += json.loads(line[6:]).get("text", "")
elif line.startswith("data: ") and event == "done":
done = json.loads(line[6:])
text = (done.get("output") or {}).get("output") or raw
print(done.get("status"), done.get("charged_credits"), done.get("truncated"))
const res = await fetch(`${BASE}/run-stream`, {
method: "POST",
headers: { Authorization: `Bearer ${TOKEN}`, "Content-Type": "application/json", "Idempotency-Key": key, Accept: "text/event-stream" },
body: JSON.stringify(INPUT),
});
const reader = res.body.getReader();
const dec = new TextDecoder();
let buf = "", raw = "", event = null, done = null;
for (;;) {
const { value, done: end } = await reader.read();
if (end) break;
buf += dec.decode(value, { stream: true });
let i;
while ((i = buf.indexOf("\n")) >= 0) {
const line = buf.slice(0, i); buf = buf.slice(i + 1);
if (line.startsWith("event: ")) event = line.slice(7);
else if (line.startsWith("data: ") && event === "delta") raw += JSON.parse(line.slice(6)).text || "";
else if (line.startsWith("data: ") && event === "done") done = JSON.parse(line.slice(6));
}
}
const streamed = done?.output?.output || raw; // browsers may get only ticks + done
console.log(done, streamed.length);
req, _ = http.NewRequest(http.MethodPost, base+"/run-stream", bytes.NewReader(body))
req.Header.Set("Authorization", "Bearer "+token)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", key)
req.Header.Set("Accept", "text/event-stream")
res, err = http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer res.Body.Close()
var raw strings.Builder
event := ""
sc := bufio.NewScanner(res.Body)
sc.Buffer(make([]byte, 1<<20), 1<<20)
for sc.Scan() {
line := sc.Text()
switch {
case strings.HasPrefix(line, "event: "):
event = line[7:]
case strings.HasPrefix(line, "data: ") && event == "delta":
var d struct{ Text string `json:"text"` }
_ = json.Unmarshal([]byte(line[6:]), &d)
raw.WriteString(d.Text)
case strings.HasPrefix(line, "data: ") && event == "done":
fmt.Println("done:", line[6:])
}
}
HttpRequest stream = HttpRequest.newBuilder(URI.create(BASE + "/run-stream"))
.header("Authorization", "Bearer " + TOKEN)
.header("Content-Type", "application/json")
.header("Idempotency-Key", key)
.header("Accept", "text/event-stream")
.POST(HttpRequest.BodyPublishers.ofString(input)).build();
HTTP.send(stream, HttpResponse.BodyHandlers.ofLines()).body().forEach(line -> {
// "event: delta" lines are followed by "data: {\"text\":...}"; "event: done" by the status.
if (line.startsWith("data: ")) System.out.println(line.substring(6));
});
uri = URI("#{BASE}/run-stream")
req = Net::HTTP::Post.new(uri)
{ "Authorization" => "Bearer #{TOKEN}", "Content-Type" => "application/json",
"Idempotency-Key" => key, "Accept" => "text/event-stream" }.each { |k, v| req[k] = v }
req.body = JSON.generate(INPUT)
raw, event = +"", nil
Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) do |h|
h.request(req) do |res|
res.read_body do |chunk|
chunk.each_line do |line|
line = line.chomp
if line.start_with?("event: ") then event = line[7..]
elsif line.start_with?("data: ") && event == "delta" then raw << JSON.parse(line[6..])["text"].to_s
elsif line.start_with?("data: ") && event == "done" then puts line[6..]
end
end
end
end
end
<?php
$raw = ""; $event = null;
$ch = curl_init(BASE . "/run-stream");
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_POSTFIELDS => json_encode($input),
CURLOPT_HTTPHEADER => ["Authorization: Bearer " . TOKEN, "Content-Type: application/json", "Idempotency-Key: " . $key, "Accept: text/event-stream"],
CURLOPT_WRITEFUNCTION => function ($ch, $chunk) use (&$raw, &$event) {
foreach (explode("\n", $chunk) as $line) {
if (str_starts_with($line, "event: ")) $event = substr($line, 7);
elseif (str_starts_with($line, "data: ") && $event === "delta") $raw .= json_decode(substr($line, 6), true)["text"] ?? "";
elseif (str_starts_with($line, "data: ") && $event === "done") echo substr($line, 6), PHP_EOL;
}
return strlen($chunk);
},
]);
curl_exec($ch);
curl_close($ch);
var sreq = new HttpRequestMessage(HttpMethod.Post, "https://api.skillsafe.ai/v1/app-api/run-stream");
sreq.Headers.Add("Authorization", $"Bearer {Environment.GetEnvironmentVariable("SKILLSAFE_TOKEN") ?? "YOUR_TOKEN"}");
sreq.Headers.Add("Idempotency-Key", key);
sreq.Headers.Add("Accept", "text/event-stream");
sreq.Content = new StringContent(json, System.Text.Encoding.UTF8, "application/json");
using var sres = await new HttpClient().SendAsync(sreq, HttpCompletionOption.ResponseHeadersRead);
using var sr = new StreamReader(await sres.Content.ReadAsStreamAsync());
var raw = new System.Text.StringBuilder(); string? ev = null, line;
while ((line = await sr.ReadLineAsync()) != null)
{
if (line.StartsWith("event: ")) ev = line[7..];
else if (line.StartsWith("data: ") && ev == "delta") raw.Append(JsonSerializer.Deserialize<JsonElement>(line[6..]).GetProperty("text").GetString());
else if (line.StartsWith("data: ") && ev == "done") Console.WriteLine(line[6..]);
}
7. Parse the reply
The reply is one JSON object, delivered as a string in data.output.output; you must
JSON.parse it. The model is told to send no code fences, but tolerate them: strip a
leading ```json and a trailing ```, keep everything from the first
{ to the last }, and parse that outer object.
# reply.json holds data.output.output from step 5. Strip any fence, keep the object:
python3 - <<'EOF'
import json, re
t = open("reply.json").read().strip()
t = re.sub(r"^```(?:json)?\s*", "", t, flags=re.I)
t = re.sub(r"\s*```\s*$", "", t)
r = json.loads(t[t.index("{"):t.rindex("}") + 1])
print(r["status"], "-", r["headline"])
for f in r["findings"]:
print(f["ref"] or "-", f["severity"], "|", f["finding"], "->", f["fix"])
for h in r["hierarchy"]:
print(" " * h["level"] + h["path"], h["title"], "*new" if h["is_new"] else "", h["count"] or "")
for x in r["redirects"]:
print("301", x["from"], "->", x["to"])
print([l["label"] for l in r["navigation"]["header"]], "CTA:", r["navigation"]["cta"]["label"])
EOF
import re
t = re.sub(r"^```(?:json)?\s*", "", text.strip(), flags=re.I)
t = re.sub(r"\s*```\s*$", "", t)
reply = json.loads(t[t.index("{"):t.rindex("}") + 1])
print(reply["status"], reply["headline"])
print([(f["ref"], f["severity"]) for f in reply["findings"]])
print([(h["path"], h["level"], h["count"]) for h in reply["hierarchy"]])
print([(x["from"], x["to"]) for x in reply["redirects"]])
let t = text.trim().replace(/^```(?:json)?\s*/i, "").replace(/\s*```\s*$/, "");
const reply = JSON.parse(t.slice(t.indexOf("{"), t.lastIndexOf("}") + 1));
console.log(reply.status, reply.headline);
console.log(reply.findings.map((f) => [f.ref, f.severity]));
console.log(reply.hierarchy.map((h) => [h.path, h.level, h.count]));
console.log(reply.redirects.map((x) => `${x.from} -> ${x.to}`));
t := strings.TrimSpace(jobOutput)
t = strings.TrimPrefix(strings.TrimPrefix(t, "```json"), "```")
t = strings.TrimSuffix(strings.TrimSpace(t), "```")
start, end := strings.Index(t, "{"), strings.LastIndex(t, "}")
var reply struct {
Status string `json:"status"`
Headline string `json:"headline"`
Findings []struct {
Ref, Severity, Finding, Fix string
} `json:"findings"`
Hierarchy []struct {
Path, Title, Parent, Nav, Priority string
Level int
IsNew bool `json:"is_new"`
Count *int `json:"count"`
} `json:"hierarchy"`
Redirects []struct {
From, To, Reason string
} `json:"redirects"`
}
if err := json.Unmarshal([]byte(t[start:end+1]), &reply); err != nil {
panic(err)
}
fmt.Println(reply.Status, reply.Headline)
for _, h := range reply.Hierarchy {
fmt.Println(h.Level, h.Path, h.Title)
}
// output is data.output.output from step 5: a string holding the reply JSON.
String t = output.strip().replaceFirst("^```(?:json)?\\s*", "").replaceFirst("\\s*```\\s*$", "");
String json = t.substring(t.indexOf('{'), t.lastIndexOf('}') + 1);
// With Jackson: JsonNode r = new ObjectMapper().readTree(json);
// r.get("status"): sound | tidy_up | restructure
// r.get("findings"), r.get("hierarchy"), r.get("redirects"), r.get("navigation"), r.get("linking")
System.out.println(json);
t = text.strip.sub(/\A```(?:json)?\s*/i, "").sub(/\s*```\s*\z/, "")
reply = JSON.parse(t[t.index("{")..t.rindex("}")])
puts reply["status"], reply["headline"]
reply["findings"].each { |f| puts "#{f['ref']} #{f['severity']} #{f['finding']}" }
reply["hierarchy"].each { |h| puts "#{' ' * h['level']}#{h['path']} #{h['title']}" }
reply["redirects"].each { |x| puts "301 #{x['from']} -> #{x['to']}" }
<?php
$t = preg_replace('/\s*```\s*$/', "", preg_replace('/^```(?:json)?\s*/i', "", trim($text)));
$reply = json_decode(substr($t, strpos($t, "{"), strrpos($t, "}") - strpos($t, "{") + 1), true);
echo $reply["status"], " ", $reply["headline"], PHP_EOL;
foreach ($reply["findings"] as $f) { echo $f["ref"], " ", $f["severity"], " ", $f["finding"], PHP_EOL; }
foreach ($reply["hierarchy"] as $h) { echo str_repeat(" ", $h["level"]), $h["path"], " ", $h["title"], PHP_EOL; }
foreach ($reply["redirects"] as $x) { echo "301 ", $x["from"], " -> ", $x["to"], PHP_EOL; }
var t = System.Text.RegularExpressions.Regex.Replace(output.Trim(), @"^```(?:json)?\s*", "");
t = System.Text.RegularExpressions.Regex.Replace(t, @"\s*```\s*$", "");
var json2 = t.Substring(t.IndexOf('{'), t.LastIndexOf('}') - t.IndexOf('{') + 1);
using var parsed = JsonDocument.Parse(json2);
var r = parsed.RootElement;
Console.WriteLine($"{r.GetProperty("status")} {r.GetProperty("headline")}");
foreach (var h in r.GetProperty("hierarchy").EnumerateArray())
Console.WriteLine($"{h.GetProperty("level")} {h.GetProperty("path")} {h.GetProperty("title")}");
Invariants worth asserting
The web page holds every reply to the browser's facts and to the inventory before it shows it
(recon.js, reconcile()). Do the same before you hand a redirect map to a
server:
- Flags: every flag in
facts.flagshas exactly oneprescan_responsesentry; no response names a flag that was not sent; adismissedflag carries a reason (read it and decide whether you agree). - Status: one of
sound,tidy_up,restructure, and never looser than the flags left standing (those not dismissed): anyhighmeansrestructure, else anymediummeans at leasttidy_up. - Findings: every standing
highormediumflag appears in somefindings[].ref, and every id named there was sent. - Hierarchy: it has
/; every node's parent is in the plan and its path starts with the parent's path (so breadcrumbs mirror the URL); every path is root-relative, lowercase, hyphenated, with no extension, query string, date or numeric id, and one trailing-slash style; no node is listed twice. - Nothing invented: every node not marked
is_newis in your inventory, or sits under a/*node that covers inventory URLs, or receives a redirect. - Navigation: 4-7 header items (a house guideline); every header, CTA and footer path is a node in
hierarchy. - Redirects: every
fromis a URL from the inventory, copied exactly (a wildcardfrommust cover some); everytois a node in the plan (or under a/*node); no chains, loops, self-redirects or duplicate sources; no URL is both kept and redirected. - Coverage: every inventory URL you sent is kept (its path is a node, or it sits under a
/*node) or redirected, except utility URLs and URLs already answering 3xx, 4xx or 5xx.
# assert-reply.py - the core checks, for body.json and reply.json from the steps above.
import json, re
body = json.load(open("body.json")); facts = json.loads(body["facts"])
t = open("reply.json").read(); r = json.loads(t[t.index("{"):t.rindex("}") + 1])
S = lambda p: p if p == "/" else p.rstrip("/")
wild = lambda p: p.endswith("/*")
flags = {f["id"]: f for f in facts.get("flags", [])}
ids = [p["id"] for p in r["prescan_responses"]]
assert sorted(ids) == sorted(flags), "every flag answered exactly once"
dismissed = {p["id"] for p in r["prescan_responses"] if p["status"] == "dismissed"}
standing = [f for i, f in flags.items() if i not in dismissed]
rank = ["sound", "tidy_up", "restructure"]
floor = 2 if any(f["severity"] == "high" for f in standing) else 1 if any(f["severity"] == "medium" for f in standing) else 0
assert rank.index(r["status"]) >= floor, "status looser than the flags left standing"
covered = {x for f in r["findings"] for x in re.findall(r"F\d+", f["ref"])}
assert all(f["id"] in covered for f in standing if f["severity"] != "low")
H = r["hierarchy"]; plan = {S(h["path"]) for h in H}
assert "/" in plan, "no homepage node"
for h in H:
if h["path"] == "/": continue
par = S(h["parent"] or "/")
assert par in plan, h["path"] + " has a parent outside the plan"
assert par == "/" or (S(h["path"]) + "/").startswith(par.removesuffix("/*") + "/"), h["path"] + " does not mirror its parent"
assert not re.search(r"[A-Z_?]|\.(html?|php|aspx?)$", h["path"]), h["path"] + " breaks the URL rules"
inv = [l.split("\t")[0] for l in body["inventory"].splitlines() if l.strip()]
under = lambda p, w: (p + "/").startswith(w.removesuffix("/*") + "/") and p != w.removesuffix("/*")
in_plan = lambda p: S(p.removesuffix("/*") if wild(p) else p) in plan or any(wild(w) and under(p, w) for w in plan)
R = r["redirects"]
for x in R:
assert (any(under(p, x["from"]) for p in inv) if wild(x["from"]) else x["from"] in inv), "redirect from outside the inventory: " + x["from"]
assert in_plan(x["to"]), "redirect to a page not in the plan: " + x["to"]
assert x["from"] != x["to"], "self-redirect"
froms = [x["from"] for x in R]
assert len(froms) == len(set(froms)), "a URL is redirected twice"
assert not any(x["to"] in froms for x in R if not wild(x["to"])), "redirect chain"
nav = r["navigation"]
for l in nav["header"] + [nav["cta"]] + [l for g in nav["footer"] for l in g["links"]]:
assert not l.get("path") or S(l["path"]) in plan, "nav link outside the plan: " + l["path"]
skip = set(facts.get("utility_urls", [])) | {l.split("\t")[0] for l in body["inventory"].splitlines() if "\tstatus=" in l}
lost = [p for p in inv if p not in skip and not in_plan(p.split("?")[0]) and not any(p == x["from"] or (wild(x["from"]) and under(p, x["from"])) for x in R)]
assert not lost, "inventory URLs neither kept nor redirected: " + ", ".join(lost[:5])
The output contract
Every key is always present. Arrays may be empty and strings may be "" when there is
nothing to say. An enum is written "a|b|c": the reply carries exactly one of the
values. Text fields are plain prose: no Markdown.
{"lane":"architect",
"site":{"name":"...","type":"saas|content|ecommerce|docs|hybrid|local"},
"status":"sound|tidy_up|restructure",
"headline":"...","summary":"...",
"findings":[{"ref":"F1, F4","severity":"high|medium|low","finding":"...","fix":"..."}],
"hierarchy":[
{"path":"/","title":"Homepage","parent":"","level":0,"nav":"none","priority":"high","is_new":false,"count":null},
{"path":"/features","title":"Features","parent":"/","level":1,"nav":"header|header_dropdown|footer|sidebar|none","priority":"high|medium|low","is_new":true,"count":null},
{"path":"/blog/*","title":"Blog posts","parent":"/blog","level":2,"nav":"none","priority":"medium","is_new":false,"count":212}],
"redirects":[
{"from":"/Product/Analytics","to":"/features/analytics","reason":"..."},
{"from":"/resources/*","to":"/library/*","reason":"the whole folder moves, rest of the path kept"}],
"navigation":{"header":[{"label":"Features","path":"/features"}],
"cta":{"label":"Start free trial","path":"/signup"},
"footer":[{"group":"Company","links":[{"label":"About","path":"/about"}]}],
"breadcrumbs":"..."},
"linking":{"hubs":[{"hub":"/features","spokes":["/features/analytics"]}],
"cross_links":[{"from":"/features/analytics","to":"/customers/acme","anchor":"..."}],
"orphans":[{"path":"/old-page","link_from":["/blog"]}]},
"next_steps":["..."],
"prescan_responses":[{"id":"F1","status":"confirmed|dismissed","reason":"..."}]}
| key | shape | what it holds |
|---|---|---|
lane | string | Always "architect". |
site | {name, type} | The site's name and its type, one of the six site_type values. |
status | enum | The state of the structure; see the next table. |
headline | string | One sentence: the state of the structure and the single most important change. |
summary | string | 3-5 sentences: what is wrong, what the plan does, what it costs to migrate; answers question when there is one. |
findings | array of {ref, severity, finding, fix} | At most 12. One per confirmed high or medium flag (several ids may share one, "F3, F5"), plus any the model found itself in the input (ref ""). |
hierarchy | array of {path, title, parent, level, nav, priority, is_new, count} | At most 60 nodes: the homepage, every section hub and every page at level 1 or 2. level is steps below the homepage; parent is "" for /; is_new marks a proposed page. A section of many similar detail pages is its hub, up to 3 representative children, and ONE collection node whose path ends in /* ("/blog/*"): every page under that folder keeps its pattern. count is a number only on /* nodes, otherwise null. |
redirects | array of {from, to, reason} | At most 80 permanent (301) redirects. from is copied exactly from the inventory (same case, trailing slash and query string); to is a node in the plan. A wildcard rule moves a whole folder: "from": "/resources/*", "to": "/library/*" keeps the rest of the path, "to": "/library" sends the whole folder to one page. Host, protocol and trailing-slash-policy fixes are one server rule each and appear in next_steps, not here. Empty when nothing moves. |
navigation | object | header ({label, path}, 4-7 items), cta ({label, path}, rightmost and separate), footer (columns of {group, links}) and breadcrumbs (a sentence or two). Every path is a node in hierarchy. |
linking | object | hubs ({hub, spokes}; spokes link back to the hub), cross_links ({from, to, anchor} with descriptive anchor text), orphans ({path, link_from}, one per facts.orphans entry that stays). |
next_steps | array of strings | 3-8 ordered, concrete migration steps. |
prescan_responses | array of {id, status, reason} | Exactly one per flag in facts.flags: confirmed or dismissed, with the reason. Empty when no flags were sent. |
Status
| status | meaning |
|---|---|
sound | No confirmed high or medium flags; the plan only suggests additions (navigation, links, new pages). |
tidy_up | The hierarchy is basically right, but conventions, hubs, duplicates or links need fixing (confirmed medium flags). |
restructure | Any confirmed high flag, or the hierarchy itself must change: sections merged, moved or renamed, many URLs changing. |
The model may be stricter than facts.browser_status, never looser, unless it
dismissed the flags that set it.
Enums
| where | values | the page's fallback |
|---|---|---|
status | sound, tidy_up, restructure | restructure |
findings[].severity | high, medium, low | medium |
hierarchy[].nav | header, header_dropdown, footer, sidebar, none | none |
hierarchy[].priority | high, medium, low | medium |
prescan_responses[].status | confirmed, dismissed | confirmed |
Worked example
The Brindlecott Plumbing sitemap from the page's examples: a local business with 14 URLs, two
folders (/services and /service-areas) that both have hub pages, one
consistent no-trailing-slash style, and both key pages one click from the homepage. The prescan
raised no flags, so browser_status is sound and
prescan_responses comes back empty. This request is complete and sendable as it
stands; its Idempotency-Key from the page is
sitemap-desk:architect:212ua9axq9u:a1.
The request body:
{
"task": "architect",
"site_name": "Brindlecott Plumbing",
"site_type": "local",
"goals": "Phone calls and booking requests from homeowners in the service area.",
"audiences": "Homeowners in Eugene and Springfield with an urgent or planned plumbing job.",
"key_pages": "/services\n/book",
"current_nav": "Header: Services, Service areas, About, Book a visit. Footer: Contact, Privacy.",
"inventory": "/\n/services\n/services/drain-cleaning\n/services/water-heaters\n/services/leak-repair\n/services/emergency-plumbing\n/service-areas\n/service-areas/eugene\n/service-areas/springfield\n/about\n/reviews\n/book\n/contact\n/privacy",
"facts": "{\"host\":\"brindlecottplumbing.example\",\"hosts\":[\"brindlecottplumbing.example\"],\"format\":\"urlset\",\"url_count\":14,\"sent_count\":14,\"clipped\":\"\",\"slash_policy\":\"none\",\"max_depth\":2,\"depth_counts\":{\"0\":1,\"1\":7,\"2\":6},\"typical_max_depth\":2,\"sections\":[{\"path\":\"/services\",\"pages\":5,\"hub\":true,\"children\":4},{\"path\":\"/service-areas\",\"pages\":3,\"hub\":true,\"children\":2},{\"path\":\"/about\",\"pages\":1,\"hub\":true,\"children\":0},{\"path\":\"/reviews\",\"pages\":1,\"hub\":true,\"children\":0},{\"path\":\"/book\",\"pages\":1,\"hub\":true,\"children\":0},{\"path\":\"/contact\",\"pages\":1,\"hub\":true,\"children\":0},{\"path\":\"/privacy\",\"pages\":1,\"hub\":true,\"children\":0}],\"folders_without_hub\":[],\"has_inlinks\":false,\"orphans\":[],\"key_pages\":[{\"path\":\"/services\",\"found\":true,\"depth\":1},{\"path\":\"/book\",\"found\":true,\"depth\":1}],\"expected_missing\":[],\"browser_status\":\"sound\",\"flags\":[]}",
"question": ""
}
Its facts, decoded:
{
"host": "brindlecottplumbing.example",
"hosts": [
"brindlecottplumbing.example"
],
"format": "urlset",
"url_count": 14,
"sent_count": 14,
"clipped": "",
"slash_policy": "none",
"max_depth": 2,
"depth_counts": {
"0": 1,
"1": 7,
"2": 6
},
"typical_max_depth": 2,
"sections": [
{
"path": "/services",
"pages": 5,
"hub": true,
"children": 4
},
{
"path": "/service-areas",
"pages": 3,
"hub": true,
"children": 2
},
{
"path": "/about",
"pages": 1,
"hub": true,
"children": 0
},
{
"path": "/reviews",
"pages": 1,
"hub": true,
"children": 0
},
{
"path": "/book",
"pages": 1,
"hub": true,
"children": 0
},
{
"path": "/contact",
"pages": 1,
"hub": true,
"children": 0
},
{
"path": "/privacy",
"pages": 1,
"hub": true,
"children": 0
}
],
"folders_without_hub": [],
"has_inlinks": false,
"orphans": [],
"key_pages": [
{
"path": "/services",
"found": true,
"depth": 1
},
{
"path": "/book",
"found": true,
"depth": 1
}
],
"expected_missing": [],
"browser_status": "sound",
"flags": []
}
The reply, parsed from data.output.output and abbreviated: 2 of 4 findings, 6 of 14
hierarchy nodes, 1 of 3 footer columns, 1 of 2 hubs, 2 of 6 cross links and 2 of 6 next steps are
shown, and the summary ends in …. The status is sound: nothing moves, so
redirects is empty, and the plan is navigation and linking work only. A site that must
restructure gets the same keys filled with redirects (exact and wildcard) and, for large sections,
/* collection nodes, as in the contract above.
{
"lane": "architect",
"site": {
"name": "Brindlecott Plumbing",
"type": "local"
},
"status": "sound",
"headline": "The 14-page structure is already flat, consistent and within two levels, so no URL should change; the most valuable move is to put Book a visit in a separate rightmost header CTA and cross-link every service page with the Eugene and Springfield area pages.",
"summary": "The site is in good shape: every page sits at depth 1 or 2, both folders (/services and /service-areas) have hub pages, all paths are lowercase, hyphenated and follow one no-trailing-slash style, and both key pages (/services and …",
"findings": [
{
"ref": "",
"severity": "low",
"finding": "The current header lists Book a visit alongside ordinary sections, and Reviews and Contact are missing from it even though phone calls and bookings are the site's goals and reviews are the main trust signal for a local trade.",
"fix": "Header: Services, Service areas, Reviews, About, Contact, with Book a visit as a separate button at the far right."
},
{
"ref": "",
"severity": "low",
"finding": "The footer holds only Contact and Privacy, so the service and area pages that a homeowner scans for at the bottom of a page are not reachable from it.",
"fix": "Group the footer into Services, Service areas and Company columns, linking all four service pages, both area pages, and About, Reviews, Contact and Privacy."
}
],
"hierarchy": [
{
"path": "/",
"title": "Homepage",
"parent": "",
"level": 0,
"nav": "none",
"priority": "high",
"is_new": false,
"count": null
},
{
"path": "/services",
"title": "Services",
"parent": "/",
"level": 1,
"nav": "header",
"priority": "high",
"is_new": false,
"count": null
},
{
"path": "/services/emergency-plumbing",
"title": "Emergency plumbing",
"parent": "/services",
"level": 2,
"nav": "footer",
"priority": "high",
"is_new": false,
"count": null
},
{
"path": "/service-areas",
"title": "Service areas",
"parent": "/",
"level": 1,
"nav": "header",
"priority": "high",
"is_new": false,
"count": null
},
{
"path": "/service-areas/eugene",
"title": "Eugene",
"parent": "/service-areas",
"level": 2,
"nav": "footer",
"priority": "high",
"is_new": false,
"count": null
},
{
"path": "/book",
"title": "Book a visit",
"parent": "/",
"level": 1,
"nav": "header",
"priority": "high",
"is_new": false,
"count": null
}
],
"redirects": [],
"navigation": {
"header": [
{
"label": "Services",
"path": "/services"
},
{
"label": "Service areas",
"path": "/service-areas"
},
{
"label": "Reviews",
"path": "/reviews"
},
{
"label": "About",
"path": "/about"
},
{
"label": "Contact",
"path": "/contact"
}
],
"cta": {
"label": "Book a visit",
"path": "/book"
},
"footer": [
{
"group": "Service areas",
"links": [
{
"label": "Eugene",
"path": "/service-areas/eugene"
},
{
"label": "Springfield",
"path": "/service-areas/springfield"
}
]
}
],
"breadcrumbs": "Show breadcrumbs on level-2 pages only, mirroring the URL, for example Home > Services > Water heaters and Home > Service areas > Eugene; level-1 pages do not need them."
},
"linking": {
"hubs": [
{
"hub": "/service-areas",
"spokes": [
"/service-areas/eugene",
"/service-areas/springfield"
]
}
],
"cross_links": [
{
"from": "/service-areas/eugene",
"to": "/services/emergency-plumbing",
"anchor": "emergency plumbing in Eugene"
},
{
"from": "/service-areas/springfield",
"to": "/services/water-heaters",
"anchor": "water heater repair and replacement in Springfield"
}
],
"orphans": []
},
"next_steps": [
"Keep all 14 URLs unchanged; no redirects are needed and the no-trailing-slash style is already consistent.",
"Rebuild the header as Services, Service areas, Reviews, About, Contact, with Book a visit as a separate rightmost button on every page."
],
"prescan_responses": []
}
Truncation and partial results
When the balance sits between min_credits and hold_credits, the run is not
refused: it executes with a reduced output cap and reports truncated: true, in the
done event of /run-stream and on the job from GET /jobs/{id}.
What you hold is then a prefix of the reply. The web page closes the cut-off JSON
(Recon.closeJson in recon.js) and shows the sections that arrived, out of
eleven: site, status, headline, summary,
findings, hierarchy, redirects, navigation,
linking, next_steps and prescan_responses. A truncated plan
can be missing redirects or whole branches of the hierarchy, so never deploy a redirect map from
one: check the flag, top up, and resubmit with the attempt suffix on the
Idempotency-Key incremented (sitemap-desk:architect:<hash>:a2).
If a complete reply will not parse as one JSON object, the page retries once, as the next attempt,
with a retry_note saying what was wrong and asking for only the JSON object for task
architect. Do the same: keep the hash, bump the attempt, add retry_note.