Methodology · rubric v0

How we score what an agent can read.

SEO got you found. AEO & GEO get you cited and recommended. The Prolectio Score measures the layer after that: how well an AI agent — not a human — can actually experience your public documentation and complete a real task against your API. Published, versioned, and reproducible from stored inputs.

How to read a grade

Every score is 0–100, banded: green 80+ (an agent can reliably self-serve most first tasks), amber 50–79 (workable, with gaps that cost the agent tries), red <50 (an agent is likely to stall or abandon common tasks).

The three categories

CategoryWeightThe question it answers
Deterministic checks40%Are the machine-checkable fundamentals present and well-formed? Valid root llms.txt; a discoverable OpenAPI spec that validates; sane heading structure; code samples on reference pages; auth documented with an example; pages within an agent-friendly token budget; working links; crawler/meta hygiene.
LLM-rubric checks35%Is the prose clear, complete, and consistent enough for an agent to act on? A pinned model (claude-sonnet-5) grades key pages on clarity, instruction completeness, and terminology consistency. The model ID, exact prompt, and raw structured output are stored with every score.
Micro-task probes25%Can a real agent actually finish? Three standardized tasks — find the auth method, find the rate limits, construct a first valid request — graded deterministically on substance and citation, never on the agent's self-claim. Full transcripts are saved.

Task completion (probes + rubric, 60%) deliberately outweighs static structure (40%): you cannot top the Index by adding an llms.txt and nothing else.

The Agent Journey — how each stage is scored

The Prolectio Score above is specifically the Experience stage, at docs depth — that's what Index #1 benchmarks. The full Agent Journey an agent moves through against any business has four stages; here is exactly what each one measures and where its evidence lives.

1 · Discover — can an agent even reach you?

A pure, no-network read over signals a normal scan already captures: robots.txtAllow/Disallow rules per known-agent user-agent token, the emerging Content-Signal:declared-usage directives, and the response a known agent user-agent actually receives (status code, cf-mitigated/X-Robots-Tag headers). A block is reported as a scoreable, fixable finding — never a mystery, and never treated as evidence the edge is open just because robots.txt is silent. Live on every scan, every tier, free included.

2 · Experience — can an agent understand what you offer?

The Prolectio Score (above) at full depth for docs, plus the AX Audit: the same deterministic + model-graded blend applied across all eight surface classes — marketing, pricing, auth, checkout, account, support, legal, docs — combined into one composite score and a heatmap. Paid plans additionally get hosted llms.txt export and scan history/diffs.

3 · Interact — can an agent actually do the thing?

Two layers of task completion. First, the micro-task probes already described above (find-auth, find-rate-limits, construct-first-request) — deterministic substance grading, full transcript saved. Second, task journeys: a live agent runs one of four vetted templates (sign up, get a key, first API call, find the cancel path) end-to-end and is graded on completion, not self-claim. R1 — Live Interaction-Testing extends this to your real, running product rather than just documented steps; it is early access / private beta today, Owner-reviewed, not self-serve.

Agentic commerce is scored right here too — a called-out capability model inside Interact, not a blended sub-score. The Agentic Commerce Readiness Score (ACRS) grades eight capabilities an agent needs to complete a purchase — discovery, price & availability accuracy, cart construction, checkout initiation, payment acceptance, order confirmation, post-purchase, and anti-bot admission — as 21 deterministic checks over the whole site's crawled pages plus a bounded, robots-honoring probe of discovery paths (/.well-known/ manifests, product feeds) that live outside the page taxonomy. Each check is tagged to one of two sub-axes reported alongside the composite: protocol-native readiness (structured feeds, schema.org Offer/Product, ACP/UCP/commerce-MCP discovery, delegated-payment and x402 signals) and browser-agent readiness (non-JS-gated prices and cart/checkout paths, guest checkout, no CAPTCHA wall). A separate protocol-conformance pass then genuinely parses and validates what the checks detected — JSON-LD structured data, product-feed field coverage, x402response shape, ACP/UCP descriptors — and is attached as additive evidence with per-field findings. ACRS is a live per-scan score with its own red/amber/green band, reported on its own, never folded into the task-completion number above.

4 · Analyze — do you know when an agent shows up?

Not a graded score — live traffic. Agent sessions are classified from request signals, mapped to journey steps so stalls are visible per-step, and reconciled against conversions via reconciliation keys (order-intent tokens, API client IDs, payment-mandate IDs, IP/time-window fallback) so a conversion an agent drove is attributed to it, not lost. See your own data at /dashboard/analytics.

Transactability — the rung ladder, the SKUs, and certification

Agentic commerce (ACRS, above) grades whether your catalog and checkout are structurally reachable. Transactability is the live, repeated version of that question: did a real agent actually get there, on a schedule, and what happens the moment it stops? This section documents the rung ladder every Transactability claim on this site traces back to (pricing's Monitor / Monitor+ / Certified tiers, the monitoring dashboard, and every /verify/{domain} page).

The rung ladder, per protocol (ACP, UCP, commerce-MCP, x402)

RungNameWhat it means
0Not presentNo signal this protocol is exposed by this origin at all.
1Declared / staticThe protocol is present and structurally conformant — a one-time, deterministic read of the manifest/descriptor/feed. “The merchant says this works,” unverified live. This is the free one-time commerce-conformance check every Prolectio Score includes.
2Live-reachedA live agent run actually reached checkout, or the payment brink, for real — not a static inference. This is the Monitor / Monitor+ workhorse: it needs no merchant payment credentials and applies to the large majority of stores that have a checkout but no live payment capability yet.
3Live-confirmed (eligible-only)A live test-mode purchase was actually confirmed in the merchant's payment sandbox. Structurally opt-in and per-store: it requires the merchant's own sandbox credentials and an explicit consent record, and it is only ever attempted where a PSP already exposes agent-payable checkout. Never asserted as a default capability.

A store's overall rung is the minimum across every protocol it's scanned for — one missing or failing protocol caps the whole verdict, the same worst-status-wins posture the static ACRS report already takes. Every run persists a timestamped, signed evidence transcript (the agent's reasoning trace, not just a pass/fail bit) and, when it fails, a bounded failure-taxonomy tag — manifest_invalid, catalog_unfetchable, cart_create_failed, checkout_js_wall, agent_blocked, rate_limited, auth_challenge, payment_token_rejected, settlement_timeout, spec_version_drift, or other — so every failing state links to the evidence that produced it, never a bare red dot.

The SKUs and cadences

Agent-Blocking Audit (SKU A′) — a per-agent admission verdict (ChatGPT, Perplexity, Rufus, Claude, Google, and other named shopping agents) against your own bot-management, CAPTCHA, and rate-limit posture, with a fix. Free, one-time, on every plan; run again on every scheduled cycle once monitoring is on.

Monitor (SKU A, per store, one protocol) — Rung-1 static conformance plus scheduled live Rung-2 runs at a daily, weekly, or monthly cadence you choose per store (UTC hour, per-property — the same scheduling primitive Prolectio's dashboard-configurable scans already use). Regressions (a green run turning amber/red, or a rung downgrade) raise an email break alert within one cycle. Protocol-drift alerts (SKU B) watch the spec version itself: the moment ACP/UCP/commerce-MCP churns in a way that breaks your specific integration, you get a taxonomized alert plus a proposed fix — distinct from a generic “the spec changed” notice.

Monitor+ — the same spine across every protocol you run, plus Rung-3 sandbox purchase-verification the moment it's eligible, and an evidence-transcript export for your own records or a customer's procurement review.

Honesty note on cadence and status: this monitoring spine (scheduler, regression detector, drift detector, and the alert/evidence tables) is built and security-reviewed code — it is not yet self-serve for every customer while its scheduling migration and a signed-webhook key-management decision remain Owner-gated. Where a feature above is not yet live for your account, the dashboard says so plainly rather than fabricating a run.

Certification and /verify/{domain}

Certification is derived on read, every time — there is no separate stored “certified” flag and no revocation job to run. A store certifies when its fused rung has held at Rung-2 (checkout-reachable) or better, with no interruption, for a full trailing window (default 7 days). A store revokes the instant a scheduled run drops back below that bar after having qualified — the very next evaluation reflects it, with no stale badge in between. A store that has never yet qualified reads as monitoring, not revoked — revocation means losing something you actually held. And a store whose most recent run is older than the staleness bound reads as unknown rather than a frozen old status, so an abandoned store can't keep vouching for itself indefinitely.

prolectio.com/verify/{domain} and its JSON endpoint render the exact same computed payload, byte-for-byte — built for machines and B2B trust (procurement, investor diligence, an agent-operator registry), not as a consumer conversion seal. Publishing requires the store owner's own opt-in, the same publish-consent gate the Index uses; an unpublished or never-monitored domain always reads unknown, never a failing grade.

Competitive landscape

The entry layer here — scoring a manifest, watching it, showing a status page — is occupied and commoditizing. We name who does what below, and where Prolectio is actually different, rather than claiming a gap that doesn't exist.

PlayerWhat it doesWhere Prolectio still wins
UCP Checker (free, 17,700+ domains)24h monitoring, a live-agent playground to checkout, alerts, public /status/{domain} pages, a Verified Badge — single protocol (UCP).No fix delivery; single-protocol only; no agent-blocking analysis; a pass/fail status, not a signed reasoning transcript.
UCP.tools ($9/mo)ACP + UCP checkers, four-level checks, portfolio monitoring.A hobbyist tier — no fix delivery, no enablement path.
Google Lighthouse — Agentic Browsing (free, default)Static accessibility-tree / WebMCP / llms.txt checks, including checkout pages; under active development.Static only — never executes a live run; no monitoring, alerting, or fix. The biggest long-term watch item: if it adds real transaction execution, the entry layer changes.
ACP/Stripe, UCP/Shopify test harnessesPartner/developer tools to simulate agent checkout.Internal/dev-facing tooling, not a neutral third-party monitor, fix, and certification service.

Across all of them, the layers none currently touch are the ones Prolectio leads with: shipping a fix, not just a measurement; spanning multiple protocols with one cross-protocol comparison; centering the agent-blocking problem (bot-management turning away the agents a merchant wants to sell to); and a signed evidence transcript with the agent's reasoning, not a bare pass/fail.

What we will not claim

A trust commitment, stated plainly rather than buried in fine print:

  • We do not claim purchase completion for a store whose payment rails don't exist yet. Rung-3 is labeled eligible-only and only ever asserted where a merchant's PSP actually exposes agent-payable checkout.
  • We do not claim the Certification badge drives conversion. It is positioned as a machine-readable, B2B/procurement trust signal — an AI buyer parses structured data, it doesn't “see” a badge — never as a consumer trust-seal conversion lift.
  • We do not claim uptime or reliability beyond what we've actually observed in persisted runs. A rung grade, a reliability stat, and a certification status are all only ever as strong as the run history behind them.
  • We do not enter the money path. Prolectio never handles live funds, never stores live payment credentials, and is not a PSP or money-transmitter. Any payment-enablement work (roadmap, Owner-gated) only configures the merchant's own Stripe/Shopify integration, shipped as a reviewable fix — the merchant's PSP moves the money, never Prolectio.

How we crawl — good agent citizenship

  • Public pages only. No logins, no paywalled or private content.
  • We obey robots.txt. A site that disallows automated access is excluded, not scored. (Index #1 excluded: Cohere, WorkOS, Salesforce.)
  • Polite by default. One request per second, an identifying user-agent, a hard page cap.
  • No dark-pattern discovery. Ordinary same-origin links only; we don't probe hidden endpoints.

Scored as a non-JavaScript agent

Rubric v0 reads the HTML an agent receives without executing JavaScript — the behavior of a large class of real agents. Docs that render client-side arrive as near-empty shells and score accordingly. That is a real finding about that agent class, not a browser-quality verdict.

We verified this axis with a controlled JavaScript-render comparison on the ten lowest-rubric sites. Where rendering decisively changed the verdict — docs that are genuinely strong but invisible without a browser — we applied the rendered score and flagged the row † JS-REQUIRED. Sites that moved somewhat but not decisively are flagged ; their fetch-only score stands. For half the sites, rendering changed nothing: the low score is real.

Anti-gaming

  • Task completion outweighs static checks (60/40) — structural box-ticking can't carry a low-substance site to the top.
  • Hidden holdout checks rotate quarterly and are not published.
  • Every scoring change ships as a dated, public changelog entry — no silent re-weighting.

Reproducibility & corrections

Each score is reproducible from its stored inputs: the crawl snapshot, pinned model IDs, rubric version, exact prompts, and raw outputs. If you believe a score reflects a crawl error — a page we couldn't reach that a normal agent can — email corrections@prolectio.com and we'll re-run against a corrected snapshot and publish the update. The full process — re-scans, factual corrections, replies, and removal — is public at prolectio.com/disputes. Accuracy is the product; we'd rather fix a real miss than defend a number.

What this Index is not

Every Prolectio Score is our assessment — the output of the described process applied to observable inputs on a given date — not a statement of fact about any company's products, security, or legal standing. Specifically, it is:

  • Not a rating of the product or the API. Excellent engineering can ship agent-hostile docs, and vice-versa.
  • Not a security assessment. When a class such as authentication scores lower — for example because a CAPTCHA or bot-defense makes a page agent-incompatible — we are describing only that the flow is not optimized for automated agents, never that the site's security is weak, insecure, or deficient. A control that is good for human security is often, by design, simply not built for agents; we measure the latter and take no position on the former.
  • Not a legal or compliance opinion. When we score a site's legal or account pages, we measure only whether an AI agent can locate and parse them — never whether those documents are legally sufficient, enforceable, or compliant. We take no position on their legal adequacy.
  • Not a claim about every agent. Scores come from a pinned Claude model reading like an agent would; we publish the model so results are interpretable, not universal.
  • Not permanent. Docs change; scores are dated snapshots (this edition: 2026-07-19) and we re-scan on a published cadence.
  • Not a penalty for excluding bots. Respecting robots.txt is legitimate; we respect it back.

Changelog

  • 2026-08-31 — added the Transactability rung ladder (0–3), the Monitor/Monitor+/Certified SKU + cadence definitions, certification criteria and the /verify/{domain} page, a competitive-landscape section, and a “what we will not claim” commitment. Purely additive documentation of already-scored/shipped mechanics — no change to axScore, ACRS, or any rubric weight.
  • v0 (2026-07-19) — initial public rubric: deterministic checks (golden-set tested), LLM rubric pinned to claude-sonnet-5, micro-task probe harness with deterministic substance grading. JS-render comparison pass added; †/‡ flags introduced.