Nevermined Catalog · Scoring

Catalog Scoring System

Every listed x402 / MPP service now carries two independent scores. The readiness score (0–100) tells a builder how complete and agent-ready the listing is. The rank score orders discovery from measured market behaviour. They read some of the same signals, but neither score feeds the other.

Readiness rubric
readiness/1.0.0
Readiness refresh
7 days TTL
Rank signal window
30 days
Sybil floor
2 distinct payers

What changed since the September model

  • NewReadiness score. A 0–100 rubric across six dimensions, with caps, bands and ranked improvements. It is advisory and never gates listing (#4070, #4124, #4172).
  • NewCanary verification factor on rank. ×1.00 verified, ×0.85 unverified, ×0.50 degraded. It affects ordering only and leaves qualityScore untouched (#3775).
  • RemovedPaid organization tier. It no longer multiplies rankScore. API pins below 1.55 receive rankScore: null (#3920).
  • ChangedLiveness is harder to inflate. Partial-health evidence caps it at its confidence (#3977). A listing that several buyers tried to pay without one settlement loses its probe uptime (#3758).
  • ChangedRouter selection blends quality. /router/select re-ranks payable matches by relevance × (1 + 0.25·quality) instead of using quality as a tiebreak (#3816, #3832).
01

Two scores at a glance

They answer different questions for different people. A service can be excellent on one and weak on the other; §05 shows both cases.

Readiness scorecatalog_scoring_resultsRank scorecatalogServices.rankScore
QuestionIs this listing complete enough for an agent to find, understand, pay and call?Among matching services, which should a buyer see first?
AudienceThe builder, plus a public tile and badgeBuyers and harnesses, through ordering
RangeInteger 0–100 and a bandOpaque number ≥ 0 that can exceed 1. It is a sort key, not a percentage.
InputsListing metadata, endpoints, pricing, docs and operations facts, plus a live probeSettlement ledger (router_payments), monitor probes, quoted prices, editorial levers
Computed byAPI rubric over evidence the catalog-monitor worker probes (queue + ops API)API health sweep worker, every tick; immediately on a Featured edit
RefreshOn publish, on manual re-analysis, after 7 days, and after a rubric MAJOR/MINOR bumpContinuous, over a rolling 30-day window
EffectNone on listing, moderation or rankingOrders the catalog browse view and Router selection
02

Readiness score

24 rules across six weighted dimensions add up to 100 points. A failed precondition caps the total, and the result lands in one of four bands. Every result comes with a ranked list of the fixes worth the most points.

readiness = min( Σ rule points , lowest triggered cap )   // rounded to a whole number, then banded
Metadata 15 Endpoints 25 Pricing 20 Documentation 15 Operations 10 Health 15

Metadata

15 pts
  • M1Specific, non-placeholder title3
  • M2Short description ≥ 20 chars, not the title3
  • M3Detailed description ≥ 60 chars3
  • M4Logo that actually loads3
  • M5Category and homepage URL3

Endpoints

25 pts
  • E1At least one endpoint published5
  • E2Method, path and descriptionper endpoint5
  • E3Request schema, or “no parameters”per endpoint5
  • E4Response schemaper endpoint5
  • E5Safe request exampleper endpoint5

Pricing

20 pts
  • P1Price or quote publishedper payable endpoint6
  • P2Price is machine-readableper payable endpoint6
  • P3Declared fixed or dynamicper payable endpoint4
  • P4Dynamic price: range, driver, unitper dynamic endpoint · full marks if none4

Documentation

15 pts
  • D1Docs reachable at a public URL4
  • D2Setup, auth and payment walkthrough4
  • D3Schemas and examples in the docsper endpoint4
  • D4Machine-readable discovery document3

Operations

10 pts
  • O1Auth preconditions declared3
  • O2Network and funding requirements3
  • O3Limits, timeouts, sync/async behaviour2
  • O4Errors, idempotency, support channel2

Health

15 pts
  • H1Reachable, with a valid payment challenge5
  • H2Liveness5 × liveness (same formula as rank)5
  • H3Settleabilitysettleable 5 · unknown 3.33 · unpayable 05

Rules marked per endpoint pay pro rata. If 3 of 4 endpoints carry a response schema, E4 earns 3.75 of 5. All other rules pay in full or not at all, except H2 and H3, which are measured. Duplicate endpoints (same method and path) count once.

Caps and bands

A missing precondition outweighs good documentation

When a precondition fails, the total is capped at the lowest triggered value. Both cap values sit below 60, so a capped service always reads Getting started, however polished the rest of the listing is. The fix appears first in its improvement list, flagged as a blocker.

39noReachableEndpoint
No reachable endpoint

No endpoints, or none of them can actually be called.

39knownUnpayable
Known unpayable

The settlement ledger’s verdict is unpayable: repeated failed payments from distinct buyers.

59noPaymentMechanism
No payment mechanism

A caller can’t discover how to pay before invoking (no x402 or MPP challenge found).

Improvements are ordered: blockers first, then most recoverable points, then most endpoints affected, then lowest effort.

How an analysis runs
  1. Trigger

    An owner submits a URL, a listing is published, or the staleness sweep re-queues a 7-day-old score.

  2. Queue

    Durable Postgres row: deferred → queued → running → succeeded | failed. One active analysis per submitter.

  3. Probe

    The catalog-monitor worker claims the lease through the ops API and probes the target (x402 + MPP, SSRF-guarded).

  4. Score

    The API applies the rubric to the returned evidence, applies caps and ranks the improvements.

  5. Store

    A new immutable result row, stamped with its rubric version. The listing pointer moves; the tile and badge read it.

Provisional estimate

A score before the probe

A provisional run scores the listing without live health evidence. H1–H3 count as “not observed” (0 points) and no cap is evaluated. The owner sees it labelled provisional; it is never published, never shown on the badge and never marked stale.

Governance

Versioned, immutable, signed off

Versions follow readiness/MAJOR.MINOR.PATCH, and a released version never changes. A MAJOR or MINOR bump re-scores every listing (version mismatches first) and shows the builder “Scoring criteria updated… this reflects the rubric, not your service.” A PATCH is reserved for changes that move no score. Changes need CODEOWNERS sign-off, CI enforces a changelog entry, and the monitor must support a version before the API serves it.

03

Rank score

The measured quality score keeps the September weights. The rank score then applies the editorial Featured lever and the new canary verification factor on top. Paid organization tier is no longer part of it.

qualityScore = SettleGate × ( 0.25·Liveness + 0.45·Traction + 0.30·PriceIntegrity )  // 0–1, measured, internal
rankScore    = qualityScore × Featured × Verification  // public sort key, may exceed 1
Liveness: reachable and up Traction: real, distinct buyers Price integrity: advertised price matches what buyers pay
Measured gate

SettleGate

Does a payment to this listing actually clear?

settleable× 1.00
unknown× 0.85
unpayable× 0.25

Settled within 30 days → settleable. Ledger-proven failures → unpayable. Neither → unknown.

Editorial lever

Featured

A Nevermined moderator sets level 0–5, giving 1 + 0.1 × level.

level 0 (default)× 1.00
level 3× 1.30
level 5× 1.50

Neutral at 0, so it is never a penalty. It is forced to ×1.00 when the gate says unpayable. Internal only, so buyers can’t tell promotion from quality.

New · canary

Verification

Have Nevermined’s canary runs verified every listed endpoint?

verified-by-canary× 1.00
unverified× 0.85
degraded× 0.50

Verified means every listable endpoint passes the canary and carries a proven request example. One degraded, broken or hidden endpoint marks the service degraded.

DimensionHow the raw signals become 0–1
Liveness7-day uptime ÷ 100, falling back to 30-day. With no uptime yet, a baseline from health: operational 0.9, unverified 0.6, degraded 0.5, unavailable 0.1.
Capped at the monitor’s health confidence, so partial evidence such as a structured 400 cannot score as fully live. If buyers have tried to pay and none settled, the probe uptime is withheld and the listing falls back to its baseline.
TractionDistinct payers in 30 days ÷ 50, capped at 1.
Fewer than 2 distinct payers scores 0. One wallet, possibly the owner’s, cannot manufacture demand.
Price integrityWith ≥ 2 distinct payers: 1 − divergence between the settled and the advertised price. Otherwise 0.7 if a price is advertised (a quote under 30 days old, or the seller’s price label), or 0.3 if none is.
The advertised price is checked against what buyers actually pay.
Removed · paid tier

The September rankScore also multiplied by the seller organization’s subscription tier. That factor was undisclosed and contradicted the catalog’s neutrality commitment, so it was taken out. Neither score now depends on what a seller pays Nevermined. The catalog keeps its explicit, visible curation levers (tier, sortWeight). Clients pinned below API version 1.55 receive rankScore: null rather than a number whose meaning changed.

Where rank decides the order

Catalog browse

The default view with no search term.

tier ↑
→ sortWeight ↓
→ rankScore ↓ (unscored last)
→ seeded shuffle, rotates every 6 h

Search

Catalog ?search=, /ard/search, MCP. Results are ordered by relevance (hybrid semantic + lexical), and quality is not blended in.

relevance ↓

Router selection

/router/select re-ranks the payable matches. Semantic matches stay ahead of keyword-only ones.

relevance × (1 + 0.25·q)
q = rankScore ÷ best in set
04

Why the scores hold up

Scores built on payments invite self-payment, and scores built on probes invite abuse of the probe. These floors make faking a signal cost more than it gains.

MIN_DISTINCT_FAILING_BUYERS = 2

Distinct-buyer floor

Unpayable verdicts and price signals need ≥ 2 distinct buyers. One broken wallet can’t sink a listing, and one owner wallet can’t certify it.

Rank · Readiness H3
MIN_SETTLE_ATTEMPTS = 3

Attempt floor

An unpayable verdict needs ≥ 3 failed attempts, so a couple of transient failures never sink a working service.

Rank · Readiness cap
settledPrice withheld below floor

Self-pay resistance

Traction is 0 and the public settled price is hidden until ≥ 2 distinct payers exist. Paying yourself proves nothing.

Rank
settlementUnprovenSince

No uptime without a settlement

If several buyers tried to pay and none settled, the unpaid probe’s uptime is withheld. A service nobody can complete a call against isn’t “100% up”.

Rank
quote age < 30 days

Quote freshness

A stale quote never feeds price integrity, so a legitimate re-price can’t be read as price drift against the seller.

Rank
uniq_catalog_scoring_active_subject

Bounded, guarded analysis

One active analysis per submitter. Target URLs that carry credentials, a query string or a fragment are refused before any probe. Results are never rewritten.

Readiness
05

Both scores on real shapes

illustrative

Listing shapes run through both formulas. The numbers are illustrative, not live data. Quality and rank are shown ×100 to make them readable; rank is still a sort key, not a percentage.

ReadinessRank
Listing shapeScoreGateLTPQualityFeatVerifRank
Mature and honest44 payers · settled ≈ quoted · full docs except the walkthrough
96Excellent
1.009588969201.00
92
Mature and honest, featuredsame service · Nevermined featured at level 3
96Excellent
1.009588969231.00
120
Popular but thin listing60 payers · no schemas, examples or docs · not canary-verified
57Getting started
1.0092100909500.85
81
Well built, brand new0 payers (floored) · quoted price only · featured at max · no walkthrough, discovery doc or limits
87Ready
0.85900703750.85
47
Trafficked but unpayablehad demand · 80% price drift · fails to settle · endpoint degraded
39Getting started · capped
0.255570201350.50
6

The scores disagree by design. The popular-but-thin service ranks well because buyers keep paying it, but its readiness is low because an agent arriving cold has nothing to work from. The well-built newcomer is the reverse: ready on day one, with no demand yet to rank on.

Guards hold under promotion. Max Featured on the unpayable listing does nothing, because the gate forces ×1.00 and its readiness is capped at 39 (raw 80). Featuring the newcomer lifts it from 31 to 47, still well below an unfeatured 92: featuring amplifies quality and never overrides it.

06

Who sees what

Readiness is public as a score and band only. Rank is published as a single opaque number, so measured quality can’t be separated from editorial promotion.

SurfaceReadinessRank
Owner readiness pagewebapp · signed inFull report: dimensions, checks, ranked improvements, provisional label, Stale chip, rubric-change notice—
Public catalog detaillive listings onlyScore, band, rubric version and analysis date. Provisional scores are never shown.Not shown
Readiness badgesigned SVG · 4 stylesScore and band, linking to the catalog page once published—
/router/selectNevermined API key—rankScore and relevance per shortlist item (API ≥ 1.55)
Keyless feedsai-catalog.json · ARD—Order only, no number
Internal / moderationadminRequest queue, evidence, historyqualityScore, featuredLevel, verification status
07

Limits and rollout status

Stating these boundaries keeps the scores honest.

Status · 7 Oct 2026
  • Readiness scoring is merged and running on staging (argocd #714). Production enablement is pending (argocd #716 and #721). Until it lands, production analyses wait in the queue.
  • Canary verification needs the paid canary, and production’s is still held (argocd #706). Expect most production listings to read unverified (×0.85), which applies to everyone equally and so leaves their relative order unchanged.
  • Dynamic-pricing checks (P3, P4) depend on capturing MPP session intent and populating the price driver (#4081, #4082). Both are deferred.
  • Traction depth. router_payments history is still short, so traction and settled-price signals stay near zero for most listings. For now, rank leans on liveness and the quoted price.

Output quality

Neither score grades whether a service’s answers are good. Readiness measures completeness, and rank measures market behaviour.

Revenue in fiat

No FX rate is captured at purchase time, so there is no single money total across mixed assets. Counts and per-asset prices are available.

Whole-market rank

Standing across the whole x402 market needs external chain data that can’t be attributed to Nevermined. It is context, never a score input.