Public methodology
How the notation pipeline actually runs
Not a single LLM pass. Not a mega-score. Five public sources stay unfused; stats produce an envelope; template NLI reads claim vs evidence; a lab desk argues both sides from the same numbers. Indicative. Contestable. Not a licensed CRA.
The diagram below is the product — coded from the same rules as the scores, not a slide pasted on the page. Lab and research sit on dashed / dotted lanes so they cannot be mistaken for live coverage.
AUROS rating pipeline — from public file to contestable note
Sources
liveFive inputs, never fused. The lab desk reads this bus — no sixth feed.
Named scores
liveSeparate layers. Each 0–100 + badge. Never a mega-rating.
S = { Activity, Gap, Freshness, Depth, Friction }Envelope
liveLeave-one-dimension-out → disagreement band. Stats object, not a vibe.
method = lodo_disagreement · band = [min(point, LOO) − σ̂, max(point, LOO) + σ̂]Conformal 95% (research)
researchNot live — needs a labeled holdout (split / CV+).
coverage target 95% — not implementedClaim vs evidence (NLI)
liveTemplate labels on Promise Gap. Not a legal verdict. Not a chat pass.
label ∈ {entailment, contradiction, neutral}Adversarial desk
labBull, bear, arbiter — same inputs, reproducible.
3 reads(S) → {bull, bear, arbiter} — not a live 3-LLM panelCite and contest
livePublic files. A note you cannot recompute is marketing.
cite getauros.com/technology/methodologyExploratory research, not live: zkTLS · entity graph · wash clustering · Bayesian changepoint · satellite / IoT
Why this is not “just AI”
- Envelope = leave-one-dimension-out + residual spread. Deterministic. Same inputs, same band.
- NLI live = labelled rules on structured fields (entail / contradict / neutral), not a chat completion you cannot replay.
- Desk = three briefs from those metrics. A live multi-agent LLM would be a later overlay, never a substitute for the arithmetic.
Technology hub · Activity · Promise Gap · Score envelope