RESPONSIBLE AI · MEASUREMENT

Calibration Report

How AGENTPR™ claim-confidence levels and review signals perform in production — measured against analyst feedback, updated live.

Auto-refreshing every 60sLast update: 1:59:13 AM

Building baseline

Calibration baseline in progress

We publish accuracy rates once we've collected at least 50 analyst feedback signals. Until then, we're showing progress only — to avoid sharing numbers that aren't yet statistically meaningful.

0 of 50 signals0%

Claim-confidence and review-signal calibration

Percent of findings where analyst feedback confirmed the assigned assigned claim-confidence level or review signal.

High

—

No signals yet

Medium

—

No signals yet

Low

—

No signals yet

Uncertain

—

No signals yet

Sarcasm-flagged

—

No signals yet

Module breakdown

Accuracy by AGENTPR™ module — Intelligence (VoC) and Narrator.

ModuleSignalsConfirmedAccuracy
Intelligence (VoC)00—
Narrator00—

Methodology

Every material AGENTPR™ claim carries High, Medium or Low claim confidence. Uncertain and Sarcasm-flagged are review signals, not additional confidence levels. After review, analysts mark each finding as confirmed, incorrect, or uncertain.

The confirmation rate for a label or review signal is the share of findings carrying that label that analysts subsequently confirmed. Feedback flows directly from the live platform into the confidence_feedback store and is aggregated into the calibration view this page reads from.

No manual curation, no cherry-picking — the numbers you see are the same numbers our engineering and product team see internally.

Benchmark context

High-confidence targets

Industry-grade sentiment systems typically report ≥ 85% precision on high-confidence claims. We track to the same bar.

Low-confidence by design

Low confidence triggers Human Review Recommended. Uncertain is a separate review signal, not a fourth confidence level.

Sarcasm and cultural nuance

We oversample sarcasm and code-switching cases. Sarcasm-flagged items receive Human Review Required and pause for analyst sign-off.

Report history

PeriodStatusNote
Q2 2026CurrentLive — auto-refreshing
Q1 2026ArchivedPre-launch baseline

ENGINEERING OF TRUST™

Read the full Responsible AI commitments

Seven commitments behind every brief — methodology, oversight, data handling, and the Glassbox AI Policy.

View Responsible AI