RESPONSIBLE AI · MEASUREMENT
Calibration Report
How AGENTPR™ claim-confidence levels and review signals perform in production — measured against analyst feedback, updated live.
Building baseline
Calibration baseline in progress
We publish accuracy rates once we've collected at least 50 analyst feedback signals. Until then, we're showing progress only — to avoid sharing numbers that aren't yet statistically meaningful.
Claim-confidence and review-signal calibration
Percent of findings where analyst feedback confirmed the assigned assigned claim-confidence level or review signal.
High
—
No signals yet
Medium
—
No signals yet
Low
—
No signals yet
Uncertain
—
No signals yet
Sarcasm-flagged
—
No signals yet
Module breakdown
Accuracy by AGENTPR™ module — Intelligence (VoC) and Narrator.
| Module | Signals | Confirmed | Accuracy |
|---|---|---|---|
| Intelligence (VoC) | 0 | 0 | — |
| Narrator | 0 | 0 | — |
Methodology
Every material AGENTPR™ claim carries High, Medium or Low claim confidence. Uncertain and Sarcasm-flagged are review signals, not additional confidence levels. After review, analysts mark each finding as confirmed, incorrect, or uncertain.
The confirmation rate for a label or review signal is the share of findings carrying that label that analysts subsequently confirmed. Feedback flows directly from the live platform into the confidence_feedback store and is aggregated into the calibration view this page reads from.
No manual curation, no cherry-picking — the numbers you see are the same numbers our engineering and product team see internally.
Benchmark context
High-confidence targets
Industry-grade sentiment systems typically report ≥ 85% precision on high-confidence claims. We track to the same bar.
Low-confidence by design
Low confidence triggers Human Review Recommended. Uncertain is a separate review signal, not a fourth confidence level.
Sarcasm and cultural nuance
We oversample sarcasm and code-switching cases. Sarcasm-flagged items receive Human Review Required and pause for analyst sign-off.
Report history
| Period | Status | Note |
|---|---|---|
| Q2 2026 | Current | Live — auto-refreshing |
| Q1 2026 | Archived | Pre-launch baseline |
ENGINEERING OF TRUST™
Read the full Responsible AI commitments
Seven commitments behind every brief — methodology, oversight, data handling, and the Glassbox AI Policy.