A Jury, Not a Judge: How to Make AI Lead Scoring Reliable
Key takeaways
- 87% of sales organizations now use some form of AI, according to Salesforce's 2026 State of Sales report.
- A panel of three models from different families agreed with human experts up to 22% more than a single large AI judge (Verga et al., 2024).
- When language models state their own confidence, it separates right answers from wrong ones at 0.52 to 0.61, where 0.5 is a coin flip (Xiong et al., ICLR 2024).
- Language models score their own outputs higher than other models' outputs, even when humans rate them as equal (Panickssery et al., NeurIPS 2024).
- Sending only the cases where models disagree to a human cut manual review time by 40 to 90% in a 2026 study.
Why is AI lead scoring only as good as the model behind it?
Because most AI lead scoring is one model's opinion. A single model reads a company, picks a number, and that number decides who gets a call this week.
AI is now part of how most sales teams work. Salesforce's 2026 State of Sales report, a survey of 4,050 sales professionals across 22 countries, found 87% of sales organizations use some form of AI. AI scoring sits next to firmographic filters, intent data and engagement scores as one more way to qualify leads.
The question is no longer whether to use AI scoring. It is how much to trust it. And the honest answer, for a score that comes from one model, is less than the dashboard suggests.
What goes wrong when one model scores your leads?
Three things, each documented in published research.
It can give two answers to the same question. A 2025 study, "Measuring Determinism in Large Language Models," ran identical prompts five times with randomness switched off. Responses still varied. On the least stable question, the correlation between repeated runs fell as low as 0.53. A lead that scores Strong on Monday can score Possible on Tuesday with nothing changed.
It can't tell when it's wrong. Researchers at ICLR 2024 asked language models to state their confidence alongside their answers. The models leaned overconfident, and their stated confidence separated right answers from wrong ones at 0.52 to 0.61, on a scale where 0.5 is a coin flip. A "confidence: high" badge from the same model that made the call is the model marking its own homework.
It has favourites. A NeurIPS 2024 paper found that language models score their own outputs higher than other models' outputs, even when human reviewers rate them as equal. In a separate 2024 study, a single AI judge ranked a model from its own family 2nd when its true position was 4th. One model brings one set of blind spots, and it brings them to every lead on your list.
None of this makes AI scoring useless. It makes single model scoring fragile.
What is a jury of models, and does it actually work better?
A jury replaces one large model with several models from different AI labs, each scoring independently, with the results combined. Different labs train on different data with different methods, so their blind spots don't line up.
The clearest evidence comes from "Replacing Judges with Juries," published by researchers in 2024. They compared a panel of three smaller models from three different families against a single large model acting as judge, across six datasets. The panel came out ahead.
| Measure | Single large AI judge | Jury of three models |
|---|---|---|
| Agreement with human experts, Natural Questions (Cohen's kappa) | 0.627 | 0.763 |
| Agreement with human experts, TriviaQA | 0.841 | 0.906 |
| Agreement with human experts, HotpotQA | 0.830 | 0.867 |
| Correlation with human rankings, Chatbot Arena (Pearson) | 0.817 | 0.917 |
| Bias toward its own model family | Present | Lower |
On Natural Questions, the jury's agreement with human experts was 22% higher than the single judge's. It beat the single judge on every test. It also had the smallest spread in scores of any judge tested, a standard deviation of 2.2, which means it was the most consistent.
The principle is old. It is why courts use juries, why hiring panels beat single interviewers, and why forecasting tournaments average many forecasters. Independent judgements, combined, cancel out individual error.
How does Sideroom score with a jury?
Every company on a Sideroom list is scored by three models from three different AI labs. Each model reads the same evidence and rates the company against your ICP. None of them sees the others' answers.
The verdicts are then combined with a simple rule:
| Jury result | What happens |
|---|---|
| 3 of 3 agree | The score stands |
| 2 of 3 agree | The card is flagged so you can see the split |
| All 3 differ | A person reviews it before it reaches your list |
Disagreement is the most useful signal in the system. It shows exactly where a human should look. A 2026 study of AI grading found the same thing: sending only the cases where repeated model answers disagreed to a human reviewer cut manual review time by 40 to 90%, and ensembles of three models were the most reliable configuration tested. Human attention goes where the models are unsure, not spread thinly across every row.
Why does the evidence need checking before the jury sits?
Because a jury is only as good as the case put in front of it. Three models agreeing on bad data is still a bad score.
Gartner estimates poor data quality costs organizations at least $12.9 million a year on average. Sales teams feel it directly. Salesforce found 74% of sales professionals are focused on data cleansing, and high performing teams prioritize data hygiene at 79% against 54% among underperformers.
So before any model scores a company, Sideroom checks the inputs. Our own checks gather the best data available from many sources, then test it for relevance, for correctness and for whether it reflects the latest news. A company that changed its name, entered a new market or made a hire last month should be scored on who it is now.
How do you know how much to trust a score?
Every score on Sideroom carries an evidence grade. We never ask a model how confident it is, because the research above shows that number is unreliable. We grade the evidence itself, using rules that give the same grade for the same inputs every time.
| Evidence grade | What it means |
|---|---|
| Very high | Multiple sources, verified and in agreement |
| High | Verified, fewer sources |
| Medium | Partly verified |
| Low | Limited evidence |
Put the two verdicts together and the score becomes a decision:
| Jury verdict | Evidence grade | What to do |
|---|---|---|
| Strong fit, unanimous | Very high | Book the meeting |
| Strong fit | Low | Spend ten minutes checking before you reach out |
| Split jury | Any | A person has already reviewed it |
Most scoring tools give you one number and ask you to trust it. A jury verdict plus an evidence grade tells you what the score is, how many independent judges agreed, and how solid the case behind it was.
Does this replace rules based and predictive lead scoring?
No. It makes the AI part of your scoring trustworthy enough to sit alongside them.
Rules based scoring is transparent but rigid. Predictive scoring trained on your own outcomes is powerful once you have the data: a 2025 study in Frontiers in Artificial Intelligence trained models on 16,600 CRM records from a B2B software company, and a gradient boosting model reached a ROC AUC of 0.989. But predictive models need a history of wins and losses on accounts like the ones you are scoring.
AI scoring fills the gap where that history doesn't exist: new markets, new segments, and especially a room full of companies you have never sold to. That is exactly where a single model's blind spots do the most damage, and where a jury and an evidence grade matter most.
Where does this matter most?
Anywhere the cost of chasing the wrong lead is high and the list is long. Conferences are the extreme case. One event can put thousands of companies in a room, and a sales team has a few days and a fixed number of meeting slots. If the scoring is wrong, the slots go to the wrong people and there is no second chance until next year.
That is the problem Sideroom was built for. We score the full room against your ICP before you land, with the evidence checked, a jury on every company, and a grade on every score. For how to define the ICP the jury scores against, see how to build a conference ICP. For the research behind a strong shortlist, see how to research conference attendees before you go.
See how Sideroom scores your next event before you land
Sources
- Salesforce, State of Sales report, 2026: https://www.salesforce.com/news/stories/state-of-sales-report-announcement-2026/
- Verga et al., "Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models," 2024: https://arxiv.org/abs/2404.18796
- Xiong et al., "Can LLMs Express Their Uncertainty?", ICLR 2024: https://arxiv.org/abs/2306.13063
- Panickssery et al., "LLM Evaluators Recognize and Favor Their Own Generations," NeurIPS 2024: https://arxiv.org/abs/2404.13076
- "Measuring Determinism in Large Language Models," 2025: https://arxiv.org/abs/2502.20747
- SURE, LLM grading with self consistency and selective human review, 2026: https://www.mdpi.com/2504-4990/8/3/74
- Gartner, Data Quality: https://www.gartner.com/en/data-analytics/topics/data-quality
- "The relevance of lead prioritization: a B2B lead scoring model based on machine learning," Frontiers in Artificial Intelligence, 2025: https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2025.1554325/full
FAQ
What is AI lead scoring?
AI lead scoring uses a language model to read information about a company or contact and rate how well it fits your ideal customer profile. It works alongside rules based scoring and predictive scoring trained on your own CRM outcomes.
Why not just use the most powerful single model?
A single model, however capable, has consistent blind spots and favours outputs that resemble its own. Research published in 2024 found a panel of three smaller models from different labs agreed with human experts more often than one large model judging alone, and showed less bias.
What happens when the models disagree?
At Sideroom, a two to one split is flagged on the company card, and a three way split is reviewed by a person before it reaches your list. Disagreement shows where human judgement is needed most.
Can I trust an AI model's own confidence score?
Not on its own. Research at ICLR 2024 found that when models state their confidence, it separates right answers from wrong ones barely better than chance. Sideroom grades the evidence behind each score instead, using fixed rules.
Does a jury of models make scoring slower?
The three models run in parallel, so a jury adds little time to scoring. The scores are ready before your team needs them.
Is AI scoring a replacement for predictive lead scoring?
No. Predictive scoring works best when you have a history of wins and losses on similar accounts. AI scoring with a jury is most useful where that history does not exist yet, such as a new market or a conference full of companies you have never sold to.