Independent methodology

Evidence before verdicts.

Product facts trace to first-party research records. Recommendations are deterministic. Unknown earns no credit. Benchmark results stay empty until real, reproducible testing exists.

Finder score

Frozen at 100 points.

Affiliate commission is absent from the data passed into the scoring engine and is never a tie-breaker.

Must-have capability match30
Pricing / budget fit20
Industry fit15
Integration compatibility15
Call-volume economics10
Language fit5
Multi-location / scale3
Setup simplicity / trial2

Hard gates

Verified false on a required feature makes a product ineligible for Best Overall. Unknown stays visible, earns zero, and creates a warning.

Tie-breakers

More verified must-haves, better integration match, lower verified cost, newer verification, then alphabetical name.

Cheapest rule

Only exact or estimated results from verified public pricing can be called cheapest. Base price alone is excluded.

AI Receptionist Benchmark 2026

20 frozen calls. Zero published results.

Same fictional business: Northstar Heating & Air. Same knowledge base, facts, booking availability, escalation number, language expectations, and scripts.

Benchmark data is intentionally empty.

0 product results exist. No winner, score, or performance claim will appear before real testing and review.

  1. 01Opening-hours factual question
  2. 02Service-area factual question
  3. 03Supported-service factual question
  4. 04Appointment booking — simple
  5. 05Appointment booking — unavailable requested time
  6. 06Reschedule existing appointment
  7. 07Cancel appointment
  8. 08Human transfer request
  9. 09Urgent no-cooling HVAC call
  10. 10Ambiguous problem description
  11. 11Pricing question where exact answer exists
  12. 12Pricing question where answer is absent
  13. 13Unknown / out-of-scope factual question
  14. 14Interruption mid-response
  15. 15Caller changes mind mid-call
  16. 16Full Spanish call
  17. 17English to Spanish language switch
  18. 18Controlled background noise
  19. 19Long caller story requiring summary
  20. 20Question + booking + SMS request

Benchmark score

100 points after real testing.

Each call stores pass, partial, fail, or not applicable; notes; audio/reference ID; latency where measurable; expected and actual behavior; and hallucination flag.

20

factual Accuracy

Scored only from retained test evidence.

15

hallucination Avoidance

Scored only from retained test evidence.

15

booking Rescheduling

Scored only from retained test evidence.

10

transfer Escalation

Scored only from retained test evidence.

10

conversational Robustness

Scored only from retained test evidence.

10

spanish Handling

Scored only from retained test evidence.

5

sms Follow Up

Scored only from retained test evidence.

5

summary Intake

Scored only from retained test evidence.

5

latency Responsiveness

Scored only from retained test evidence.

5

setup Reliability

Scored only from retained test evidence.