TIL comparing AI voice bots side by side changes everything
Last Tuesday I ran ElevenLabs against a cheap open source model for my customer service demo, same script and 50 calls each. The cheap one sounded fine on paper but dropped context after 3 exchanges, while the other held the thread the whole time. Has anyone else done a head to head test like that and found the gap shows up way more in live calls than in samples?
Ran the same kind of test last month with two voice bots for a booking system. The big one kept the whole conversation straight, while the cheaper one started guessing wrong dates after the second question. What really stood out though was how the cheap bot sounded GREAT in the first 30 seconds. It was only when people asked follow ups or changed their mind that it fell apart. You could almost see callers get frustrated because the voice was so smooth but the answers were so off. So yeah, sample clips can fool you. Live testing is the ONLY way to know what you actually have.