Call4me's 49-call test finds AI-first calls took longer to reach a human
The founder's recordings put the median at 3:07 for AI-first calls and 1:56 for menu-first calls, in a small, non-random sample that includes caller mistakes.
By Ryan Merket · Published
Primary source: call4me
Why it matters
Call4me's figures are a small, first-party sample, but the recordings expose the extra step AI callers must navigate when customer-service systems ask for account details or route callers before the human hold queue begins.

Call4me founder Nick Khami (@skeptrune) says his AI phone-calling service reached a live person more slowly when an automated assistant answered first than when a keypad or speech menu did. In a report published October 8th, Call4me put the median time to a human at 3:07 across 11 AI-first calls, versus 1:56 across 19 calls that began with a plain phone menu. The figures come from Call4me's own calls, not an independent or controlled benchmark.
Khami built Call4me to let an AI agent place business calls, navigate menus and hold queues, then report the result to the user. In the new tally, Call4me's AI caller is also the subject of the evaluation: the company says its AI made most of the calls being measured. Five were placed by users' agents for their own errands. Khami previously founded Trieve, which RuntimeWire reported was acquired by Mintlify in July 2025.
The dataset covers 49 calls to customer-service lines at 30 companies between September 28th and October 8th, 2026. Call4me says it measured from the start of each published recording to the first moment a live person spoke. Across 31 calls with a reported time to a person, the median was 2:00; 14 reached someone in under two minutes, while six took more than five. Pottery Barn was fastest at 0:49. United took 19:09 and Delta 16:15, with most of those two waits occurring after callers were already in the hold queue.
The Delta recording shows why the overall time can misstate what an AI assistant did. Call4me says Delta's AI transferred its caller at about 1:04 after asking for a confirmation number the caller did not have. The caller declined a callback and messaging, then waited on hold; a representative answered at 16:15. The recording therefore documents a long wait after the handoff, not a 16-minute conversation with the AI.
The test also includes a call from someone without account details. In one DIRECTV call, the caller said they wanted to cancel, said they had no account number and selected satellite service; a representative answered at about 2:47. A first attempt did not reach a person. That two-call example shows a route that worked once, not proof that DIRECTV's system only transfers callers who say "cancel service."
Call4me's own account complicates the comparison. Its methodology says seven of the 17 calls that did not reach a person involved at least some caller error: examples included choosing the wrong option, failing to ask for a person, staying silent when an assistant responded, abandoning a transfer or a bug on Call4me's side. Those calls remain in the tally. The result mixes the phone systems being tested with the choices and performance of the AI making the calls.
The sample has several limits. Call4me says it called each company on a single day, at whatever time the company was being written about; most companies received one call. It explicitly says the results are not a ranking of customer-service quality and do not show how waits vary by time or day. Fast routes sometimes reached sales or quote staff rather than the support team an existing customer would need. The report also says 31 calls had a published time to a human and 17 did not reach one; its methodology notes that a successful call with no reported answer time was left untimed.
The recordings also serve as product demonstrations. Call4me markets an AI caller that works through MCP connections with assistants such as Claude Code, Codex and ChatGPT. Its site lists calling at $0.25 per minute, with prepaid credits starting at $10. The customer-service report gives prospective users examples of the menus and account prompts the product is designed to handle. The recordings also show the limits of that task: some calls stall, some routes depend on what the caller says, and some waits begin only after an automated system has transferred the call.
In this sample, calls classified as AI-first had a higher median time to a human than calls classified as menu-first. The data does not establish that AI customer-service systems generally lengthen waits. The recordings show how an automated front door can add questions or account checks before a queue, while a separate AI caller attempts to get through on a customer's behalf.