OpenAI Is Selling the Support Desk It Built for Itself
Presence resolves 75% of OpenAI's own English-language phone support without a human — and the product being sold is the guardrail layer, not the model.

OpenAI launched Presence on July 22 — an enterprise platform for building and running voice and chat agents. The proof point it led with is unusual for a launch post, because it is about OpenAI's own operations rather than a customer's: Presence handles OpenAI's English-language phone support line, where it resolves about 75% of inbound calls without a human, and it cut handoffs by 15 percentage points in 10 days.
The model underneath is GPT-5.6. But the model is not the product. Every serious lab now has a realtime voice stack good enough to hold a phone conversation. What Presence sells is the apparatus that makes a company willing to point that stack at its actual customers: policies, permissions, guardrails, escalation rules, evaluations, simulation, and telemetry — packaged and supported.
That is a meaningful shift in what OpenAI thinks it is selling.
The gap Presence is aimed at
The distance between a convincing voice agent demo and a deployed one is not capability. It is liability. A support agent that is right 95% of the time and confidently wrong 5% of the time, with authority to look up accounts and take actions, is not a 95% solution — it is an unbounded risk surface attached to a phone number your customers already have.
Enterprises have spent two years discovering this. The pilots work. The rollouts stall at the review where someone asks what the agent does when a caller is angry, or lying, or asking about something regulated, or trying to talk it into a refund it should not give.
Presence is built around that review rather than around the demo. It ships a simulation tool that stress-tests an agent against generated edge cases before it goes live, guardrails that constrain what actions the agent may take and which systems it may touch, defined escalation rules for when to hand to a human, and telemetry that watches production signals — how often humans get pulled in, where quality drops — and feeds them back. Codex, running on GPT-5.6, monitors output quality and proposes improvements.
Strip the branding and this is a QA and governance harness with a model in the middle. That is precisely the part enterprises could not build themselves in a quarter, and precisely the part that determines whether an agent ever leaves pilot.
Eating the dog food first is the argument
OpenAI's support line is a genuinely hard deployment. Callers are frustrated by definition. The product surface is enormous and changes weekly. Billing questions carry real money. Running your own agent against that traffic and publishing the resolution rate is a stronger claim than any benchmark, because the failure mode of a bad support agent is not a lower score — it is a public complaint.
The 15-point handoff reduction in 10 days is the more revealing number of the two. A 75% resolution rate is a snapshot. A steep improvement curve over a week and a half says the feedback loop — telemetry in, Codex-suggested changes out, re-evaluate — is doing something. If that loop generalizes to customer deployments, it is the actual moat: not a better voice, but a system that gets measurably better at a specific company's calls faster than a human-staffed team could be retrained.
Generalizing is the open question. OpenAI's support corpus is OpenAI's. A bank's edge cases are regulatory. An insurer's are adversarial in a different way.
The early customers say something about the ambition
Presence is in limited general availability, with OpenAI providing hands-on support during rollout. Named early customers: BBVA Mexico, SoftBank Corp., and Retail Insurance Australia, part of IAG.
The pattern there is deliberate. A retail bank, a Japanese telecom running Japanese-language agents, and an insurer — three of the most heavily regulated, most escalation-sensitive, most reputation-exposed categories of inbound call volume in existence. These are not soft launches into low-stakes chat widgets. They are the categories where contact centers are largest, most expensive, and most defended by compliance.
SoftBank's Japanese-language deployment is worth separating out. Voice quality in non-English languages has been the reliable weak point of every agent platform, and it is where incumbent contact-center vendors have kept their footing. Naming it first suggests OpenAI knows where the objection will come from.
Pricing has not been disclosed, which for a limited-GA enterprise product usually means it is still being negotiated deal by deal.
What it means for the rest of the stack
Presence is OpenAI moving down from the model layer into the application layer of the enterprise — the same move that will put it in direct competition with the contact-center platforms and the agent-orchestration startups currently building on its API. That tension is not new for this company, and it will not be the last time a partner discovers it is now a competitor.
For anyone building agents, the useful signal is what OpenAI concluded is worth productizing. Not reasoning. Not voice. Guardrails, simulation, evaluation, escalation, and a feedback loop. The lab with the best model in the building looked at why enterprise agents don't ship and decided the missing piece was everything around the model.
That is worth internalizing whether or not you ever touch Presence. The demo has been solved for a while. The review is what is left, and the review is now the product.
