Why OpenAI’s New AI Agents Are Struggling Outside the Tech Industry.
OpenAI Wants Everyone to Have an AI Agent. Getting There Means Asking How Much Control You'll Give Up.
Inside ChatGPT Work, the harness wars with Anthropic, and why coding agents don't automatically translate to the rest of the white-collar world.
98%: of OpenAI employees use Codex internally
17%: of org subscribers actually use it
1%: adoption among individual subscribers
1: The Pitch: An Agent for Every Job, Not Just Every Coder:
OpenAI is trying to do for accountants, doctors, and salespeople what coding agents already did for software engineers.
Its newest product, ChatGPT Work, is a general-purpose evolution of Codex, its coding-agent tool. Rather than writing software, Work is meant to plug into the everyday tools of white-collar work — email, Slack, calendars, cloud drives, Notion, Figma, and more — and complete multistep tasks on its own. OpenAI frames it as core to its mission of bringing AI's benefits to everyone, not just technical users.
The business logic is straightforward: agents that run longer and touch more systems consume more tokens, and reaching entire professions beyond software engineering is essential if AI labs are going to justify the scale of their infrastructure spending.
Rival approaches from vertical-specific startups — legal-focused, sales-focused, and so on — are chasing the same customers with a model-agnostic strategy, adding pressure on the large labs to prove their own harnesses are worth the lock-in.
2: Why No n-Engineers Are a Different Problem:
Getting a chatbot to write code and getting it to safely run your calendar, inbox, and file system are not the same challenge.
**OpenAI engineers described early internal use of Codex by non-technical teams as rocky **— the tool would ask communications or finance staff about source code and flag empty diffs, concepts meaningless outside engineering. The team spent months generalizing the product so it would make sense to people who had never touched a terminal.
That generalization runs into a harder problem than interface design: unlike code, which either runs or doesn't, most white-collar output — a strategy memo, a sales pitch, a client deck — has no clean pass/fail signal. OpenAI says it evaluates performance using its GDPval benchmark, built from tasks across 44 occupations, along with user feedback and internal usage patterns.
Real-world friction shows up in smaller ways too: permission flows for connecting a cloud drive proved confusing, some settings exist only on the web app and not mobile, and the tool reportedly underperforms unless users manually raise its 'effort' or reasoning setting — guidance that isn't well surfaced to newcomers yet.
"Discoverability matters in this phase, and at some point we won't have the button." — Andrew Ambrosino, lead engineer, OpenAI desktop app
3: The Real Rivalry: Model vs. Harness:
OpenAI's engineers were reluctant to compare their product to Anthropic's Claude Cowork — but the resemblance, and the history, are hard to miss.
When OpenAI first built Codex as a web app, it bet heavily on the model handling tasks with minimal user guidance. Anthropic's Claude Code took the opposite approach, walking users through options and checking in at each step — a slower but more reliable pattern that reportedly outperformed OpenAI's early design and pushed OpenAI to add more back-and-forth interaction of its own.

Anthropic’s Surprise Double Launch Directly Targets OpenAI and Google’s AI Dominance
That history feeds into an open industry debate about whether a proprietary 'harness' — the layer of tools, prompts, and guardrails wrapped around a model — is a durable advantage at all. Some engineers argue harness complexity is a short-term patch that better underlying models will eventually make unnecessary; independent benchmarking has shown open-source harnesses matching or beating both Codex and Claude on identical underlying models,suggesting the model, not the wrapper, may matter more in the long run.
Support our research
Independent analysis fueled by you.
4: What This Means for Businesses Evaluating AI Agents:
The lesson from OpenAI's own adoption numbers isn't that agentic AI doesn't work — it's that generic, DIY-configured agents ask too much of non-technical teams.
Real deployments described by OpenAI and its early users skew toward narrow, recurring, data-heavy tasks: metrics reports, dashboards, calendar and inbox cleanup, and pulling scattered information into one place. Token costs for casual use have also run well above subscription price, an economics problem enterprises will need to account for as usage scales.
For most businesses, the winning approach isn't handing every employee an open-ended agent and a stack of permission prompts to sort through — it's deploying AI that's already configured around how the business actually runs.
The Real Lesson for Enterprise Buyers.
OpenAI's own numbers show the gap between a model demo and a workflow your finance team will actually trust: near-total adoption inside a lab full of engineers, and barely a foothold everywhere else.
That gap is exactly what Agent+ is built to close. Instead of asking non-technical teams to configure permissions, wrangle plug-ins, and guess at 'effort settings,' Agent+ is built and tuned around the specific workflows a business already runs — so the agent shows up ready to work on day one, not after weeks of trial and error.
See how Agent+ turns everyday business AI into results at otherworldsai.com







