In this briefing
- 01Activity volume is visible; productivity needs a different body of evidence
- 02Define the complete task before discussing tool coverage
- 03Human intervention is part of the cost and risk, not noise
- 04The control plane is catching up, but a feature list is not operational proof
- →What to watch next
- ↗Sources and verification
- OpenAI's internal data shows that Agent use can substantially expand research activity, but experiment counts, code volume and runtime are mainly activity proxies. Its success-rate sample covers only tasks for which a ground-truth outcome could be found, and excludes uncertain classifications and small-sample points, so 3.1 Agent workdays cannot be converted directly into 3.1 times the productivity.
- Cohere's analysis of 696,291 public MCP tools found that, under a strict mapping based on tool names and descriptions, only 2.6% were classified as able to execute one O*NET work task from end to end. Tool counts, interface coverage and complete business outcomes are three different layers, so enterprises must first define task units that can be accepted.
- LiteLLM v1.100.0 brings upload entry points, caller credentials, MCP sessions, shared budgets and routing costs into the control plane, but release notes are not evidence of production outcomes. A real operating ledger must also record human intervention, rework, failures and the results that are ultimately accepted.
Activity volume is visible; productivity needs a different body of evidence
OpenAI's internal research measurement, published on 6 September, says that by mid-August the research organisation generated 3.1 standard 8-hour Agent workdays of runtime for every human workday. At API list prices, the median researcher used more than $600 worth of inference a day, while the P90 researcher exceeded $7,000; experiments per active experimenter also reached their highest level since 2025-01. OpenAI notes that experiment volume correlates with Codex use, but available compute grew at the same time, so the observations cannot establish causation.
These figures show the scale of activity that can run in parallel, not that business output has increased by the same multiple. Code commits, experiment counts, token consumption and Agent occupancy are easy to measure, yet they can include failed attempts, repeated runs and explorations that did not produce a decision. OpenAI calls the measurements preliminary; success rates cover only tasks for which a ground-truth outcome could be found, and the charts omit uncertain classifications and points with fewer than 50 sessions or unique users. If an enterprise copies only runtime or call-cost metrics, it may record inputs precisely while still not showing which results were adopted.
Enterprises should separate activity metrics from outcome metrics. The former support capacity, budget and queue management; the latter should record at least acceptance-pass rates, rework counts, human-review time, whether a decision was adopted and the cycle from task start to usable result. Only when the two sets can be linked may growth in Agent use be interpreted as a change in productivity.
Define the complete task before discussing tool coverage
OpenAI applies Epoch AI's framework to divide AI research and development into six stages: deciding, designing, building, running, analysing and communicating. It says that high-level planning currently accounts for only a small share of Agent activity. Epoch AI's original project lists more than 60 AI research and development tasks, then rates maturity from 0 to 5, ranging from not used to autonomous execution; the authors explicitly describe the classification as preliminary, subjective and incomplete. Its value is not to give every enterprise a fixed template, but to remind teams that the single label 'doing research' contains work units with very different evidence, tools, risks and acceptance methods.
For its 2026-05 collection, Cohere gathered 696,291 tools from 123,069 MCP server listings across seven public directories. For each tool, the team used its name and description, together with a language model, to judge whether it executed the nearest O*NET task. Under this strict standard, only 2.6% were classified as independently completing a whole work task, about one in forty, while 419 occupations showed no corresponding activity. The researchers also stress that public directories do not include bespoke internal enterprise servers, making 2.6% closer to a floor for the visible scope, and that the analysis does not demonstrate actual use or reliable execution. Many tools are fine-grained actions such as querying, uploading or converting; a rich interface does not mean that complete work has been automated.
Procurement and platform teams should first create their own task register: what the input is, which systems may be called, what constitutes completion, who can accept the result and how failures are rolled back. They can then map each tool to a task step, rather than use server, tool or connector counts as a substitute for business coverage.
Human intervention is part of the cost and risk, not noise
Among the successful tasks in OpenAI's statistics over the past six months, more than half of those estimated to require a skilled human 4 to 8 hours involved at least one human intervention. Intervention means that an apparently long-horizon task still relied on a researcher for direction at a critical point, although the source does not break down the reasons or time spent. The proportion covers only internal tasks for which a ground-truth outcome could be found and cannot be extrapolated directly to contract review, customer service or production operations. It is nevertheless enough to show that task difficulty, Agent runtime and human time saved are three different measures.
A more robust measurement method treats each human touchpoint as a first-class event: record the state before intervention, its reason and duration, what changed and whether the final result improved. Open-ended tasks with uncertain outcomes should not be silently removed; they can be recorded separately as 'expert judgement required', with adopted, rejected and awaiting validation as distinct states. This avoids overstating success when automated scoring is difficult and avoids misclassifying necessary human judgement as a system failure.
Business owners can set a human-machine boundary for each task type: which steps may run unattended, which require review by two people, and which outputs can only be recommendations. The budget model should include inference charges, human intervention, quality validation and rework, using the total cost per accepted result as the basis for comparing models and processes.
The control plane is catching up, but a feature list is not operational proof
LiteLLM v1.100.0, released on 6 September, adds or fixes several operational controls. A vector-store upload entry point can enforce security restrictions; credential-less Vertex passthrough no longer leaks the caller's virtual key, and forwarded credentials are no longer copied into retry breadcrumbs. MCP gateway session tokens support RS256 signing and RFC 7662 introspection, while shared budgets can be bound to model-access groups and spend recorded by time window. The release also counts auto-router classifier cost in savings and benchmarks, and fixes a case in which unmasked personal information was still forwarded in Lakera monitor mode.
These changes belong to the same operating problem as outcome measurement: an enterprise must know who used which identity and tools, when, at what cost, whether boundaries were crossed and whether the output ultimately passed business acceptance. Release notes, however, can only show that functionality entered a version; they cannot demonstrate secure defaults, a regression-free upgrade or effective controls in production. A release containing many changes still needs isolated validation of permissions, session expiry, concurrent budgets, log redaction and rollback paths.
A workable Agent ledger should connect four classes of records: the task definition, each run and its cost, human interventions and final acceptance. Platform controls provide identity, permission, budget and traceability, while business quality gates judge results. Missing either side can leave an organisation with rising call volumes but no clear account of value or risk.
What to watch next
- Whether OpenAI continues to publish results broken down by task stage, difficulty and type of human intervention, and adds the scale of excluded uncertain tasks.
- How complete-work-task coverage, actual usage, failure rates and permission incidents change when internal enterprise MCP servers are included.
- Whether LiteLLM v1.100.0's shared-budget, MCP session and credential-separation capabilities pass independent validation under high concurrency, key rotation and rollback scenarios.
Sources and verification
Golden Data has edited this briefing from the public materials listed above. The original sources govern facts and figures. The enterprise relevance sections are Golden Data editorial analysis and do not constitute an endorsement of any third-party product.
← Back to AI Daily Briefing