For its fourth annual developer conference, OpenAI presented dots, its personal agents, new working tools and a less expensive model. While these advances expand what users can delegate to AI, they also make more pressing an economic question: how to finance computing consumption that increases with product capabilities? The reorganization of subscriptions provides a first response, by further controlling volumes and charging for speed.
On September 29, in San Francisco, OpenAI held its fourth major mass for developers. More than twenty announcements, according to the group, to further install its models in daily work: software development, collaboration, connected applications and personal agents. With Dots, the company gives concrete form to this ambition; the user can entrust a mission to an agent who continues it between two conversations, including when their computer is turned off.
Powered by GPT-6 Astra, these agents have their own cloud computer and browser. They can search for information, prepare documents or work on code, then return to their user when a decision requires their intervention. The movement initiated with Operator and the automation of online tasks thus takes on another dimension.
This continuity has a counterpart, each mission can trigger a succession of operations of which the user only sees the result. As OpenAI reduces the human time needed to initiate and track a job, the amount of compute consumed by a single subscriber increases. Behind the new features of DevDay, the new commercial grid organizes the management of this consumption.
An invoice that had already caught up with the package
As early as January 2025, a few weeks after the launch of ChatGPT Pro, Sam Altman declared that OpenAI was losing money on this subscription. Customers were using it much more than expected, he explained on X, in comments reported by The Verge. A high price was therefore not enough to guarantee the profitability of intensive use.
Financial data reported since gives a measure of this pressure, OpenAI told its investors that its inference expenses had quadrupled in 2025. Its adjusted gross margin would have fallen from 40% in 2024 to 33% in 2025.
The succession of models is also accompanied by a change in the way of producing a response. With o1, from 2024, OpenAI explained that performance improved when the model had more calculation to reason with at the time of its use. Agents extend this logic by multiplying the steps: consult a source, interpret a document, execute a program, check the result and start again if necessary.
In a software development example, requesting a fix can thus result in reading multiple files, modifying the code, running tests and making new fixes. A short instruction can trigger a long job. The monthly plan absorbs this variability until its limits become visible.
More efficient models, more ambitious uses
This development does not imply that each new generation costs more to accomplish exactly the same task. The efficiency gains can be significant; OpenAI also highlights the ability of its new models to obtain better results with fewer resources on certain evaluations.
But the user adapts his requests to the possibilities of the product. A summary becomes a documented study, code assistance becomes the development of a feature, a one-off operation becomes recurring monitoring. Thus a reduction in the cost per task can be absorbed by the increase in the number of tasks and their complexity.
Dots reinforce this dynamic by allowing work to continue between exchanges. Their permanent availability does not mean that they calculate without interruption, but it does increase the opportunities to launch, resume and extend a mission.
For OpenAI, the challenge therefore consists of maintaining a legible subscription while controlling variable expenses.
The Pro 200 keeps its price, but reduces its promise
The most noticeable change concerns historical subscribers. Thus the allocation of Pro 200 in Codex and ChatGPT Work increases from twenty to ten times that of Plus. In ChatGPT, the GPT-6 Pro cap drops from 200 to 100 messages per week. OpenAI maintains the old allocation of eligible subscribers until October 29, 2026, before they switch to the new regime.
At the top of the range appears Pro 500, with an announced allocation of twenty-five times that of Plus and access to Ultrafast. American prices represent approximately 176 euros per month for Pro 200 and 440 euros for Pro 500. These amounts do not constitute the French selling prices.
The former top of the range thus becomes an intermediate level. OpenAI asking for a higher contribution from users who wish to have more capacity, while reducing what the historic package finances. This reorganization brings the subscription closer to models combining flat rate and usage pricing.
The group officially justifies the reduction in allocations by the increasing efficiency of its models. The argument deserves to be assessed at the task scale, because although a reduced allocation may be sufficient if each operation consumes less, it becomes more restrictive if the user continues complex work or multiplies missions.
Speed becomes an expense to be arbitrated
Ultrafast makes this arbitration particularly visible. OpenAI announces up to 300 tokens per second in Codex, generating up to eight times faster than standard mode for GPT-6 Astra. In return, the documentation specifies that this mode consumes the included allocation eight times faster. With purchased credits, the billing multiplier is six.
These figures relate to distinct mechanisms: generation speed, package billing and credit billing.
Sol, or the interest of reserving power for the right tasks
The launch of GPT-6.1 Sol sheds light on the other side of this policy. OpenAI presents the model as close to Astra in several software development, computer use, and professional work assessments, with standard API rates five times lower for uncached incoming and outgoing tokens.
This price difference does not reveal the internal costs of the laboratory and does not guarantee identical performance in all professions. It nevertheless provides a basis for a more economical organization of uses: entrusting current operations to a less expensive model and reserving the most powerful for the stages where it brings a measurable gain.
A processing chain could, for example, use a first model to extract information and prepare a document, then another to examine ambiguities or resolve a difficulty. The interest depends on the result because a less expensive model which increases the number of errors, rework and human verifications can ultimately increase the bill.
The choice of model therefore becomes a production decision. This is also the challenge of the AI routing and billing infrastructures that Stripe is developing: determining which resources to commit to obtain a given level of quality. For users, part of this arbitration could gradually be taken care of by the products themselves.
More entry doors, a common allocation
At the same time, OpenAI is expanding the places where this capacity can be consumed. After integrating third-party applications into ChatGPT, the group allows, with “Sign in with ChatGPT”, to mobilize its allocation from sixteen announced partners, including Notion and Devin de Cognition. Eligible uses are deducted from the plan limits.
The same distinction appears with Dots: conversations with the agent are not counted against the ChatGPT limits, but the tasks it launches in Work and Codex consume the corresponding allocations. At launch, Dots for Pro offers also remain unavailable in the European Economic Area, the United Kingdom and Switzerland
For companies, this development invites us to measure the cost of a successful task, including rework and control. How much does a validated analysis, a usable proposal or an accepted correction cost? The next advance expected from agents will also be knowing what power to engage, when to ask for help and when to stop. Delegating work will increasingly require assigning a budget to it.