Thibault Sottiaux on Dots, Codex Cloud, and achieving a 10x computer-use speed improvement in one year
Key Points
- OpenAI has achieved a 10x speed improvement in computer-use over the past year through reduced model reasoning, parallel safety monitoring, and harness optimization, enabling roughly 15 actions per second.
- Dots, OpenAI's always-on personal agent, removes the model picker entirely and modulates its own computational effort based on user constraints like deadlines and budgets.
- The Codex harness underpinning Dots operates as a continuous feedback loop with model capabilities, shedding complexity as models improve rather than adding it.
Summary
Thibault Sottiaux on Dots, Codex, and the 10x computer-use speed jump
Thibault Sottiaux leads AI agent development at OpenAI and was among the most closely watched figures on the DevDay floor — his X account, as an aside near the end of the conversation, draws enough public attention that he has to actively manage real-time feedback at a scale few product leaders at any company have faced.
The headline product he is most focused on is Dots, OpenAI's always-on personal agent. The core pitch is behavioral rather than functional: Dots never stops, occasionally sleeps, does proactive research in the background, and adapts to user preferences over time through natural feedback. Sottiaux says that after a few days of use, the experience tips into something that feels indispensable. The investment required from the user is real, though — Dots needs to be taught before it delivers.
“I do think we have seen a 10x for one year speed up. And it's not that we made the model sample faster... When you think about the model it's like how much does it need to think before it can take the next action reliably... We have a second agent which we call Guardian which watches over this primary agent. So we also made this auto review model faster.”
Harness and model
The Codex harness underpins both Dots and OpenAI's managed agents API. Sottiaux describes the relationship between harness and model as a continuous dance: teams build the harness slightly ahead of model capabilities to compensate for gaps, then the next model generation catches up and parts of the harness have to be deleted because they're holding the model back. Right now the harness is primarily a hardened, safety-focused layer rather than a complexity multiplier, because current models are capable enough that tooling complexity is being stripped out rather than added.
Specialist Dots inside OpenAI are provisioned exactly like employees — assigned identities, IAM roles, and dedicated hardware (Mac minis). One internal example Sottiaux gives is procurement. The shared memory system that supports them is designed to persist preferences and institutional context across the organization, so a team can write standing instructions ("we're a Rails shop, here's the documentation") and the agent reads and retains them rather than needing to be re-briefed.
Computer-use speed
Sottiaux says OpenAI has achieved a 10x speed improvement in computer-use over the past year, and it did not come from making the model sample faster. Three levers drove it: reducing how much the model needs to think before taking each action reliably; accelerating Guardian, a second agent that runs in parallel to watch the primary agent and intervene on unsafe actions (for example, flagging a domain mismatch before sensitive information is entered); and harness-level improvements, including parallelization and removing stalls, informed by running Astra at scale.
On raw throughput today, Ultrafast on Astra runs at 300 tokens per second. Because roughly 20 tokens are required per action, that translates to around 15 actions per second in practice. The newly announced Decisions API tightens this further, enabling one decision approximately every 200 milliseconds. Sottiaux expects real-time computer-use actions within the next year and believes a stock OpenAI model with computer-use will beat OpenAI's original Dota competition model — the ten-year anniversary of which falls this year.
Model picker obsolescence
Sottiaux argues the model picker is a transitional interface. The right analogy, in his framing, is how humans communicate deadline and budget constraints in natural language — "I need this by Monday" or "I have $100 to spend on this" — and the system should allocate resources accordingly without requiring the user to select a model tier. Dots already removes the picker entirely; the agent modulates its own effort based on the ask. Sottiaux says users will eventually find it hard to believe a model picker ever existed.
Every deal, every interview. 5 minutes.
TBPN Digest delivers summaries of the latest fundraises, interviews and tech news from TBPN, every weekday.