When ChatGPT Work Fails, How Much Paid Usage Does the Failure Consume?
A contributor reports rapid depletion of a limited paid Work allowance during a failure-heavy period, and new state evidence shows that the visible Working indicator cannot be used as a reliable proxy for whether metered execution should still be active. In one Cats Work occurrence, Working remained visible at 11:47 AM even though the saved completion report existed at 11:26. HXR does not know whether any usage accrued after completion, but the stale status prevents the user from inferring metering state from the interface.
- Evaluation ID
- HXR-EVAL-0055
- Status
- Open — gathering evidence
- HXR Incident connections
- 0
- Last updated
- September 28, 2026
Evaluation scope
ChatGPT Work usage consumed during failed, interrupted, retried, or recovery execution; user visibility into per-task and failure-attributable consumption; included allowance depletion; optional credits; and remedies when confirmed service failures materially consume usage. Excludes unsupported claims that every failed attempt is charged, exact failure-attributable usage without provider data, and ordinary successful-task consumption.
Question being evaluated
When ChatGPT Work fails, retries, repeats tool work, or requires recovery after a system-side interruption, can the user determine how much paid usage was consumed by successful work versus failure overhead, and is system-caused failed usage protected, restored, or otherwise accounted for?
Current findings
What works well or deserves recognition
OpenAI provides remaining-allowance and usage controls for supported plans, documents that Work participates in included agentic usage and optional credits, and provides a Support route for usage/balance questions. OpenAI also explicitly positions Work for longer, multi-step delegated tasks. These controls establish that usage is measurable at some level, but HXR has not verified a personal-plan per-failure breakdown sufficient to quantify failure overhead.
Difficulties and opportunities to improve
The contributor reports rapid depletion of a scarce paid allowance during a period with repeated Work failures but cannot determine the failure-attributable share. Repeated execution, retries, context reconstruction, tool readback, and manual recovery can create economic/opportunity cost if they consume included usage that could otherwise be used for productive work.
Mixed findings and limits
The stale Working label proves a status mismatch, not post-completion charging. Provider telemetry is required to determine when allowance metering actually stopped relative to durable completion.
Organization information and response status
OpenAI Support case 15955492 has been asked to preserve and review usage/metering telemetry for the documented failed Work runs, quantify failure/retry/recovery consumption where possible, state whether system-caused failures consume included allowance or purchased credits, and identify any allowance-restoration or billing-adjustment mechanism. No specialist response on usage attribution is recorded yet.
How HXR would test performance
Run equivalent multi-step Work tasks under normal completion and under injected platform/reasoning/stream failures. Measure total included-usage/credit consumption, retries, and recovery overhead. Pass if failed-run usage is identifiable, internal retries are not silently indistinguishable from productive work, confirmed service-caused waste has a documented allowance/credit remedy, and task-level records reconcile to the aggregate usage meter.
Evidence coverage and limits
Current evidence combines the contributor's percentage-level allowance report, five direct September 28 Work-reliability screenshots under HXR-EVAL-0053, and first-party OpenAI documentation describing Work/Codex agentic usage and optional credits. Exact failed-run usage records have not yet been obtained from OpenAI.
Next observation or verification
Ask OpenAI case 15955492 to correlate the task's 11:26 durable completion timestamp, 11:47 stale Working indicator, and usage/metering timeline before assigning any post-completion usage claim.
Revision history — Public Revision 4
Added stale post-completion Working-state evidence as a usage-accounting context: the UI cannot be used to infer whether metered execution has actually ended.
Technical record details
- Evaluation type
- Mechanism
- Benchmark posture
- Developing Evidence
- Blue Score readiness
- Not Assessed
- Reform / positive practice
- Reform Proposed
- Public revision
- 4