ChatGPT Work Can Stop Mid-Task Without Telling You What Finished
The documented ChatGPT reliability pattern now includes direct evidence that Work’s elapsed “Working for …” timer can remain active after the underlying task has stopped. In an October 8 desktop occurrence, the UI displayed “Working for 3h 43m 52s” while the assistant’s same-thread status report explicitly said no task was currently running in the background. This demonstrates that the timer is not an authoritative execution-state indicator.
- Evaluation ID
- HXR-EVAL-0053
- Status
- Open — gathering evidence
- HXR Incident connections
- 0
- Last updated
- October 08, 2026
Evaluation scope
ChatGPT Work and similar multi-step tool-mediated workflows after task execution begins: partial external actions, status visibility, interruption/error transitions, terminal accounting, checkpointing, verification, and safe resume. Excludes model-answer quality, unsupported population prevalence, unsupported root-cause attribution, and the pre-first-assistant-event failed-turn boundary evaluated separately in HXR-EVAL-0025.
Question being evaluated
When ChatGPT accepts and begins a multi-step agentic task that performs tool or external-system actions, does it preserve and surface a durable execution ledger so that interruption or failure ends in a clear terminal state identifying completed, failed, unverified, and remaining work without requiring the user to detect the stall and prompt for recovery?
Current findings
What works well or deserves recognition
The task's completion could be independently verified from saved report and Gmail state even while the UI remained stale. That means a durable completion signal existed and could potentially be synchronized to the client automatically.
Difficulties and opportunities to improve
The user cannot infer from the elapsed Work timer whether execution is actually active. The timer can continue after the task has stopped, forcing manual status interrogation and making it impossible to infer from the UI whether paid usage should still be accruing.
Mixed findings and limits
The evidence establishes task-state/timer divergence but not post-stop charging. It does not show the exact execution stop timestamp or whether allowance metering continued during the stale timer interval.
Organization information and response status
OpenAI support remains in case 16205320 / original 15955492. On October 8 HXR sent one consolidated update rather than another piecemeal screenshot message. It supplied the clean-reinstall result, current app/device details, cross-network diagnostics, ordinary-Chat continuation-without-progress, the post-reinstall Work share URL, and representative screenshots of distinct stopped segments, and requested backend task/run correlation and engineering disposition.
How HXR would test performance
Run representative multi-step tasks with consequential external writes and inject failures after some actions commit. The product should automatically expose a terminal or resumable state within a documented timeout; distinguish Running, Interrupted, Failed, and Complete; accurately categorize completed-and-verified, completed/unverified, failed, and remaining actions against independent external readback; survive reload or reconnect; offer checkpoint-aware resume without duplicate writes; handle deliberate user interruption with the same partial-state accounting; and require no extra user prompt to learn that execution stopped or to obtain the ledger.
Evidence coverage and limits
Evidence now extends through October 8 and includes ordinary-Chat resume-without-progress testimony alongside the earlier direct Work/Chat screenshots of interrupted streams, Reasoning failed, stale Working, unstable Retry state, post-failure execution, missing terminal reports, and fresh-thread recovery failure.
Next observation or verification
Ask OpenAI to correlate the documented Work run’s actual execution stop time, elapsed-timer state, and usage/metering timeline, and explain which event is supposed to stop the timer.
Revision history — Public Revision 17
Added direct desktop evidence that the Work elapsed timer continued to 3h43m52s after the assistant reported no task was running.
Technical record details
- Evaluation type
- Mechanism
- Benchmark posture
- Developing Evidence
- Blue Score readiness
- Not Assessed
- Reform / positive practice
- Reform Proposed
- Public revision
- 17