HXR All evaluations
Menu

ChatGPT Work Can Stop Mid-Task Without Telling You What Finished

The documented ChatGPT reliability pattern now includes direct evidence that Work’s elapsed “Working for …” timer can remain active after the underlying task has stopped. In an October 8 desktop occurrence, the UI displayed “Working for 3h 43m 52s” while the assistant’s same-thread status report explicitly said no task was currently running in the background. This demonstrates that the timer is not an authoritative execution-state indicator.

What this page means: Evaluations examine how a product, organization, or practice affects people. Cases and other evidence can establish strengths, difficulties, mixed outcomes, and useful improvements. Counts describe the record; they do not establish prevalence or responsibility.
Evaluation ID
HXR-EVAL-0053
Status
Open — gathering evidence
HXR Incident connections
0
Last updated
October 08, 2026

Evaluation scope

ChatGPT Work and similar multi-step tool-mediated workflows after task execution begins: partial external actions, status visibility, interruption/error transitions, terminal accounting, checkpointing, verification, and safe resume. Excludes model-answer quality, unsupported population prevalence, unsupported root-cause attribution, and the pre-first-assistant-event failed-turn boundary evaluated separately in HXR-EVAL-0025.

Question being evaluated

When ChatGPT accepts and begins a multi-step agentic task that performs tool or external-system actions, does it preserve and surface a durable execution ledger so that interruption or failure ends in a clear terminal state identifying completed, failed, unverified, and remaining work without requiring the user to detect the stall and prompt for recovery?

Current findings

What works well or deserves recognition

The task's completion could be independently verified from saved report and Gmail state even while the UI remained stale. That means a durable completion signal existed and could potentially be synchronized to the client automatically.

Difficulties and opportunities to improve

The user cannot infer from the elapsed Work timer whether execution is actually active. The timer can continue after the task has stopped, forcing manual status interrogation and making it impossible to infer from the UI whether paid usage should still be accruing.

Mixed findings and limits

The evidence establishes task-state/timer divergence but not post-stop charging. It does not show the exact execution stop timestamp or whether allowance metering continued during the stale timer interval.

Organization information and response status

OpenAI support remains in case 16205320 / original 15955492. On October 8 HXR sent one consolidated update rather than another piecemeal screenshot message. It supplied the clean-reinstall result, current app/device details, cross-network diagnostics, ordinary-Chat continuation-without-progress, the post-reinstall Work share URL, and representative screenshots of distinct stopped segments, and requested backend task/run correlation and engineering disposition.

How HXR would test performance

Run representative multi-step tasks with consequential external writes and inject failures after some actions commit. The product should automatically expose a terminal or resumable state within a documented timeout; distinguish Running, Interrupted, Failed, and Complete; accurately categorize completed-and-verified, completed/unverified, failed, and remaining actions against independent external readback; survive reload or reconnect; offer checkpoint-aware resume without duplicate writes; handle deliberate user interruption with the same partial-state accounting; and require no extra user prompt to learn that execution stopped or to obtain the ledger.

Evidence coverage and limits

Evidence now extends through October 8 and includes ordinary-Chat resume-without-progress testimony alongside the earlier direct Work/Chat screenshots of interrupted streams, Reasoning failed, stale Working, unstable Retry state, post-failure execution, missing terminal reports, and fresh-thread recovery failure.

Next observation or verification

Ask OpenAI to correlate the documented Work run’s actual execution stop time, elapsed-timer state, and usage/metering timeline, and explain which event is supposed to stop the timer.

Revision history — Public Revision 17

Added direct desktop evidence that the Work elapsed timer continued to 3h43m52s after the assistant reported no task was running.

Technical record details
Evaluation type
Mechanism
Benchmark posture
Developing Evidence
Blue Score readiness
Not Assessed
Reform / positive practice
Reform Proposed
Public revision
17

Help test this Evaluation

Add your experience

Your experience may support, challenge, update, or add a positive example to this Evaluation. A short account is enough to begin; screenshots and files are optional.

Include this Evaluation number in your submission: HXR-EVAL-0053.

Open the submission form

The form opens on Jotform and sends your material to an HXR operator by email. It does not automatically create a case, publish anything, or contact the organization.

Remove passwords, authentication codes, payment details, and unnecessary personal or confidential information. Describe sensitive evidence rather than uploading it here.

A different experience? Share a new topic.