Human Experience Reform

Independent public-interest evaluations of how systems respect dignity and support agency.

← Evaluations & Cases

HXR Incident Case

ChatGPT Work Stopped After Partial Execution Without Saying What Finished

The ChatGPT reliability Incident remains active across both Work and ordinary Chat. The contributor reports recurrence on home Wi‑Fi, cellular, other Wi‑Fi networks, and cellular away from home with no VPN/proxy/filtering. The most common recent symptom is silent cessation after a short burst of activity: the assistant says it will continue, then no further useful progress occurs until another user turn, even after long waits.

How to read this record: This Case preserves one bounded experience. Related Evaluations examine what it and other evidence may show about a system or practice. Testimony remains testimony; cause, prevalence, and responsibility require their own supporting evidence.
Case ID
HXR-2026-0928-0001
Status
Published — Active Evidence Collection
Reform status
Reform Proposed; no verified implementation recorded
Resolution
Open. Failures remain active through October 8 across Chat and Work and across multiple network conditions. OpenAI has requested and received additional diagnostics but has not yet provided a substantive specialist finding or verified mitigation.
Public revision
16
Last updated
October 08, 2026

Contributor objective

Delegate substantial multi-step work to ChatGPT and, if anything interrupts it, automatically know what completed, what failed, what remains, and whether it is safe to resume without redoing successful external actions.

Observed sequence

1. The contributor gives a sufficiently specified multi-step instruction and ChatGPT begins Work-mode or tool execution.
2. Some actions complete or appear to complete, including actions against external systems.
3. A later operation errors, times out, is interrupted, or stops making visible progress.
4. The interaction does not reliably transition to a terminal report separating completed, unverified, failed, and remaining work.
5. The contributor has to ask whether anything is still happening or instruct ChatGPT to continue.
6. After intervention, ChatGPT can sometimes reconstruct useful partial state and resume.
7. In the September 28 fresh-thread occurrence, the iOS stream reported interruption while canonical readback showed durable HXR records had already committed, demonstrating that visible response state and task state can diverge.

Current finding

A partially executed delegated workflow can become state-opaque at the point of failure. The user cannot reliably distinguish still running, stopped, partially completed, completed-but-unverified, or safe-to-retry without manually supervising the agent and asking it to reconstruct state.

Evidence and limits

HXR preserves five private direct September 28 reliability screenshots across two Work threads. The newest, HXR-EVID-2026-0135, is post-protocol-hardening evidence. The same-day Help My Pet intake reliability audit documents the mitigation steps and explicitly states that they did not establish or fix the underlying streaming/platform fault. This weakens a procedure-only explanation without proving one shared root cause across all failure labels.

Human impact

Friction

The user must monitor apparent progress, detect that delegated work has stopped or become ambiguous, ask for status, reconstruct what happened, and decide whether retrying will duplicate actions that may already have committed.

Reported burden or pain

The contributor repeatedly loses 30–90 minute blocks, and sometimes overnight elapsed time, waiting for work that appeared to be continuing but did not materially advance. He must return, detect the stall, and wake the thread again, often repeatedly across multiple conversations.

Actual harm

Demonstrated harm includes repeated time and attention burden, manual state verification, and at least about 21 minutes of stale Working status after the saved completion report in this occurrence. The user had to ask the system whether the batch was actually finished.

Potential harm

If the system continues external writes after displaying a failure state, a user who taps Retry or starts over may create overlapping executions or duplicate actions. If continued background execution also consumes limited paid usage, the failure can additionally create hidden economic/opportunity cost.

Invisible taxes

Progress watching, re-prompting, state reconstruction, external-system readback, duplicate-write avoidance, repeated verification, manual transcript recovery, and unquantified paid-usage consumption from retries or recovery.

Requested reform

Maintain a durable operation ledger for multi-step tool tasks and automatically surface it whenever execution completes, fails, times out, is interrupted, or loses client continuity. Categorize actions as completed-and-verified, completed or attempted but unverified, failed, and not attempted or remaining. Clearly distinguish Running, Interrupted, Failed, and Complete, and resume from the checkpoint without duplicating committed actions.

Acceptance test

Inject failures into representative multi-step tasks after some external writes commit. The user should automatically receive an accurate terminal or resumable ledger, confirmed against external readback, without typing another prompt. Reload or reconnect should preserve the state. Resume should continue from the checkpoint without repeating committed actions. Deliberate user interruption should produce the same partial-state accounting.

What you can try now

Ask ChatGPT whether it is still working, request a status summary or 'continue,' independently inspect external systems, then retry or split work into smaller prompts. OpenAI also recommends general freeze/error troubleshooting such as reload, new chat, different device/network/model, and diagnostics for Support.

Classification: Available but user-borne; manual supervision and reconciliation required

Organization response

OpenAI case 16205320 / original 15955492 remains open. HXR has now supplied the requested device and network diagnostics and clarified that recent failures include ordinary Chat continuation-without-progress, not only Work-mode streaming/state issues. A substantive human specialist analysis remains pending.

Public record details
Domain
AI assistant delegated-work reliability / task-state transparency / failure recovery
Jurisdiction
ChatGPT Work and connected external-system workflows; not geographically bounded
Publication basis
Anonymized With Contributor Approval
Framework
Current HXR canonical framework as of 2026-09-28