Accuracy scores do not show what it is like to use an AI tool for a real deliverable. F4 therefore changed the unit of evaluation. Instead of comparing two products on one-shot answers, we used ChatGPT for a small operations workflow: upload two synthetic files, produce a decision brief, challenge a risky financial boundary, and preserve the corrected result.
The useful observations were not a single score:
- the first brief already used the corrected Harborline values and identified the missing CSV cost
- it already refused to turn 314 estimated staff hours into monetary ROI
- one correction request made the separation between the CSV cost subtotal and the time-saved estimate more explicit
- the corrected brief could be fully copied from ChatGPT's built-in editor into local Markdown
- the UI displayed Markdown, PDF, and DOCX download choices, but the controlled session did not capture a completed file-download event
- Temporary Chat created an evidence-durability problem: after tab loss, an explicitly authorized recovery rerun was required
This is a workflow review of one controlled task, not a general review of ChatGPT.
The workflow we tested
We uploaded two synthetic files: the F2 Harborline operations packet and the F3 structured CSV. The prompt requested an export-ready weekly operations decision brief with exactly five sections:
- Executive decision
- Operational performance
- Cost and data-quality limits
- Risks and exceptions
- Recommended human follow-up
It required corrected values, aggregate-versus-subgroup distinctions, explicit missing-cost handling, no invented customer satisfaction or ROI claims, and an evidence list tying major conclusions to the supplied files. No real company, customer, or financial data was used.
Test conditions
| Condition | Observed |
|---|---|
| Product | ChatGPT Web |
| Plan | Pro |
| Exact model shown | Not shown in UI |
| Session | Temporary Chat |
| Web search | Off |
| Built-in data analysis | Observed |
| Files uploaded | 2 in one multi-file action |
| Steps to first result | 4 |
| Planned correction messages | 1 |
| Export steps | 3 |
| Verified preservation method | Built-in editor full copy to local Markdown |
The F2 source SHA-256 was 0CDEC648D9CAF891B59628C5A56A853124659C4F580BA4B0A431BE03CA24445E.
The F3 CSV SHA-256 was 3DD5D56F100DA9B98A02FD363701EE30CBA0EBEBE87A3C4C3F531C5605C0B5B8.
What the first brief got right
The first preserved brief already handled several high-risk facts correctly. It used the corrected 6,240-ticket total and 6.9% overall reopen rate rather than superseded values. It separated aggregate success from the weaker Technical queue and disclosed the severity-2 privacy incident.
On the structured CSV, it calculated 1,602 successes out of 1,860 requests and did not treat Epsilon's missing February cost as zero. It called USD 1,343 a known-cost subtotal, not a complete total.
The first brief also kept the Harborline 314-hour figure in the right category: an estimated time saving, not payroll savings or monetary ROI.
Why we sent one correction
The planned correction asked ChatGPT to recheck every statement about cost completeness and monetary ROI, explicitly reminding it that one CSV cost value was missing and that estimated time saved must not be converted into cash savings without source support.
This was not a repair for a wrong arithmetic result. The first brief already stated the essential limits. The question was whether a focused correction could make the boundary clearer without breaking supported parts of the document.
What changed after the correction
The corrected export strengthened three points.
First, it stated that the CSV contained 18 records but only 17 reported cost_usd values. USD 1,343 was therefore a known-cost subtotal, not a complete total.
Second, it explicitly separated that subtotal from the Harborline 314-hour estimate. The source packet did not provide an approved labor-cost figure, and nothing established that the CSV costs were labor costs that could be applied to those hours.
Third, it made the prohibited inference concrete: no cash savings, ROI percentage, payback period, or complete total-cost conclusion could be calculated from the available data.
That is a useful correction pattern for an operations workflow: challenge one risk boundary without rewriting the whole recommendation.
A capture limitation we are not hiding
The corrected document changed on screen and a corrected editor export was preserved. However, the file copied into the f4-corrected-response.md raw slot was byte-identical to the first-response raw file.
We therefore do not use that mislabeled corrected-raw file as evidence of the correction. For the correction result, the authoritative evidence is the preserved corrected export and the correction screenshots.
This distinction matters because an evidence-backed review should not silently promote an invalid capture into primary evidence.
Preservation and export: what actually worked
The verified path was:
- open the generated brief in ChatGPT's built-in editor;
- use the full-copy control; and
- save the copied content locally as Markdown.
The editor also displayed download choices for Markdown, PDF, and DOCX. We did not capture a completed controlled download event, so this review does not claim that those file-download options completed successfully. The verified claim is narrower: full editor copy to local Markdown worked.
Friction we observed
Temporary Chat recovery
The largest workflow weakness was recoverability. A temporary/private conversation is not a durable evidence store. After the tab was lost, the evaluation needed one explicitly authorized recovery rerun to recreate a preservable session.
For auditable work, preserve the output outside the temporary conversation before closing it.
Long-response navigation
The decision brief required vertical navigation. That is ordinary for a long answer, but it adds friction when a reviewer wants to compare one clause against its source while keeping the whole document visible.
Exact wording matters for export claims
Seeing a download control is not the same as observing a completed download. The run verified full-copy preservation; it did not verify a completed download event.
Who this workflow suited in the test
The workflow was useful for a user who can review a generated operations brief, identify a risk boundary such as missing cost or unsupported ROI, send a focused correction request, and preserve the final result outside the chat session.
It is a weaker fit when the workflow expects a temporary conversation to remain recoverable without external preservation, or when a user expects the model to establish financial ROI from incomplete source data.
What this review does not prove
F4 used one product, one synthetic workflow, and one planned correction. It does not compare ChatGPT with Claude and does not measure production integrations, collaborative editing, very large files, or repeated weekly use over time.
The recovery rerun also means this was not an untouched one-shot session from beginning to end. That limitation is part of the result rather than something to erase.
Practical takeaway
For a decision-brief workflow, the most useful checks were different from a benchmark score:
- Did the first draft keep corrected values and important exceptions?
- Could a focused correction strengthen one risk boundary without introducing a new error?
- Could the corrected result be preserved outside the chat?
- What evidence remains if the temporary session disappears?
In this controlled task, ChatGPT handled the cost-and-ROI boundary carefully and supported a usable correction-and-copy loop. The main operational caution was evidence durability, not the arithmetic itself.
For the separate structured-data test, see ChatGPT vs Claude on CSV Analysis.
