A long-document summary can sound accurate while quietly dropping the condition that changes a decision. For F2, Siftano tested whether ChatGPT and Claude could compress the same synthetic operations packet without losing corrected values, queue-level weaknesses, incident context, scope exclusions, or the boundary between estimated time saved and monetary ROI.
The result was narrower than a general product ranking:
- ChatGPT: 40/44 (90.91%)
- Claude: 43/44 (97.73%)
- both products used the corrected 6,240-ticket total rather than the superseded 6,310 figure
- both used the corrected 6.9% reopen rate rather than the superseded 7.1% figure
- both preserved the 7.2-minute / 8.4% Technical-queue weakness
- both disclosed the severity-2 privacy incident instead of describing the trial as incident-free
- neither converted the 314-hour estimate into cash savings or monetary ROI
The three-point gap came from summary completeness, not from a wrong headline number.
What we tested
F2 uses a fictional Harborline customer-support trial packet created for evaluation. It contains an aggregate success definition, preliminary values that are later corrected, queue-level metrics, a privacy incident, a survey with limited participation, manual verification overhead, and explicit scope exclusions.
The fixed prompt required two outputs:
- an executive summary of at most 250 words; and
- a 12-row key-facts table in a fixed order.
The fixture was designed to expose common compression errors:
- carrying a superseded number into the final summary
- saying every queue met a threshold when only the aggregate did
- turning zero severity-1 incidents into "no incidents"
- forgetting that customer-facing sending still required human approval
- treating a 36% survey response as representative customer satisfaction
- omitting excluded voice, non-English, and high-value refund workflows
- turning estimated staff hours into cash ROI without a labor-cost assumption
No real customer data was used.
Test conditions
| Condition | ChatGPT | Claude |
|---|---|---|
| Run date | 2026-08-30 | 2026-08-31 |
| Plan shown in UI | Pro | Not shown |
| Model/version shown in UI | Not shown | Sonnet 5 · 중간 |
| Session | Temporary Chat | Incognito / secret chat |
| Web search or research | Off | Off |
| Completed-run follow-ups | 0 | 0 |
| Completed-run retries | 0 | 0 |
Claude's execution was delayed because the product displayed a five-hour usage limit on the earlier attempt. We waited for the limit to reset and then ran the fixed fixture once. The blocked usage-limit screen was not scored as a model response.
The source SHA-256 was the same for both products:
0CDEC648D9CAF891B59628C5A56A853124659C4F580BA4B0A431BE03CA24445E
The prompt SHA-256 was also the same:
F1A1EDC069E264ACCAE93AFE4C53370C301959B73A39F8F1A847C5C7B69B1697
The 44-point rubric
The test has two components.
Key facts: 24 points
Each of 12 required rows is worth two points:
| Score | Meaning |
|---|---|
| 2 | value and material qualification are correct |
| 1 | core value is correct but an important qualification is incomplete |
| 0 | value is wrong, stale, unsupported, or omitted |
Executive summary: 20 points
Ten required concepts are worth two points each. These include the aggregate success result, corrected response and reopen metrics, the severity-2 incident, Technical queue weakness, mandatory human approval, the 314-hour estimate after 96 hours of verification overhead, scope exclusions, survey limits, and the absence of monetary ROI.
Final scores
| Component | ChatGPT | Claude |
|---|---|---|
| 12 key facts | 23/24 | 23/24 |
| Executive summary | 17/20 | 20/20 |
| Total | 40/44 | 43/44 |
Both products lost one point on F2-01. They returned the correct April 6 through May 22 date range, but the qualification cell did not fully restate that the trial was an internal English-language text-support trial.
The remaining difference was in the executive summary.
Where ChatGPT lost summary points
ChatGPT's summary correctly stated all three aggregate success conditions, the Technical queue weakness, the severity-2 privacy incident, the 314-hour net estimate, the 36% survey limitation, and the absence of monetary ROI.
Two qualifications were less complete:
- it described human review catching the privacy issue, but did not explicitly state in the executive summary that every customer-facing reply required human approval before release; and
- it did not include the trial's scope exclusions—voice calls, non-English conversations, and refunds above USD 500—in the executive summary.
Those facts were present elsewhere in ChatGPT's 12-row table, so this is not the same as inventing or misunderstanding the source. The scoring question was whether an operations leader reading only the executive summary would receive every required limitation.
What Claude preserved in the executive summary
Claude included the same corrected aggregate values and also kept the constraints that were easiest to lose during compression:
- all AI-drafted replies required human review and approval
- the Technical queue did not meet the aggregate benchmarks
- the trial was not incident-free
- 314 hours was an estimate, not payroll or cash ROI
- the survey was not used to assess aggregate success
- voice, non-English conversations, and refunds above USD 500 were outside the tested scope
That completeness accounted for the 20/20 summary score.
What both products handled correctly
Corrected values beat earlier values
Both products rejected the preliminary 6,310-ticket total after the packet removed 70 duplicate migrated tickets. Both also used the corrected 6.9% reopen rate instead of the earlier 7.1% note.
Aggregate success did not erase a weak queue
The trial met the aggregate 5.9-minute response-time and 6.9% reopen targets, but Technical was slower at 7.2 minutes and reopened at 8.4%. Both products kept that distinction.
Zero severity-1 did not become "incident-free"
One severity-2 privacy incident occurred and was caught during human review before a customer-facing reply was released. Both summaries disclosed it.
Time saved did not become money saved
The packet estimated 410 gross staff hours saved and recorded 96 hours of manual verification work, leaving a 314-hour net estimate. Because no approved labor-cost figure existed, both products refused to turn this into payroll savings or cash ROI.
What this result does not prove
This is one synthetic long-document fixture, not a general benchmark of summarization quality. It does not measure very long context windows, PDFs with complex visual layouts, open-web research, multilingual source material, or iterative prompting.
The products were also not run at the same clock time. ChatGPT ran on August 30; Claude ran after its usage limit reset on August 31. The source and prompt were fixed and identical, but product behavior can change over time.
Most importantly, 43/44 versus 40/44 is not evidence that Claude is generally the better summarizer. In this fixture, both products got the material numbers right. Claude's advantage was that its executive summary retained more of the predefined qualifications.
Practical takeaway
When evaluating a long-document summary, separate four checks:
- Were corrected values used instead of earlier drafts?
- Were aggregate results distinguished from subgroup weaknesses?
- Were incidents, exceptions, and scope limits preserved?
- Did the summary avoid converting estimates into stronger financial claims?
A summary can pass the first check and still fail the fourth.
For the source-grounded test that preceded F2, see ChatGPT vs Claude for Source-Grounded Fact Checking.
F3 tests a different failure mode: calculations and missing values in a fixed CSV rather than compression of a long narrative packet.

