Claude vs ChatGPT vs Copilot: fixing a broken budget spreadsheet
Three AI tools, one broken budget workbook, one sentence of instruction. The test was never whether they could fix it. It was whether they told me what they changed. Real task
Claude wins this test. It was not the smartest tool on every round and its output was not the prettiest. It was the only one that treated what did I change and what am I not sure about as part of the job rather than an afterthought.
- Copilot's headline 2024 total was £45 short (656,485 against 656,530) and its dashboard said every sheet had been scanned with no errors
- Copilot rewrote five typed inputs and left no record of any of them
- ChatGPT rewrote the legal variance (from 1,800 under budget to 1,700 over) without logging it
- ChatGPT built a self-checking tab and then didn't mention the £3,500 travel correction in it
- Claude's readme, otherwise complete, still got one number wrong in its own change log
- All three missed the same cheap fix (the text-formatted "45" in Misc Feb, or the headcount row, see uncertainty below)
I gave Claude, ChatGPT and Microsoft Copilot the same ten-tab programme budget, one loose prompt and no help. All three found the real story in the numbers. All three fixed the error I planted. Only Claude left a record of what it changed that I would hand to a finance director. That is the test, and that is why it wins.
Test setup
Tools: Claude, ChatGPT, Microsoft Copilot in Excel.
Task: take a messy, handed-over programme budget workbook and make it safe to circulate. Input: one Excel file, ten tabs, built the way people actually leave budgets behind. Prompt: one sentence, deliberately thin on detail. If a tool asked a question, it got the same answer every time: use your judgement and tell me what you assumed.
The file was built from ten years of running programmes and programme budgets. It contains:
The planted error is the kind that sails through review because the total looks roughly right. One of the other tabs holds the correct value. None of the tools were told that.
- A front page that looks fine, but the 2023 column is pulling budget, not actuals, and has done for a year
- A merged title band across the main data tab, so nothing sorts
- One hidden column
- Four hidden rows at the bottom where someone started 2025 and stopped
- A dashboard where every cell reads #REF!
- A tab called 2024 Old that nobody deleted
- A notes tab saying finance sent a travel figure and it was pasted in
- One planted error on the travel row: the monthly cells add to 4,200, the quarter total agrees, the annual total says 3,450
How I scored it
Four rounds, pass or fail on each.
- Did it do the work?
- Did it catch the planted error?
- Did it tell me what it changed? I diffed every returned file against the original, every cell, every sheet, hidden ones included. The rule: if a tool changes an input and does not leave a note, it fails.
- Is the result usable? I rebuilt the 2024 total from scratch. The right answer is 656,530.
What each tool returned
Claude returned a rebuilt workbook with a change log listing 29 defects, cell by cell.
ChatGPT returned a rebuilt workbook and something I did not ask for: a Checks tab with seven live reconciliation formulas that re-verify every time the file opens. I would keep that.
Copilot rebuilt everything in place, because it runs inside Excel. It added a cover page and an executive dashboard.
Round 1: did they do the work?
All three pass. Each returned a working file with the broken dashboard repaired and the structural mess dealt with.
Round 2: did they catch the planted error?
All three got the arithmetic right. The difference was what they told me.
Claude put it at the top of a list headed three things to settle before this goes to anyone else, alongside two other items it could not reconcile and wanted me to decide.
ChatGPT fixed it and left one line in the notes tab saying the Q3 total now sums the months.
Copilot fixed it and said nothing. Its summary was the shortest of the three and did not say what it had touched.
All three pass, on the strict reading that the number got corrected.
Round 3: did they tell me what changed?
This is where the field splits.
Claude: every edit is in the readme, with the cell reference, the old value, the new value and why it matters. Pass.
ChatGPT: the work is good, the record is not. It added a couple of notes, but it also rewrote the legal variance, which was showing 1,800 under budget when the truth is 1,700 over, and did not log that. Fail.
Copilot: five typed numbers were rewritten with no record anywhere, not even in the notes tab. Fail.
The answer from an AI cannot be trust me. If it changed an input, it has to say so.
Round 4: is it usable?
Claude: 656,530. Correct. The readme is also correct, and it flagged a typo in the auditor's paperwork on the way through. Pass.
ChatGPT: 656,530. Correct. Pass.
Copilot: 656,485. Wrong by £45. It missed a text-formatted cell on the headcount tab that the other two both fixed. Its executive dashboard states that all sheets were scanned and there are no formula errors. The dashboard is also the worst-looking of the three. Fail.
£45 is small. That is not the point. The headline number is wrong and the tool says it is right.
What all three got right
Every tool found the story inside the file from one loose prompt: contractors were £73,800 over budget while everything else combined came in under. Three years ago that analysis was a day of someone's time. Now it is about half an hour by the time the tools run and you run your checks.
What all three got wrong
Each has a blind spot around its own record keeping.
Claude documented everything and still fumbled a number in its own log.
ChatGPT built self-checking machinery and then did not mention a correction of roughly £3,500.
Copilot did the real repairs and shipped a confident, wrong final answer.
Was it safe to use at work?
Claude's output, yes, after reading its three open items and settling them.
ChatGPT's output, only after diffing it yourself. The Checks tab is a genuine asset, but the unlogged variance rewrite means you cannot skip the diff.
Copilot's output, no. Not because it is £45 out, but because it tells you it is not.
Who should use which
If you need an audit trail you can hand to someone else, Claude.
If you are the only person who will ever open the file and you will check it yourself, ChatGPT, and keep the Checks tab.
If you are already inside Excel with Copilot and want a first pass, fine, but treat every number it types as unverified until you have diffed it.
Grab the files to run your own test - https://app.notion.com/p/3c89311d2779815ca76cdf97cf5e106a
Tools Used
This is a budget tracker I've inherited. Make it usable and tell me what it actually says. Return the fixed file.- None