At minimum, I have ADHD. Well, I’m technically on the waiting list for a diagnosis (what confused, hapless millennial isn’t), but I know I do. It’s been… ok. Accidentally manageable, even. Most of the time. But when you’re a few years shy of 40, a dad, and trying to scrape by in this big ol’ world of work, smartphones, social media, tax returns, bills, yellow grass becoming the norm etc. – well, it can all really do a number on any noggin, let alone one that’s essentially being driven by an overzealous monkey with very poor impulse control.
Lately, thanks to a myriad of factors including the summer heat and having to entertain a five-and-a-half-year-old for six school-free weeks (all while working the entire time), the monkey in my brain is done. It had a good run, stepping up to the plate when push came to shove, but halfway through the summer holidays, the poor little guy just needs a frozen banana, and a three-day nap.
So, at just over 160 words into this article, you know about my brain. But why? Well, it’s because I think (or at least, I hope), that AI can help. You see, the most frustrating thing I find about all this is that I know there’s so much potential under the hood. Somewhere, nestled among the bundles of neurons, is a V8 engine that’s constantly revving away.
I feel bright, exciting ideas bubbling away at the surface even as I write these words – start my own YouTube channel, get back into drumming, learn to meditate, rehab my knee for running. And then there’s all the boring stuff that’s annoyingly crucial for modern life. Like booking school dinners, answering emails, and changing broadband providers for the fourth time in as many years because they all crank the bloody prices up. And that’s all just half an hour’s worth of Adult Admin on any given Tuesday morning.
The list is endless. But while it constantly feels like I’m on the verge of actually ploughing forward in life with all the admin, parenting, self-improvement, etc., something’s missing. The engine is revving, but there’s a whole bunch of empty space where crucial gears should be. Things are spinning at ultrafast speeds, but without those gears in place, nothing is connecting. I feel like I’m standing still. And yes, at times, it absolutely sucks.
It’s even worse when something requires multiple levels of planning. Forgetting to grab coriander for tonight’s chilli is one thing. But messing up, say, a kid’s birthday party? Yeah, you really don’t want to do that.
For the record, I haven’t! But only because my partner organises it each year. But this is exactly the kind of thing that I’m hoping AI can help me with. If AI can slot into the gaps where those missing gears are, maybe things can connect better. I’m less interested in asking it to write poetry or explain the meaning of life than I am in finding out whether it can shoulder some of that mental load. If I can dump all the disorganised information bouncing around my brain into a chat and get something useful back, it could make a real difference.
That’s the theory, anyway. An AI-generated plan is only useful if it notices the important details, handles uncertainty sensibly, and doesn’t confidently fill in gaps with things I never said.
To find out which assistant I’d trust with this kind of everyday task, I gave the same fictional party-planning prompt to the free consumer versions of ChatGPT, Google Gemini, Claude, and Microsoft Copilot. And the results were more revealing than I expected…
ChatGPT vs Gemini vs Claude vs Copilot: the party-planning challenge

Giving AI deceptively simple challenges is nothing new. Generators have struggled with everything from Will Smith eating spaghetti to creating an image of a wine glass filled to the brim, while more recent “car wash tests” have caught AI recommending that people walk to a nearby car wash, somehow overlooking the need to bring the car.
Different AI assistants are also better suited to different jobs. One informal test can’t establish which is universally best at research, writing, coding, image creation, and everything else. That’s why I’ve gone for a deliberately focused, narrow look at how they handle one realistic organisation task.
The birthday party itself is fictional, but it’s exactly the kind of situation I want to be able to handle without getting stressed. I want to see which assistant could take a chaotic collection of constraints and turn it into a useful plan without being told what problems to look for.
With that in mind, each AI service received the following prompt in a fresh conversation:
Turn the notes below into a clear plan for a family birthday party:
Notes:
Party at home on Saturday, 2pm–5pm. Our child is turning six. Eight children were invited, but three parents haven’t replied. My partner says ten children are coming. Two younger siblings may also attend. Six adults are confirmed. The budget is £120, including food and decorations. One child is allergic to peanuts. Another can’t eat dairy. We’ve ordered a cake serving ten people, with collection at 11am. We need food for children and adults. Pizza delivery was quoted at £78, but that doesn’t include food for adults. We want balloons but no helium. Party bags are needed, ideally without plastic tat. We can organise garden games if it’s dry, but the forecast says rain from around 3pm. The living room only comfortably fits 14 people. We haven’t decided whether parents will stay. We need to buy candles, paper plates and drinks for the children. My partner works until noon. I’ll have our child with me all morning. Nothing has been bought yet.
There are several potential traps buried in those notes which you, with your wet, organic, mammalian brain, probably spotted immediately. In particular, I wanted to see whether each assistant would spot:
- The conflicting and incomplete guest numbers, including whether parents and younger siblings would attend.
- The fact that the possible attendance exceeds the living room’s 14-person capacity, with rain potentially ruling out the garden.
- The £78 pizza eating up most of the £120 budget without feeding the adults.
- The ten-serving cake potentially being too small, along with the peanut allergy and dairy restriction.
- The amount still to buy and prepare, despite one parent working until noon and the other looking after the child.
Crucially, I also wanted to see whether the assistants would acknowledge what they didn’t know or simply invent convenient answers.
I used the default free experience offered by each service, kept the first response, and didn’t provide corrections or follow-up prompts. I wasn’t simply looking for the longest answer either. The best response needed to spot the conflicting attendance figures, tight budget, limited space, and dietary requirements without inventing convenient solutions.
Here’s how they got on:
ChatGPT


ChatGPT produced the most comprehensive response, but importantly, it didn’t mistake understanding for certainty.
It began by identifying the headcount as the biggest issue. Rather than deciding how many guests were attending, it suggested contacting the three parents who hadn’t replied and asking whether their child was coming, whether an adult would stay, whether a sibling would attend, and whether there were any other dietary requirements.
ChatGPT also spotted that 12 possible children and six confirmed adults couldn’t all fit comfortably in a room with a capacity of 14. It rejected the £78 pizza because it would leave only £42 for adult food, drinks, decorations, plates, and party bags, and suggested a simpler supermarket spread instead.
Its allergy advice was the strongest of the four. ChatGPT recommended checking ingredients and precautionary warnings, while making as much of the children’s food as possible suitable for both affected guests. It also realised that a cake serving ten might not cover 12 children and suggested having fruit or another dairy-free dessert available.
The answer wasn’t flawless. Its suggested £85 budget didn’t include the unknown cost of the cake, so its claim that the party should remain comfortably below £120 was too confident. It identified the capacity problem without really solving what would happen if 18 people were forced indoors by the rain, and the whole response was longer than it needed to be.
Still, it offered the best combination of caution and practical detail.
Gemini


Gemini’s response looked clean and well organised. It identified the need to chase outstanding replies, decide whether parents could stay, and reconsider the pizza delivery. It also suggested using the garden before the expected rain, with indoor games and crafts ready from 3pm.
Unfortunately, it introduced a basic error almost immediately. Despite the prompt clearly stating that the party ran from 2pm to 5pm, Gemini listed the time as “2:00 PM – 3:00 PM”. It later scheduled indoor activities from 3pm to 5pm, so it clearly knew the correct finishing time somewhere along the way. That only leaves the response contradicting itself.
Its headcount was inconsistent too. Gemini said there could be “up to 18 people”, but ten children, two younger siblings, and six confirmed adults already add up to 18. Any additional parents staying would push the number higher.
Gemini recognised the peanut and dairy requirements but didn’t do much with them. It suggested allergy-safe food and a dairy-free alternative, but didn’t consider whether the cake was suitable or say much about cross-contamination. It also noted that the cake served ten without offering a clear solution.
The answer was broadly competent, but if I need to check basics such as the party’s finishing time and maximum attendance myself, it hasn’t done much to lighten my mental load.
Claude


Claude gave the shortest answer, and arguably showed the clearest judgement.
Rather than burying the important issues beneath shopping lists and activity suggestions, it led with four things that needed resolving – the headcount, indoor capacity, budget, and whether parents would stay. Its best suggestion was also the simplest: contact the three unconfirmed families immediately because their answers affect the food, cake, space, and cost.
Claude also made fewer unsupported assumptions than any rival. It correctly calculated that ten children and six adults already exceeded the living room’s comfortable capacity, and that spending £78 on pizza would leave only £42 for everything else. It also recommended preparing an indoor alternative rather than relying on the garden.
The trade-off was that Claude produced more of a planning framework than a finished plan. Its shopping list and activity suggestions were basic, and its allergy advice wasn’t strong enough. It recommended keeping peanut-containing items separate, but checking the arrangements with the child’s parent, avoiding peanut products, and taking greater care over cross-contamination would have been more useful. Health guidance warns that buffet food, plates, and cutlery can all present cross-contamination risks at parties.
Claude was also the answer I most enjoyed reading. It was concise, calm, and didn’t pretend to know things it couldn’t. Had this test been based purely on judgement and restraint, it probably would have won.
Copilot


Copilot initially produced the most impressive-looking response. It included an itemised menu, activity schedule, shopping plan, party-bag ideas, and a budget with an apparent £22 contingency.
Look more closely, however, and much of that detail rests on assumptions.
Copilot decided that 12 children and nine adults were likely to attend. It appears to have turned the three parents who hadn’t replied about their children into three additional adults who might stay, even though those are separate questions. That attendance figure is possible, but the prompt doesn’t provide enough information to justify it.
It then recommended assuming that all parents would remain, while continuing to plan for nine adults. It also assumed that the house had a suitable kitchen or dining area that could act as overflow space, despite neither being mentioned.
Its menu relied on approximate supermarket prices, yet presented a confident total of £59. Depending on how its suggested pizzas are counted, the individual figures appear to add up to somewhere between £59 and £63. Its decorations similarly ranged from £18 to £21, but the final budget used the lowest figure.
There were good ideas here. Copilot correctly rejected the expensive pizza, included dairy-free and peanut-free choices, and produced a varied mix of activities. Cutting a ten-serving cake into smaller portions could also stretch it further, although relying on that for potentially 12 children and numerous adults isn’t much of a plan.
Copilot created an impressive sense of completeness, but it often achieved that by quietly deciding what the missing information should be. It looked polished, yet it was the answer I’d trust least.
ChatGPT vs Gemini vs Claude vs Copilot verdict


ChatGPT narrowly wins this particular challenge. It wasn’t flawless, but it was the best at doing two things at once – spotting the problems hidden in the notes, and still producing a plan I could use. It caught the attendance conflict, capacity issue, inadequate cake, limited budget, and dietary requirements without simply inventing answers to every unknown.
Claude came extremely close. In fact, it showed better judgement in places, particularly in how quickly it cut through the clutter and prioritised the unanswered invitations. It finishes second because the original task was to create a clear plan, and its response stopped a little short of that. But it’s nothing that a quick follow-up prompt wouldn’t fix.
Gemini got most of the fundamentals right but undermined itself with avoidable errors. Copilot offered the most detail, but much of that apparent helpfulness rested on assumptions that weren’t in the prompt, which was a big red flag for me.
The biggest lesson, then, is that a longer, more confident AI response isn’t necessarily a better one. If I’m using an assistant to reduce my mental load, I don’t need it to hide uncertainty beneath a convincing-looking checklist. I need it to notice what matters, tell me what it doesn’t know, and still help me decide what to do next.
With all that in mind, I’d classify this experiment as a success. Even though no single response was perfect, I can clearly see how much time and stress I could save myself using AI to plan and organise things in future – especially with more practice and honing my follow-up prompts. Being given even a rough on-paper plan gives me something solid to focus on without drifting off course. But more importantly perhaps, it gives that little, loyal monkey, a well-deserved rest.
