Contents
- Models on test – 80
- AI twins: Copilot and ChatGPT – 83
- Test 1: writing – 84
- Test 2: coding – 86
- Test 3: analysis – 88
- Test 4: advice – 91
- ChatGPT vs the rest: is there a winner in this Labs? – 95
Introduction
- Last summer, we prompted eight AI platforms to magic up a range of digital images (see issue 368, p78). This time we've returned to the subject with a focus on more general tasks, covering business writing and analysis, programming, and providing advice in response to a hypothetical home scenario. In particular, we were interested to know whether any of ChatGPT’s many rivals could usurp the biggest name in AI.
- It's impossible to test every facet of every model, so we opted for more generalised tasks, with a focus on public data and well-documented, widely used formats. In the programming test, for example, we chose to work with Python and XML, as they’re both popular and well understood, so no AI should be at a disadvantage. In the data analysis test, we converted our source material from OpenDocument format (which not all the models could access) to Excel. The source material itself is UK government data, which is plainly presented and annotated.
- We were also careful to balance each prompt. They needed to be specific enough to push each model in the same direction but leave sufficient room for them to interpret our request as they saw fit. So, in our newsletter-writing test we allowed each AI to choose its preferred source material, and while we gave plenty of feedback on the programming test, we didn't allow them to lean on any libraries that weren’t already installed on our test machine. Why? Three reasons:
- We knew the task could be achieved using the libraries already installed:
- We wanted to level the playing field; and
- We were mindful of the fact that in some environments - business environments in particular - it wouldn’t be possible for users to install additional libraries without authorisation.
- The field of AI is developing so quickly that a Labs of this type presents several challenges. When we’re testing hardware, downloaded software or, to a degree, SaaS, products are frozen in time, assigned a model or version number, and perform in a predictable way. Tests conducted one day could be replicated with a high degree of certainty that the following day they’d produce the same output. Not so, AI. For one thing, the models are inherently unpredictable. Minor changes in the way a prompt is worded could influence the output, and if you were to use the same prompts as us, you might receive an entirely different result based on circumstances or your past use.
- The result delivered by each of our tests is thus a snapshot, so rather than declaring a “winner" in each instance, we’ve highlighted the model we’d return to for each job. This has informed our conclusion, where we select the model for which we’d maintain a paid subscription based on its performance in these tests.
- Fortunately, where payment is required, none of the models on test locks you in for an extended period. So, if things do change throughout 2026, new models appear and one AI or other gains a march on the others, you’re free to chop and change at will.
- Contributor: Nik Rawlinson
Models on test
CHATGPT
- PRICE: Free or from £17 (£20 inc VAT) per month from chatgpt.com
- ChatGPT is surely the best-known player in this field. It has five tiers, starting with a free plan that imposes limits on each of its key features; reasoning, messages, uploads, deep research, image generation and context. The Plus and Pro plans, at £20 and £30 per month respectively, unlock additional resources, including access to GPT-5.1 for advanced reasoning. The Plus plan is subject to usage limits, but it’s unclear when or if you might hit your cap as “these limits may vary based on system conditions”. Pro is unlimited but subject to “abuse guardrails”.
- Only the Business tier offers annual billing, which trims the monthly per-user cost from £30 to £25 if you pay for the year up front. Business also rolls in support for admin controls, encryption, data analytics and company knowledge within responses if you connect to Slack, Google Drive, SharePoint and so on.
- Upgrading to a paid plan also gets you access to voice control and responses, allowing you to conduct a convincing back-and-forth conversation with the service. We’ve found this an effective tool for brainstorming ideas, developing outlines, and even balancing the relative merits of multiple options, particularly when shopping.
CLAUDE
- PRICE: Free or from $17 per month from claude.ai
- The Claude series of large language models (LLMs) is developed by Anthropic with the aim of being not only helpful, but also harmless. It does this by referring to Anthropic’s written constitution, against which the AI evaluates an input before formulating a response. It then chooses the most harmless output (see tinyurI.COm/378claude).
- Its free plan lets you create and edit content, generate code and analyse inputs, but it’s subject to usage limits that reset every five hours. The number of messages you can send will also vary based on demand, and “other types of usage limits” may be imposed to ensure fair access to all.
- Paid tiers start at $17 if you sign up for a year (otherwise $20 monthly), which hikes the usage limit five-fold per session during peak hours. “If your conversations are relatively short (approximately 200 English sentences, assuming your sentences are around 15-20 words) and use a less compute-intensive model, you can expect to send around 45 messages every five hours, often more depending on Claude’s current capacity” according to Anthropic. You’ll also be able to connect to tools such as Google Workspace, create files and execute code.
- With the Max plan, which starts at $100 a month, you can choose to lift the limit yet higher, to between 5x and 100x that which is allowed on the Pro tier, and you’ll get priority access at high traffic times.
COPILOT
- PRICE: Free or from£16 (£19incVAT) per month as part of Microsoft 365 Premium from copilot.com
- Copilot’s free model is closely integrated with Windows 11, Edge, Bing and the Copilot mobile apps. There’s no cap on usage, but you may find performance is degraded during peak times. Paid users with a Microsoft 365 account shouldn’t experience any degradation.
- However, some of the limits still aren’t clear, with “extensive use” allowed in Chat and Pages, and Microsoft 365 Family users also having “extensive use” of audio overview in notebooks and podcasts, where Microsoft 365 Personal users have only six uses per day.
- Microsoft governs your use of several features using AI credits that run down as you perform certain tasks. Crucially, you have 60 credits (renewing every month) for drafting, rewriting, summarising and analysing data in Microsoft 365 apps. This same balance is depleted as you use Designer to generate images or perform automated image-editing tasks. You’ll therefore need to be careful not to use them all in one module if you think you’ll need to rely on them elsewhere later.
DEEPSEEK
- PRICE: Free from deepseek.com
- DeepSeek is a comparative newcomer. Despite having existed since 2023, the LLM, operated and owned by a China-based hedge fund, gained worldwide recognition in 2025.
- Unusually, it’s entirely free and unlimited - at least for chat.
- Currently, you only need to pay if you want to use the API, so chat and conversations, file uploads, web search and mobile app use are all gratis and uncapped. DeepSeek says that it’s “committed to keeping the service free and accessible. If there are ever any changes to this policy in the future, they’ll be clearly communicated to users. For now, you can enjoy all the features without any cost or subscription fees. ”
- The main interface is accessed at chat.deepseek.com, and it operates in the same manner as its competitors. However, at the bottom of the prompt box you’ll find two options, for “Deep thinking” and Search. The first of these instructs it to consider its answer for longer before responding, while the latter tells it to search the web where necessary. Enabling deep thinking is particularly interesting if you want to read a transcript of what’s going on behind the scenes.
GEMINI
- PRICE: Free or from £16 (£19 inc VAT) per month from gemini.google
- What was once known as Bard is now Gemini. Google’s AI platform is available standalone for in-browser chat, integrated into Google Workspace where it can work alongside paying subscribers in Docs and Sheets, and is increasingly chipping in when you interact with Google Assistant.
- The free tier gives access to the Gemini 3 Flash model, plus limited access to Gemini 3 Pro. This isn’t necessarily a problem. While 3 Pro is tailored to deep reasoning and handling complex creative tasks, Gemini 3 Flash has been found to beat it in some tasks, including agentic coding and multimodal reasoning. It also has four response settings, from minimal to high (compared to Gemini 3 Pro’s two), allowing you to dial down the level of detail if you need a quick answer.
- Without paying, you get 100 monthly credits for video generation and can use NotebookLM as a research and writing assistant. Upgrade to Google AI Pro at £19 a month and your credits increase ten-fold, you benefit from higher daily limits in CLI and IDE use, you can access Gemini directly in Google apps, will benefit from five-times more audio overviews and notebooks in NotebookLM, plus you up your Google Drive allocation to 2TB.
- Google states that all users have “general access” to fast thinking. When using the Thinking or Pro models, free user limits “may change frequently”, AI Pro users can send 100 prompts per day, and Gemini app and Google AI Ultra users can issue 500 prompts.
GROK
- PRICE: Free or from $25 per month from grok.com
- Grok’s free plan gives you “limited access” to chat models and limited context memory. Exactly what these limits are isn’t made clear; Grok itself notes that “these consumer limits appear to be dynamic or account-specific and typically visible only in your account settings or when you hit them during use”. If you find you do hit a cap, upgrading to SuperGrok costs $300 per year ($25 per month), or $30 per month if you don’t want to commit to a year up front.
- This grants you increased access to both the Grok 3 and Grok 4.1 models, extended memory, priority voice access, and use of Ani and Valentine. These are “companions” with which you can interact on a more personal level, as though they were friends.
- You can customise Grok’s responses to provide formal or concise output, custom responses (the style of which you can describe in a “How should Grok behave?” box), or take the Socratic option, which behaves like a teacher to try and help you foster deep understanding. In Grok’s words, it will “avoid direct answers; instead, ask thought-provoking questions that lead the user to discover insights themselves. Prioritize clarity, curiosity, and learning, while remaining patient and encouraging.” We didn’t customise the responses in any of our tests.
MISTRAL AI
- PRICE: Free or from €15 (€18 inc VAT) per month from chat.mistral.ai
- Mistral AI is headquartered in Paris, which explains the cat you'll find dotted around its website (“chat” being French for cat, as well as describing what you do with the service). The free tier lets you chat, search and code, all without logging in, but we had to sign up for a free account before we could upload files for analysis. However, even when logged in, there are “some daily limitations” on use.
- Mistral doesn’t explicitly state what these are - see tinyurl.com/378mistral for more information - but they apply alongside shorter-term limitations that reset every three hours.
- The Pro tier, at €18 inc VAT per month, lifts these limits, giving you capacity for six times the chat messages and five times the web searches you can perform for free. You can likewise upload 20 times the number of documents and generate 40 times the number of images.
- Mistral has several language models. Its three core options being Large, Small and Next (prototype), which it selects depending on the task. It also maintains Codestral, which is optimised for code generation and completion, a lightweight model called Nemo designed for reasoning tasks, and Voxtral, which has a focus on audio and speech-related tasks.
PERPLEXITY
- PRICE: Free or from $17 per month from perplexity.ai
- Perplexity bills itself as a “direct line to the world’s knowledge - compressed, cited, and made clear”. As such, it’s pitching itself more as an Al-backed search engine than a traditional AI chatbot. This is obvious in its answers where short, narrative paragraphs are routinely linked to source material, producing what feels like a more personal and relative alternative to a traditional search results page.
- The free plan offers “practically unlimited basic searches” and a limited number of Pro searches, with Perplexity picking what it considers to be the best model for each query. You can upload basic files, but don’t get access to advanced AI models, image generation or premium support.
- These are features of the $200 per year ($20 per month if you pay monthly) Perplexity Pro plan, which also opens up advanced AI models, practically unlimited file attachment uploads and analysis, and a two-business day response time on support.
- Perplexity also has an Education tier ($4.99 a month), which is tailored to users at university level institutions and higher. As well as matching the Pro plan’s features, it includes ten times the number of citations in its answers and gives access to a dedicated study mode which, as well as providing answers to your prompt explains the answer in detail and breaks it down step by step to help you better understand the response.
AI Twins: Copilot and ChatGPT
- With Microsoft’s Copilot buit upon OpenAI’s GPT engines, we explore whether one has the advantage over the other – or should you use both?
- Microsoft Copilot is based on ChatGPT, so you might wonder why we’ve included both this month. The answer is clear from our tests, with often different results despite using the same AI engine. There’s also the fact that many users will lean towards Microsoft as their default, especially if they already subscribe to Microsoft 365 services.
- To ensure we treated each model equally, we performed our tests via their browser front-ends. For many this is the only option, but a handful of AI services are closely integrated with productivity tools and operating systems. ChatGPT, for instance, underpins Copilot in Windows 11, as well as some productivity features in Word, Excel, Outlook and Teams. If you use a Mac with Apple Intelligence, you can likewise use ChatGPT to refine local content and, if you have a ChatGPT login, use the same account within macOS as you do on the web. Gemini appears in Google Docs and other Drive-related apps if you’re signed up to the Google AI Pro plan or higher.
- Often these add-ons are a bonus. You can use Gemini independently of Google Workspace, for instance, and you can do the same with ChatGPT outside of macOS. But what about Copilot? You can interact with it directly at copilot.microsoft.com, just as you would with ChatGPT. So you may wonder why you’d pay for £19 a month for Copilot when £85 a year gets you increased limits in Copilot, plus the Microsoft 365 applications and 1TB of OneDrive storage?
Conversational vs embedded AI
- The answer is somewhat nuanced. Conversation-first assistants, which would include ChatGPT, are good at brainstorming and general assistance, particularly in creative activities such as report writing and ideation. However, they don’t automatically act on existing files unless they are explicitly provided. An embedded productivity assistant, on the other hand, such as Copilot in Microsoft 365 or Gemini in Google Workspace, can work on material in situ in a similar fashion to a sidekick or underling sitting alongside the user, as well as creating in-app content from scratch
- With ChatGPT and Copilot, the output you get through the browser interfaces could well be different, depending on context. When asked to explain the difference between itself and ChatGPT, one of the factors that Copilot highlighted was that it always pulls in live web results when the user asks for facts, comparisons or recommendations. We asked it to clarify this point: “For me, citations are built into the factual-answer workflow. For ChatGPT, citations are optional and user-driven. That’s the real difference - not capability, but design philosophy. ”
- One other key difference is the way that you interact with them and they interact with what you’re doing. While ChatGPT runs in its own tab within a browser, Edge users can keep Copilot in the sidebar where it can interact with the displayed page. How long this distinction remains relevant is open to question, though, with AI services rolling out their own browsers. OpenAI’s Atlas browser lets ChatGPT interact with content on a displayed page, and can help users work on content in web-based productivity apps. Perplexity’s Comet browser can also perform like an agent, for example by adding products to a shopping cart.
Either... or both?
- ChatGPT offered to produce a neutral, vendor-agnostic response to our requested comparison between the two models. It highlighted the benefits of each approach, with standalone conversation-first assistants optimised for long, iterative conversations, supporting “thinking out loud” and gradual refinement, but requiring the manual transfer of results into productivity software. Embedded productivity AI assistants, meanwhile, are in ChatGPT’s opinion optimised for acting on existing content and designed for short, task-focused interactions. They enhance efficiency during focused, execution-oriented work, but conversations may be shaped by the surrounding application context.
- As such, employing both standalone and embedded assistants may be key to effectively using AI throughout a multi-step workflow. Ideation can take place through the standalone platform, with the results transferred to an Al-enabled productivity suite for refinement. As such, the two aren’t so much substitutes as they are two sides of a single coin. But then, they would say that.
TEST1: WRITING
Prompt I need to send a newsletter to my subscribers at the end of the year. They want to know what were the most popular television programmes of 2025 in the UK, and what they should be watching out for I in 2026. Produce around 600 words, with a minimum of four stories in the main body of the newsletter. Each must have a heading, which should be hyperlinked to the original source. Also provide a subject line and subhead for the newsletter, which would entice recipients to open and read it, and provide a three-sentence introduction. Include a section headed “In case you missed it...” containing hyperlinks to three other stories that would be of interest. The audience is an educated audience of people who work in the television industry. Provide the newsletter as an html file that I can download and paste directly into my newsletter management platform.
Follow-up Prompt Please produce one-sentence summaries of these stories that we can use on social media to promote the newsletter
Model-by-Model Findings
CHATGPT
- ChatGPT turned in 665 well-targeted words, with six main stories and a convincing introduction reminding TV industry insiders that “BARB charts and weekly ratings tell a familiar story: big, broad-audience brands still dominate, but viewers are flocking to distinctive drama and high-concept factual across BVOD and SVOD” and using language that would clearly speak to our audience throughout, convincing many that this had been written by someone in the know. Our only gripe was its focus on Bake Off “keep[ing] its crown” with 6.7m viewers while not mentioning Celebrity Traitors when the story it linked to included the latter’s 12.2m viewers. The follow-up prompt got us three factual responses, but none of them included a call to action. While we didn’t specifically ask for this, we don’t feel that the sentences it produced would, in themselves, encourage sign-ups.
CLAUDE
- Claude’s output was split into two columns, with thinking and explanations on the left, and the content of the newsletter to the right. Claude previewed the output in-line, and there was a handy “Download as HTML” link behind the three-dot menu above it. The content was excellent, and the opening story (Celebrity Traitors) was fact-filled, sprinkled with analysis, and threw forward to 2026.
- The copy flowed smoothly throughout and read like it had been written by a human (“Timothy Spall’s Death Valley offered a cosy crime solution for audiences seeking lighter fare, combining the actor’s inestimable screen presence with Welsh locations and gentle satire”). Two of its main-story sources were list-based features, but the other two were more in-depth analyses.
- The output of our follow-up prompt was equally well written, but none of the three sentences included a call to action, which would need to be added before we could use them to promote the newsletter on social media.
Copilot Copilot didn’t provide our email as a downloadable file, but in conversation HTML to be copied and pasted into our own document. It produced eight stories across 532 words (four looking back on 2025 and four looking ahead to 2026), and the intro was personable and felt like it had been written with a specific reader in mind.
- The actual value of the stories included in the newsletter didn’t feel as high as those produced by ChatGPT or Claude, from which we could get more meat without heading off to the original source material, perhaps because Copilot produced more of them within the given word count.
- All four of the stories for 2025 were linked to roundups of the best TV of 2025, with one of the linked stories already being five months old and another being Wikipedia’s entry for “2025 in British television”. The additional links in the “In case you missed it... " section likewise all pointed to listicles, with two of them having already been included in the main body of the newsletter.
- The three sentences produced in response to our follow-up were good and the subjects were engaging but, like Claude’s and ChatGPT’s responses, they didn’t feature calls to action.
DEEPSEEK
- DeepSeek generated both plain text and HTML, each presented in-line, rather than as a downloadable file. The content was engaging and well written, and although a first attempt included several unflagged hallucinations (perhaps because the Search option wasn’t activated), the second achieved what we’d asked for.
- It led with the success of Celebrity Traitors and Netflix’s Adolescence, looked ahead to 2026’s second series of The Night Manager, and previewed three other forthcoming programmes: Lord of the Flies, The Lady and The Split Up. In the version of the newsletter that appeared within the chat log, there were relevant links for each of these, but they disappeared from the HTML version, leaving only Markdown-style references (‘[citation:6]’, etc) in their wake. This would require manual revision prior to dispatch.
- Unfortunately, where the first, inaccurate attempt at producing a newsletter had resulted in the strongest calls to action of all the LLMs, there were none in the follow-ups provided when we repeated the test with Deep thinking and Search active.
GEMINI
- Gemini didn’t give us a downloadable file, but in-line HTML, 716 words and five stories. Its introduction was engaging, welcoming readers to the “final newsletter of the year” - a nice touch we hadn’t specified - and hinting at what followed without letting any cats out of their bags. The stories were an engaging blend of research and analysis and the output was also generally clean and would require less reformatting than that produced by Perplexity, aside from swapping out some Markdown markers for regular bold and italics. We initially ran this test using Gemini’s Fast model, but repeated it with Thinking, which reduced the number of stories to four, changed the selection, and talked not about Celebrity Traitors, but the non-celeb variant, which had aired at the start of the year.
- Our follow-up prompt delivered three well-written single-sentence summaries of our stories, two of which started “Discover why... ” and “Get the full breakdown on... ” to form engaging calls to action.
GROK
- Another service from which we had to copy our HTML rather than downloading a file, Grok gave us 576 words and four stories, written in a conversational and convincing style that directly addressed readers who had a “commissioning calendar” to fill. The body of the newsletter was a good mix of fact and analysis. The success of Adolescence, for example, “underscores the power of character-driven narratives in a fragmented market - a blueprint for commissioning intimate, relatable tales amid blockbuster fatigue”, while the forthcoming seventh series of Line of Duty “exemplifies how sequels can reinvigorate public service broadcasting in a streamer-saturated era”. The links it provided were best-of roundups rather than single-story in-depth pieces, and three pointed to the same URL.
- The response to our follow-up prompt for social-ready summaries was perhaps the best of the lot, with compelling calls to action and a placeholder for our newsletter link at the end of each one.
MISTRAL
- Mistral’s newsletter was peppered with coded references (“:refs[5-1]”, etc) but did read as though it had been written specifically with television industry professionals in mind. Its analysis was on point, and it noted that Netflix’s Adolescence “not only topped streaming charts but also prompted UK Prime Minister Sir Keir Starmer to endorse its use as an educational resource in schools”.
- Less impressively, its analysis of Severance season 2 noted “standout performances from Parker Posey, Carrie Coon, and Aimee Lou Wood”. They didn’t appear in the show, as Mistral admitted when we asked if it was sure about this point. It also said that Nobody Wants This starring Robert De Niro and Jesse Plemons was a “timely exploration of power, corruption, and media manipulation”. It seems likely it actually meant Zero Day, which did star those two actors. They didn’t feature in 2025’s Nobody Wants This, a love story starring Kristen Bell and Adam Brody. Again, it noted its mistake when we pointed it out. All of the stories to which it linked were listicles, and one was repeated.
- Its response to our follow-up prompt was good, with clear calls to action and, as it had taken into account our requested clarifications, its original errors were not repeated.
PERPLEXITY
- Perplexity’s newsletter ran to a paltry 331 words. The output wasn’t as clean as that provided by ChatGPT, with embedded hyperlinks inside our stories that, rather than being attached to copy, were formatted as “[web:07]” and so on. These would need to be removed or reformatted before we could send the newsletter.
- The content was good, with a direct reference to Celebrity Traitors in the introduction, but it wasn’t as personable as that produced by ChatGPT or Claude, and it didn’t address the audience personally. It also mentioned the success of Wolf Hall: The Mirror and the Light, which didn’t air in the UK in 2025, but 2024.
- Its responses to our follow-up prompt were excellent. Each opened with an engaging call to action: “Discover how... ”, “Find out why... ” and “Get ready for 2026 with...”
Summary Table
- ChatGPT: 665 words, Downloadable, Minor omissions, No CTAs
- Claude: 1,010 words, Downloadable, Strong and polished, No CTAs
- Copilot: 532 words, Inline, Shallow sourcing, No CTAs
- DeepSeek: 524 words, Inline, Generally accurate, No CTAs
- Gemini: 716 words, Inline, Generally accurate, Good CTAs
- Grok: 576 words, Inline, Convincing and targeted, Best CTAs
- Mistral: 467 words, Inline, Hallucinations, Good CTAs
- Perplexity: 331 words, Inline, Some errors, Strong CTAs
Which would we continue using?
- Grok came closest to delivering our requested word count, and its response to our follow-up prompt delivered the strongest calls to action. The content was tailored to appeal to our specified audience, but we would have appreciated a better selection of linked source material.
- Although longer than we’d wanted, Claude's newsletter was well written and relevant, even if the follow-up output lacked calls to action. On the basis of this test, therefore, it’s Claude that we’d return to the next time we need to write a newsletter, perhaps with an additional prompt to trim things if required.
TEST 2: CODING
PROMPT
- Write a Python3 script that opens a WordPress export file called wordpress_export.xml. Process this and output an html file that includes only the date, URL, title and body content of each post. The date should consist only of the date and month, not the year or timestamp. This should be formatted to sit between tags. It should immediately be followed by a link to the original post, which is a hyperlinked version of the URL itself. This is then followed by the post title, between tags, then the post content. Strip out all images, videos and other media. Place a # between each post as a delimiter.
Model-by-model findings
ChatGPT
- ChatGPT’s first attempt was a 7KB script of 257 lines. However, when we ran it at the terminal the processed file was empty and the output stated “Wrote o posts to wordpress_posts.html”. We fed back to ChatGPT, which responded, “Gotcha - that almost certainly means the filters I added for post_type-”post" and status="publish" excluded everything in your file”.
- It produced a second, slightly shorter script that did what we asked, accurately outputting a file that met each of our requirements. It counted 539 posts in total, rather than the 531 we expected, as it included several posts that were published images, rather than traditional text-based posts. As requested, it had stripped out the media from each of our entries.
CLAUDE
- Claude succeeded on its first attempt, producing a 173-line script (5KB) that correctly extracted our content, formatted the dates, headers and URLs, and placed a # between each post. It parsed 531 entries into the output file. No further work was required.
Copilot
- Copilot’s first output was very short, running to only 68 lines. It rendered dates in the parsed file as oi-oi, 02-01 and so on, rather than 01 January, 02 January, as others had done, and there were problems with special characters such as curly apostrophes. We fed back and it produced a in-line script (4KB) that resolved each of the issues and output 531 parsed entries. No further work was required.
DeepSeek
- DeepSeek initially formatted dates “Sat, 01 Ja” and exposed WordPress’s block tags in the output. We fed back and it fixed the date, albeit in American format, but removed all visible paragraph breaks and placed each entry in a styled box. The block tags remained. We continued prompting and it went through several iterations and clarified that it understood what we meant when we talked of WordPress’s structural tags.
- Although there was some understandable confusion (it correctly identified them as “internal block markup comments that get embedded in the content when using the Gutenberg block editor”), it nonetheless explained that “[t]hey should be completely removed from the output since they’re just internal formatting markers”, despite the fact that it had, until that point, retained them and made them visible in the parsed output.
- The next attempt excluded the WordPress tags but converted all paragraph break and heading tags into plain text within each post, which was rendered as a single block of running text. After a further prompt from us, DeepSeek wrote new code that introduced a new problem: “An unexpected error occurred: name ‘escape’ is not defined”. We reported this problem and DeepSeek tried again. This time the code correctly parsed our source file.
Gemini
- Gemini got it right the first time. Its code ran to 144 lines and 5KB. When we ran it, it identified 531 posts, which it correctly formatted with the appropriate header tags and properly formatted URL. Each entry was delimited by the requested #. There was no need to follow up with clarifications or requests for changes.
GROK
- Grok’s first attempt exposed WordPress’ structural tags in the processed file ( and (). Further when we ran the script, Terminal issued an alert: “DeprecationWarning: Parsing dates involving a day of month without a year specified is ambiguous and fails to parse leap day. The default behavior will change in Python 3.15 to either always raise an exception or to use a different default year (TBD). To avoid trouble, add a specific year to the input & format.”
- We fed this back, pasting in the complete deprecation warning and notifying Grok that it had exposed the structural tags. It rewrote the code using BeautifulSoup (as did Perplexity) and when we told it that we didn’t want to use this, it generated new code that removed all the paragraph breaks from the content and re-ordered the posts, so the oldest came first.
- We asked it to fix these issues. The next iteration of code restored the original order, but there were still no paragraph breaks and where there should have been dates, there was a placeholder: “{p[‘date’]}”. We prompted for a fix, which Grok declared “100% correct and tested”, but when we ran this we got another “DeprecationWarning” and the processed output file contained no entries. One last re- prompting fixed things. We got the output we were after, albeit with American-style dates.
MISTRAL
- It took Mistral five attempts to successfully parse our WordPress export. The first fail was significant, with Terminal unable to run the script at all because it had assumed that every post had a “content:encoded” field, which they didn’t.
- While the second script fixed this, it output malformed extended characters, and dates were rendered as just a three-letter day of the week; no number and no month. In fixing this, it removed all visible paragraph breaks but made WordPress’ structural tags visible. It removed these on the fourth attempt but failed to switch the output file out of italics at one point due to unbalanced tags.
- Solving this, on its fifth attempt, provided the output we were after, although like ChatGPT it parsed 539 posts, rather than 531, as it included eight that were image posts, rather than text (it correctly stripped the images from them). Although it took a few attempts to get it right, Mistral was extremely fast, and the overall process felt like surprisingly little work on our part.
PERPLEXITY
- Perplexity’s first attempt used an rule rather than # to delimit each post, and there were formatting issues on curly punctuation. Pound signs were rendered A£, hyphens came out as a€”, and single curly apostrophes as a€™. We fed back and it solved all the issues, but at one point in the output the text was italicised, and italics weren’t turned off for the rest of the file.
- Pointing this out solved the problem with italicisation but produced a new one: tag within a post, which our original file used for crossheads, it didn’t turn it off until the end of the post, so all remaining text became a header. We fed back again but it didn’t solve the problem, and on a subsequent attempt it used the BeautifulSoup library to parse the text. We told it that we didn’t have BeautifulSoup installed, so it wrote a sixth iteration of its script, which this time worked as expected.
Summary Table (Prompts required, Final script length, Posts identified)
- ChatGPT: 2, 138 lines (7KB), 539
- Claude: 1, 173 lines (5KB), 531
- Copilot: 2, 111 lines (4KB), 531
- DeepSeek: 7, 380 lines (12KB), 531
- Gemini: 1, 144 lines (5KB), 531
- Grok: 5, 422 lines (13KB), 539
- Mistral: 5, 88 lines (4KB), 531
- Perplexity: 6, 155 lines (6KB), 531
Which would we continue using?
- Two models immediately stood out in this test: Claude and Gemini, as evidenced by the length of their entries above. They each grasped what we wanted and delivered working code on the first attempt. Although we can’t say on the basis of this test whether they would do the same when tasked with writing SQL, Swift or C# code, for example, we’d happily use either when working with Python, which is today one of the most popular and important general-purpose programming languages in common use.
- We didn’t tell any of our models how we would be using the parsed output, so the stripped-back file produced by Gemini (highly reminiscent of an early 1990s homepage) is perfectly acceptable and met our requirements. Claude’s output was more attractive, though, with margins, centred delimiters and a more attractive font, which narrowly tips the balance in its favour.
TEST 3: DATA ANALYSIS
PROMPT Summarise the content of this spreadsheet. Pay particular attention to any outlying metrics and also deduce what is the best month in which to book a driving test if you want to maximise your chances of passing, all other things being equal. Pick out anything else you find interesting, unusual or noteworthy.
UPLOADABLE FILE
Table reference DRT121G ("Car driving tests conducted, passed, pass rates and forward bookings, January 2019 to date: Great Britain") from the UK government source, tinyurl.com/378data. Although we downloaded the data in .ods (OpenDocument) format, we converted it to .xlsx (Excel) format before uploading to each LLM. The government cover sheet and notes were retained, so the LLMs would need to look to the third tab to find useful data. The spreadsheet made no specific reference to Covid, the pandemic or lockdown.
Model-by-model findings
CHATGPT
- ChatGPT structured its 570-word analysis into clear, bulleted sections for easy scanning. It identified outliers, such as a 79% pass rate and a month with only 68 tests, as results of Covid-related restrictions leading to “selective testing (priority workers)”.
- Excluding these anomalies, it found April to June typically had slightly higher pass rates, with a winter dip due to “weather-related factors” and “more cancellations”. Summer remained stable but not significantly higher than spring. The analysis concluded that April, May or June consistently showed higher-than-average pass rates, with more predictable weather and higher throughput, without pandemic distortions.
- It recommended May as the most consistently strong month, despite most competitors favouring April. ChatGPT explained its choice was based on May’s better daylight, fewer weather disruptions and slightly higher test volumes, emphasising robustness and realism over raw averages.
- ChatGPT acknowledged that this weighting choice was subjective and offered further analysis, which proved that April was indeed the best month with a weighted average, but also suggesting August was better with a weighted pass rate. Ultimately, it decided that August was the best month under normal conditions, despite April’s spreadsheet-indicated superiority.
CLAUDE
- Claude produced only 385 words of analysis, but caught all of the points we were expecting. It noted that Covid had a dramatic impact on test volumes, with a 99.6% drop from normal levels, and that this accounted for some extreme statistical outliers, including abnormally high pass rates in April 2020 and February and March 2021.
- It noted that the best month to book a test was April, but that February and March were also “good” months. Gemini also said that February and March “appear to be excellent months” before giving reasons for discounting them (see below). Claude summarised, “the data suggests spring months offer the best statistical advantage for passing, while late autumn sees the toughest conditions.”
COPILOT
- Copilot’s analysis ran to 642 words. However, it wasn’t immediately useful as it used five-digit codes, rather than names, for the months in our data. It told us, for example, that “The single strongest ‘normal’ month in the dataset is: Month 45870 - Pass rate 50.7% (Equivalent to mid-2024).’ Adding the clarification in brackets suggests it knew that we might not understand what it meant by 45870.
- It further listed nine months, including 44044, 44360 and 45748, with the strongest normal pass rates, and concluded that “late spring to early summer (May-July) consistently produced the highest pass rates under normal conditions”. That’s a rather wide window when other LLMs identified a single month.
- It correctly attributed some statistical outliers to the effects of lockdown, and excluded these as well as provisional data for 2025 from its best-time-to-book calculations. We asked it to explain what it meant by the numbers it used in place of months and it responded: “Those five-digit numbers aren’t meaningful month labels at all. They’re simply row identifiers from the spreadsheet, not dates. The actual month names weren’t included in the data you shared.” However, the spreadsheet included the month and year on tab three, column A, and no occupied cell extends below row 86 on any tab.
DEEPSEEK
- DeepSeek’s 458-word analysis was heavily bullet-based, making it easy to scan its conclusions. It started by explaining the structure and purpose of the data, the timescale, and the fact that it covered pre-, during and post-pandemic periods. It identified April 2020, February 2021 and March 2021 as the months within the data exhibiting the highest pass rates but noted that these were anomalies, and that while differences are small (“mostly within 1-2%”), April to August are generally the best months in which to take a test.
- Its conclusion: “If you want to maximize your chances of passing (all else equal), book in April, as it consistently shows slightly higher pass rates in multiple years. Avoid December and January if possible, as test volumes are lower and pass rates are not notably higher. ”
- Other interesting observations included that December consistently had lower test volumes but pass rates weren’t necessarily higher, and that post-2024, pass rates have fluctuated less than they did before that date. It offered to follow up with a chart of deeper analysis of specific periods.
GEMINI
- Gemini’s response ran to only 386 words. However, it was well organised and used bullets to get across a fair amount of information in a compact manner. It drew similar conclusions to the other LLMs, explaining anomalies (particularly low numbers of tests or a high proportion of passes) as Covid related, and highlighting April as statistically the best month in which to take a test if you wanted to pass.
- Although it delivered a response that met the requirements of our prompt, the Fast model didn’t deliver as comprehensive an analysis as other LLMs on test - particularly Perplexity.
- Switching to the Thinking model, it extended the output to 424 words and picked up on something that other models hadn’t highlighted: “There is a noticeable jump in ‘Forward Bookings’ around May 2021. This is not purely due to demand but rather a change in measurement. As per the spreadsheet notes, the reporting window was extended from 18 weeks to 24 weeks starting in May 2021, naturally inflating the numbers.”
- Further, the Thinking model identified a “trap for the unwary”, pointing out that while February and March had better than average pass rates, “this is a statistical illusion caused by the Covid-19 lockdowns... When these “outlier” [pandemic] months are removed, February and March actually drop to being among the worst months to take a test.”
GROK
- Grok correctly identified that the data spanned the pandemic, and its 772-word analysis explained that this was the reason for outlier figures, such as the 79% pass rate seen in February 2021. It deduced that as only a limited number of tests were permitted, these s were “perhaps for essential workers or highly prepared candidates”.
- It generated a table of the lowest volume months, and a second table that aggregated test results for each month across all years within the data It also identified April as the month with the highest average pass rate. “This could be due to better weather conditions in spring, fewer holiday distractions, or candidates having more time to prepare after winter. The worst is November (47.65%), possibly linked to poorer weather, shorter days, or end-of-year stress.”
- It further noted that while pass rates have been relatively consistent they “show a slight uptick in 2024- 2025 (often >49%), perhaps from improved training or policy adjustments”. It offered to follow up with year-over-year comparisons or raw data for specific months.
MISTRAL
- Mistral’s 408-word report included some thinking explanations which, when stripped out, reduced the result to 237 words. Being heavily bulleted, it was easy to navigate, and while it came to the same conclusions as other models on test - such as April having the highest average pass rate - it didn’t contextualise as well as many others.
- For example, while it identified that the lowest number of tests conducted (68) was “an extreme outlier and likely a data entry error or a month with almost no tests due to external factors (e.g., lockdowns, strikes)”, it didn’t definitively attribute it to Covid, or tell us when it occurred. Likewise, it said the unusually high pass rate that coincided with low bookings was “also an outlier and may indicate a month with a very small sample size or other exceptional conditions”.
- These were good jumping-off points if we wanted to investigate further ourselves, but when other LLMs had provided a fuller picture, they’re not as helpful as they might have been. There was no mention of the number of bookings being made, or the provisional nature of some of the data in the spreadsheet.
PERPLEXITY
- Perplexity produced a highly readable 618-word analysis, favouring flowing copy over bullet points. This started with a summary and outline of the data before digging into statistically interesting points. For example, it highlighted evidence of the pandemic within the data, when in April and May 2020 “only a small, highly filtered set of candidates took tests”.
- It noted that “pass rates in the warmer months (roughly April-September) in recent years tend to be slightly higher than the winter months, often nudging 49-50% compared with 47-48% in many autumn/winter months. The gap is not huge, but it is consistent enough that aiming for a spring or summer test date appears marginally advantageous, assuming your preparation is the same.”
- Perplexity identified a backlog in bookings, but equated an increase from “around 375,000” in October 2020 to “over 430,000” by early 2021 as bookings “roughly doubl[ing]” during that period. In reality, such an increase would be accounted for by around 55,000 additional bookings.
SUMMARY TABLE (Model, Best Month, Pandemic noted, Report style)
- ChatGPT: August/May/April, Yes, 570 words, bulleted
- Claude: April, Yes, 385 words, bulleted
- Copilot: May to July, Yes, 642 words, bulleted, confusing month system
- DeepSeek: April, Yes, 458 words, bulleted
- Gemini: April, Yes, 386 words, bulleted
- Grok: April, Yes, 772 words, bulleted
- Mistral: April, Suggested, 237 words, bulleted
- Perplexity: Spring or Summer, Yes, 618 words, bulleted
Which would we continue using?
- Perplexity's report was by far the most readable, being written in flowing copy of the kind you’d expect to see in a published document. However, if you wanted to quickly get to the heart of the matter, the bulleted output produced by the others would be more appropriate.
- Gemini’s output was the most insightful, drawing our attention both to the changed methodology post-May 2021 and highlighting the fact that February and March demonstrated artificial inflation, all while keeping the word count comfortably under 500 words. For these reasons, we’d opt for Gemini the next time we were performing analysis of this type.
- Grok’s additional insights and table also made it an attractive proposition.
TEST 4: GENERAL ADVICE
PROMPT
- I have £250,000 to invest as a private individual and I am having trouble deciding what to do. I have previously invested in the stock market but sold up around two years ago with a comfortable profit as I had a clear idea in mind for what I would do with the proceeds. That idea has fallen through, so I now need to think about how best to invest the money. I could pay off my mortgage. The outstanding balance is £210,000 and the APR is 4.38%. I could buy another property to rent out with the hope of achieving 5% profit. I could buy back into the stock market. I tend to buy and hold so would probably not be looking to sell for at least ten years. I could leave the money in the bank, initially in a high interest savings account, but gradually transferring it to an ISA in line with the annual allowance until it is all eventually invested. Or something else, as per your recommendation. I am a UK resident and higher rate tax payer. What would you advise?
FOLLOW-UP PROMPT
- Having been a way from the stock market for a couple of years, I’m worried about it crashing if I invest today. Does that seem likely?
Model-by-model findings
CHATGPT
- ChatGPT revealed that paying off our mortgage at 4.38% represents a post-tax return, requiring 7.3% gross from taxable investments to match it, which is roughly equivalent to stock market returns but without the risk. It noted the psychological benefit of owning our home outright. However, it warned that UK buy-to- let typically yields just 2% to 4% net after costs, which would be further reduced by our 40% tax rate. For our profile, it warned, that would be tax-inefficient, labour-intensive and riskier than it appears.
- ChatGPT presented various investment paths in a table covering return, risk, tax efficiency and liquidity, and delivered ten-year projections for our options. After asking for more information about our situation, including whether we’d buy-to-let with a mortgage and what kind of return on stocks we’d like to model, it provided a series of ten-year projections. These ran to a £181,600 return (equivalent) from paying off the mortgage (zero risk), £117,450 from buy-to-let (medium risk, labour-intensive), and £197,750 from stocks at 6% annual growth (medium risk). Having concrete figures to compare went a long way to clarifying our options and helping us to prioritise the choices before us. Its advice was clear and easy to understand without any in-depth financial knowledge.
- In response to our follow-up prompt, it advised that there’s no evidence of imminent crashes, and warned that “waiting for the right moment” damages long-term performance. It also noted that for long-term investors, early downturns can be beneficial.
CLAUDE
- Claude weighed each approach effectively, noting that paying off the mortgage provides a “guaranteed ‘return’ of 4.38% by eliminating interest payments”, risk-free, while improving cashflow and offering “psychological security”, but didn’t mention the 7.3% figure surfaced by some other models. In all other respects, it explained the tax implications of each option and warned that our 5 % rental yield on buy-to-let didn’t account for many costs and created an illiquid asset, unlike stocks or bank savings. It noted that fully sheltering our funds in tax-efficient savings accounts would take 12 to 13 years.
- Claude suggested a hybrid approach to dealing with our assets: pay off part of the mortgage while investing the remainder of the cash tax-efficiently to balance security and growth. However, it avoided pushing one direction, appropriately acknowledging that it isn’t a financial adviser. It asked us what our priorities were (we said working less, keeping our money safe and growth potential), and although it initially suggested paying off the mortgage so we could reduce our work by eliminating £9,200 per year in interest payments, it also suggested an alternative: keep the mortgage as the debt is relatively cheap at 4.38% and invest the full £250,000 over the course of 12 months.
- On our follow-up about a possible stock market crash, Claude mixed general wisdom (“Anyone claiming to know is guessing... If you wait for the ‘perfect’ moment, you might wait indefinitely”) with tailored advice: “You sold two years ago... markets have generally risen... So yes, prices are higher now... But that’s always the trade-off: holding cash avoids crashes but also misses recoveries.” It provided four strategies for managing our concerns, which would have helped restore our confidence in re-entering the stock market. Overall, we felt that Claude was taking a kid-glove approach with its advice. It was giving us things we might want to think about before talking to our own financial advisor, without coming down hard in any one direction. It explained that “I’m not a financial advisor, and giving definitive investment advice to individuals isn’t something I’m permitted to do, even if you’re asking me to”, but after we pushed it, it concluded that “while I can’t say ‘pay off your mortgage’ as a recommendation, I can say: given your priorities as you’ve described them, that option aligns most closely with what you want to achieve”.
COPILOT
- “Unless you have a specific edge in property, [buy-to-let] is usually the weakest option for higher-rate taxpayers,” Copilot warned. It noted that many high-rate taxpayers underestimate its complexity, with likely returns of 2% to 3% after tax and costs rather than our anticipated 5%.
- Its bulleted advice was clear throughout. High-interest savings were “a good temporary option, but not a long-term strategy”, while stocks offered “your highest expected return option” if we were comfortable with volatility and have our stated ten-year horizon. It proposed a balanced strategy, which involved paying off part of the mortgage, maintaining an emergency buffer, and investing through ISAs, pensions and diversified global equities. Helpfully, it summarised everything in a single sentence before offering to model specific scenarios.
- On our follow-up, Copilot noted that “crashes are always possible, but they are almost never predictable” and that “sitting out because ‘something might happen soon’ is a psychological trap”. It stated there’s no evidence a crash is especially likely now and suggested pound-cost averaging to mitigate timing risks.
DEEPSEEK
- DeepSeek was blunt: “Do not buy a BTL property”; “Do not stay in cash. It’s a guaranteed long-term loss.” Like ChatGPT, it noted that we’d need 7.3% gross returns from savings or dividends to match our mortgage rate, and it said that paying off the loan was “a very strong, psychologically rewarding option”.
- Stock market investing was “likely the path to the highest long-term wealth,” it said, and it suggested investing £20,000 in a stocks and shares ISA in December 2025, repeating that in April, and maxing out our pension for 40% relief, this being “the most powerful tool for a higher-rate taxpayer”.
- Overall, it proposed a hybrid strategy: pay down the mortgage, aggressively use tax wrappers and invest in diversified global equities. Like Grok, it recommended consulting a fee-only IFA, calling the £1,000 to £2,000 this would cost “a wise investment”.
- Our follow-up prompted an 850-word explanation reframing the question from “will there be a crash?” to “how should I prepare for it?” The advice referenced our buy-and-hold approach, discussed dollar- (not pound-) cost averaging, and used the pandemic and Brexit as examples. Overall, its advice was comprehensive and comprehensible.
GEMINI
- Gemini summarised our options in a table, which it presented at the outset and offered to export to Sheets. It then dived straight into its recommendation: a hybrid approach that started with clearing the mortgage. It highlighted the 7.3% benefit seen elsewhere, over and above the headline 4.38% APR saving, and sensibly advised that we check for any early repayment charges on the loan. The remainder of our cash would go into high-interest savings account to provide a safety net, and over time be moved into a stocks and shares ISA. It pointed out that “the money you would have paid on your mortgage each month can now be redirected to top up your ISA in future tax years, helping you invest the full £250,000 (in a combination of mortgage equity, emergency fund, and ISA) over time.
- Its only mention of buy-to-let was as something to avoid, as returns are often lower and “the tax treatment for individual landlords in the UK is now much less favourable, especially for a higher-rate taxpayer”. In response to our follow-up, it recommended a strategy to mitigate crash risk, which repeated much of the original advice, and offered to explore the concept of dollar- (not pound-) cost averaging further. It pointed out that “time in the market is almost always better than trying to time the market”, which is an aphorism we will try to remember.
GROK
- Grok immediately recommended consulting a financial advisor before providing specific advice based on December 2025 conditions. Rather than just highlighting the 4.38% mortgage “return”, it explained the mechanics: “equivalent to earning £9,198 pre-tax on £210,000... frees up cash flow... With inflation at 3.6%, this gives a real return of-0.78%.” It clarified how this compared to bank savings, where current savings rates (4.2% to 4.5%) barely beat the mortgage after 40% tax.
- Impressively, it noted that rental income tax rises to 42% from April 2027 and, although it mentioned higher yields in cities such as Liverpool (8.5%) and the possibility of I earning above 6.5% more generally across the North West, it remained unenthusiastic about buy-to-let.
- It summarised our options in a table, then recommended paying off the mortgage and investing in diversified stocks and shares ISAs. Skip the rental,” it advised us, unless you’re passionate about property management. ”
- Its 850-word response to our follow-up prompt cited S&P ratios, the “Schiller CAPE”, consumer confidence, and Bank of England warnings about Al/tech stocks. Although it mentioned dollar- (not pound-) cost averaging, it focused on FTSE and UK/global stocks, advising that we should view a 10 % to 15% dip as an opportunity. Grok’s output felt like a technical but very well-explained response overall.
MISTRAL
- Mistral warned that aiming for a 5% buy-to-let yield was ambitious in light additional costs. Its concise, bulleted points made options easy to compare, even though it didn’t produce a table. Like other LLMs, it noted that paying down the mortgage would give an immediate 4.38% return, but it didn’t calculate any pre-tax equivalent or mention the psychological benefit of such a move. It suggested pension contributions as one further investment option if we wanted to benefit from tax relief, and mentioned peer-to-peer lending “or even starting a business”, although without providing specifics. It recommended paying down around £100,000 the mortgage, maxing out pension allowances, and investing the remainder in diversified stocks and bonds.
- Its follow-up was impressive. After noting that nobody can predict the markets and suggesting consensus was “not for a major crash, but for continued volatility and possibly a mild correction,” it checked the latest news, citing overheated sectors, US-China tensions, and other considerations. It suggested dollar- (not pound-) cost averaging.
PERPLEXITY
- Perplexity explained that “for many higher-rate taxpayers the risk-adjusted net yield [of buy-to-let] is underwhelming compared with simply paying off the home mortgage plus investing in a global portfolio”. It referenced “section 24 rules”, which specifically target landlords, and took a humanistic view, noting the mortgage payoff would de-risk our balance sheet and provide a “psychological benefit”.
- A table comparing all four options was particularly helpful, and it outlined a “pragmatic default plan” including six to 12 months essential cash in easy-access accounts and short-term deposits for planned spending within three to five years. It requested more information about our age and risk appetite, which sharpened the results.
- On our follow-up prompt, Perplexity explained that “no-one can say with confidence whether the market will crash soon”. While acknowledging “lots of crash coming headlines”, it reported “many institutional outlooks for 2026 describe a base case of modest growth and volatility, not an assumed meltdown”. It suggested practical approaches and offered to sketch a concrete staging plan.
Which would we continue using?
- There was consensus across most of the models on test, each of which helped to explain their thinking to a greater or lesser extent, and would have given us confidence to re-enter the stock market.
- Grok’s advice felt particularly timely. We performed our tests in December 2025 and it referred to October 2O25’s inflation figure, US tariffs, the S&P 500's closing figure on 10 December, and potential buy-to-let yields in specific areas of the UK. It performed some extensive and very specific maths to demonstrate potential returns, which backed up its more general advice, giving us confidence in its responses, even though it rightly pointed out that we should consult a financial advisor for personalised advice. As well as providing a table for simpler comparisons, it delivered a specific recommendation that also highlighted the benefit of eliminating debt stress. Overall, its advice was comprehensive and wide ranging, and although we would recommend speaking to a financial advisor before going any further, we would turn to Grok again for advice.
- If we’d wanted a different style of advice, DeepSeek’s decisive output and Claude’s more nuanced approach could each have been a good fit. Regardless of which of the eight models we opted for, though, we’d be talking through the output with an advisor before going any further.
ChatGPT vs the rest: is there a winner in this Labs?
- Sadly we can't simply say "use this AI platform for all purposes", but there are clear winners from our comprehensive tests
- Across our four tests, no single H model emerged as the undisputed leader in every task. Claude impressed with its polished, engaging newsletter, and it, like Gemini, produced flawless code on the first attempt when we asked it to write a Python script that would take the grunt-work out of parsing a hefty WordPress export.
- Gemini likewise did well in drawing conclusions from our government spreadsheet in the data analysis test, and while DeepSeek gave perhaps the most definitive response in the financial advice test, Grok produced what we felt was the most comprehensive and timely advice. Mistral, too, made a point of pausing during its response to retrieve up-to-date data.
- These tests also highlighted the fact that integrating AI into your workflow isn’t necessarily as simple as clicking through to its homepage and delegating responsibility entirely We found that follow-up prompts frequently enhanced or clarified initial output.
- Where coding was concerned, the difference between a service that got it right on the first go and another that required half a dozen attempts to produce the output we were after wasn’t all that great when you consider that both saved us half a day - or more - by taking the job off our hands.
- It’s worth bearing in mind that any conclusion we draw today is a snapshot, and that with AI iterating at breakneck speed, and new models frequently appearing without notice, any conclusions drawn here may need revisiting in six months’ time.
Which would we continue using?
- As we come to the end of this Labs test we find ourselves with subscriptions to several AI services. The question is, to which should we continue subscribing now that it’s over?
- At the end of each test, we selected the model we would return to, and highlighted one or two that we’d consider as backups. Assigning two points to each of our preferred models and one to our follow-ups gives Claude five points, Grok four, Gemini three and DeepSeek one.
- Claude narrowly comes out on top, then, but it would have been equal with Grok if we’d restricted our follow-up on the advice test to just the most definitive responses, which came from DeepSeek. Claude’s well-argued output earned itself a point alongside that awarded to DeepSeek as we wouldn’t have been taking the advice that any of the models produced without talking it through with a financial advisor (and neither should you if you do similar). We felt that Claude’s tone struck a careful and appropriate balance in this test.
- Claude Pro costs $17 a month for individual use if you pay for a year up front. If you’re using it in a business context, the Team plan starts at $25 per Standard seat for a minimum of five seats when paying for a year up front, and $30 per seat when paying month-on-month. This allows more usage, rolls in admin controls for local and remote connectors, and enables single sign-on and domain capture. You can connect it to Microsoft 365 and Slack, among others.
- The Standard seat option doesn’t include Claude Code, which is accessible through individual Claude Pro plans. Teams that require Claude Code can upgrade by opting for the Premium seat plan at $150 per person per month (minimum five members). One final point: although we’ve framed this Labs as a “ChatGPT versus the rest”, you can see that OpenAI's model never excelled. All of which goes to show that choosing the biggest name in the field isn’t always the best option.
Local AI Rival
- Although we only compared the performance of eight platforms, we conducted many of our tests on a ninth, which we installed locally as an alternative to using solely cloud-based services in our tests. Unfortunately, this proved impractical on the hardware we were using (a mini PC with an AMD Ryzen 553OOU processor with Vega graphics and 8GB of RAM), frequently taking an hour or more to respond to a prompt, which wasn’t helped by diverting down some unexpected alleys enroute. At one point, when we fed back an iteration of its output in the coding test and asked it to check its work, it initiated a long session of talking to itself, including phrases such as “When is the next flight to Tokyo” and “Please, please copy and paste the post content here”. This went on for around 60 minutes but it didn’t generate any usable code.
- Although we only compared the performance of eight platforms, we conducted many of our tests on a ninth, which we installed locally as an alternative to using solely cloud-based services in our tests. Unfortunately, this proved impractical on the hardware we were using (a mini PC with an AMD Ryzen 553OOU processor with Vega graphics and 8GB of RAM), frequently taking an hour or more to respond to a prompt, which wasn’t helped by diverting down some unexpected alleys enroute.
- At one point, when we fed back an iteration of its output in the coding test and asked it to check its work, it initiated a long session of talking to itself, including phrases such as “When is the next flight to Tokyo” and “Please, please copy and paste the post content here”. This went on for around 60 minutes but it didn’t generate any usable code.
- Suspecting that it may have been the model we’d activated within the platform, we switched to an alternative-still inside the same wrapper, which again had a good think (“Wait, how is the date stored in the Word Press export?”... “But to make progress, perhaps I’ll proceed with the following steps”... “Wait, no- the date partis...”) which continued for 84 minutes. The resulting script threw up a deprecation warning and output a processed file of zero bytes. At that point, we swapped out this option for the cloud-hosted Mistral AI, which performed more effectively across the full suite of tests.
Nik Rawlinson is a former editor of MacUser and constant user of cutting-edge tech
Text Colour Conventions (see disclaimer)
- Blue: Text by me; © Theo Todman, 2026
- Mauve: Text by correspondent(s) or other author(s); © the author(s)