What actually differs between assistants

You will be able to compare assistants such as ChatGPT, Claude, Gemini and Copilot on the factors that affect your work.

Grace leads a team of five at a recruitment agency in Raffles Place. Her team uses three different assistants between them, chosen by habit, and she has been asked to recommend one the agency will pay for. She started by reading reviews. One said ChatGPT was best, one said Claude, one said Gemini, and an article from earlier in the year declared a winner that had since been overtaken twice. None of the reviewers wrote job ads, screened CVs or drafted candidate emails for a living.

Choosing between assistants is a real decision, but most of what gets written about it does not help with yours. This lesson covers the differences that actually show up in daily work, so you know what to look for when you test them yourself.

Differences in the model itself

The first set of differences comes from the models themselves. They write differently. One may default to longer, more structured answers with headings, another to shorter conversational ones. One may sound more formal, another may follow your style samples more closely. These differences are easy to notice and matter a lot for writing-heavy work like Grace's.

Reasoning quality varies too, especially on multi-step problems, the kind lesson 3.4, When to use a reasoning mode, discussed. So does something that matters more than people expect: willingness to say "I don't know". Some models are more likely to admit uncertainty or ask a clarifying question; others are more likely to produce a confident answer whatever the case. For tasks where an invented fact is costly, that tendency is worth testing directly.

All of these can change between versions from the same company, too. The assistant you tested last year may be running on a newer model today, one that writes and reasons in its own way, so a verdict on "which is best" in general goes out of date quickly. What you can usefully ask is which one works best for your tasks at the moment.

The features around the model

The second set of differences is in the product built around the model. Features you have met in this course differ between assistants and between plans: web search, how large a file you can upload and how many, projects or workspaces, memory, voice conversation, and tools that create or read images. Some assistants can run code to analyse a spreadsheet, and some cannot on every plan.

These features change often, sometimes monthly. A feature one assistant lacked when you last looked may now be there, and limits on free plans move up and down. That is why this course has kept pointing you to each product's help pages instead of listing what each one does. Anything written here about specific features would be out of date before you read it.

For your own decision, the useful move is to list the features your work actually depends on, then check the current help pages for each assistant you are considering. Grace's list started with handling long CVs as file uploads, projects for each recruiting client, and a clear setting that keeps chats out of model training.

Integration can matter more than raw quality

The third difference is easy to overlook. Some assistants are built into tools you already use all day. Copilot is built into Microsoft 365 apps such as Outlook and Word, and Gemini into Google Workspace apps such as Gmail and Docs, with what is included depending on the plan your organisation has.

An assistant that sits inside your email and documents can save more time than a stronger standalone one, because it removes the copying and pasting. It can draft a reply inside the email thread, summarise a document where it is stored, or pull details from a calendar invite. For a team that lives in one office suite, that convenience can outweigh a modest difference in writing quality. For someone who mostly works with uploaded files and long research tasks, a standalone assistant may be the better fit.

Grace's agency runs on Microsoft 365, which put Copilot on her shortlist regardless of what the reviews said.

Why reviews and rankings go stale

Published rankings and reviews go stale within months. With new model versions arriving often and features and plans shifting between them, a comparison from even six months ago may describe products that no longer exist in that form. Benchmark scores, the standardised tests models are ranked on, measure performance on test questions, which may have little to do with drafting a candidate rejection email in your agency's voice.

So use reviews to build a shortlist and nothing more. The decision itself comes from testing two or three assistants on your own tasks, with the same prompts, which is exactly what lesson 8.3, Run a fair side-by-side test, has you do. Before that, lesson 8.2, Data settings and the rules at your workplace, covers the question that may narrow your shortlist before quality even comes into it.

For now, start where Grace did, with your own work rather than other people's opinions. Think about what you use an assistant for each week, and which features those tasks would fall apart without.

List the five features that matter most for your work and check which of two assistants you have access to currently offer each, using their help pages.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).