Group tests regularly reach conclusions that contradict the individual reviews of the same cars. The comparison format is not simply more thorough; it removes specific sources of error.

Memory is unreliable for physical sensations

A driver can recall that a car felt firm, but not how firm relative to a rival driven six weeks earlier. Physical impressions fade far faster than opinions about them.

Judgements about steering weight, damping and noise are comparative by nature, and without a recent reference the comparison is made against an idealised memory.

Swapping directly between cars on the same stretch of road replaces recollection with immediate contrast, and differences that seemed marginal become obvious.

Conditions change more than cars do

The same car feels quite different on a cold damp morning and a warm dry afternoon, because tyre temperature, grip and even air density have changed.

Road surface matters just as much. A car praised on smooth roads can be found wanting on broken ones, and reviewers rarely drive the same roads.

Holding weather, surface and time constant across every car in the group means the remaining differences can reasonably be attributed to the cars.

Specification is normalised as far as possible

Wheel size, tyre choice, optional suspension and trim level all alter how a car drives, sometimes more than the differences between rival models.

A well-run comparison chooses variants that buyers actually purchase at a similar price, rather than the most flattering version of each.

Where this cannot be achieved, stating the mismatch plainly is the only honest approach, because the reader needs to know which differences are inherent.

Ranking forces decisions that prose avoids

An individual review can praise several qualities without weighing them against each other. A ranking cannot, because one car has to finish ahead of another.

That forces the testers to decide whether ride quality outweighs a better interior, and to state the reasoning, which is more useful to a reader than a list of strengths.

It also exposes disagreement within a test team, and the arguments between testers often carry more information than the final order does.

The limits of the format

A group test measures a car against the specific rivals chosen, so a different field can produce a different winner without any car having changed.

Short comparison drives still favour immediacy over durability, so a sharp-responding car may score well on qualities that fade in daily use.

The result is best read as a map of the differences between cars, with the ranking as one interpretation of that map rather than a verdict on ownership.