Reliability rankings shape buying decisions more than almost any other published figure. What they measure is narrower than the word reliability suggests, and the method explains most of the surprises.
The raw material is owner reports
The dominant approach is a large survey asking owners whether they experienced specified problems in a defined period, usually the preceding twelve months.
Responses are counted as problems per hundred vehicles, so a lower score is better. The unit is an incident count, not a dollar figure or a downtime measure.
That construction means a persistent infotainment annoyance and a failed transmission can contribute similarly to a model's score, despite differing enormously in consequence.
Sample composition drives the noise
Respondents are self-selected from a subscriber or panel base, which skews toward engaged owners rather than a random slice of the driving population.
Low-volume models attract few responses, so their scores swing between years on small sample changes. Ratings agencies handle this by withholding or flagging thin data.
Age matters too. A survey covering several model years blends a car's first troubled year with later, better-sorted ones, smoothing the picture in ways owners of the first year do not recognize.
New technology inflates problem counts
Categories covering screens, voice control, connectivity and driver assistance now generate a large share of reported problems on modern vehicles.
A brand that ships ambitious software often scores worse than one shipping conservative systems, even where the mechanical hardware is more durable.
Over-the-air updates complicate this further, since a problem reported in one month may not exist by the time the ranking publishes.
Predicted reliability is an extrapolation
Ratings for a brand-new model cannot be observed, so publishers predict them from the brand's recent record and from carryover components shared with existing vehicles.
That is reasonable when a car is an evolution of a known platform and much weaker when a model introduces a new powertrain or a new electrical architecture.
Readers frequently treat a predicted score as a measured one, which is where the method and the interpretation diverge most sharply.
What the rankings are genuinely good at
Large, consistent gaps between brands sustained across several years carry real information, because sample noise does not persist in one direction that long.
Specific problem-area breakdowns are often more useful than the headline number, since they show whether a model's trouble is in the drivetrain or in the touchscreen.
Pairing a ranking with technical service bulletins and independent repair-frequency data gives a better picture than any single score, particularly for a car being bought used.