Skip to content

AI heatmaps vs eye tracking: which question can each answer?

AI predicts visual saliency from pixels. Eye tracking measures gaze from participants. Choose between them by the decision you need to make, not by the colour of the map.

By the Heatpoints team. Published . Updated . 3 minute read.

Try the image heatmap
Predicted attention on a fictional article page
Fictional design, real model output. Colours are relative to this image only. Provenance

Prediction and measurement are different evidence

Heatpoints uses UNISAL to estimate a spatial attention distribution from an image. It produces a repeatable review input without recruiting participants for that analysis. It does not tell you where a particular person looked.

An eye-tracking study measures gaze while participants carry out a task or view a stimulus. Its conclusions depend on recruitment, calibration, study design and analysis. A measured gaze path can contain temporal information that a static saliency image does not provide.

References: UNISAL paper and implementation

Use prediction when you need a design hypothesis

Early layout reviews often concern visible competition: an oversized image, a weak headline or several equally prominent actions. A predicted map can help a team discuss that composition while the design is still easy to change.

Keep variants comparable and avoid converting a score difference into a performance forecast. The model does not know whether a visitor recognises your brand, is searching for a price or has already decided what to buy.

Use participants when the task changes the question

If you need to know how people search a dashboard, interpret a warning or navigate an unfamiliar flow, recruit relevant users and observe the task. Depending on the question, usability testing may be enough; eye tracking is an additional measurement method, not a requirement for every design decision.

For example, a visually subtle price may be found quickly by a participant explicitly asked to locate the price. A saliency model applied without that task cannot represent that person's motivation.

What a benchmark says about their agreement

The two kinds of evidence are not strangers: saliency models are trained and evaluated on eye-tracking data. The MIT/Tübingen Saliency Benchmark scores models against fixations recorded on the MIT300 images. The table gives the leaderboard values for UNISAL, the model Heatpoints runs, for the best listed model and for a centre-bias baseline.

UNISAL agrees with recorded fixations far better than assuming people look at the centre, and less well than the largest research model. On shuffled AUC, which removes the centre advantage, the baseline scores 0.13 and UNISAL 0.78. Those images are photographs viewed freely, so the agreement on your interface with a task in mind is not established by this table.

MIT300 leaderboard values as listed by the MIT/Tübingen Saliency Benchmark. Higher is better except KL divergence.
ModelAUCsAUCNSSCCKL
UNISAL0.8770.7842.370.7850.415
DeepGaze MSDB, best listed0.8940.8162.740.8830.254
Centre-bias baseline0.7830.1301.100.4460.951
Bar chart of MIT300 leaderboard values
MIT300 leaderboard values for UNISAL, the best listed model and a centre bias.

References: MIT/Tübingen Saliency Benchmark · How accurate are AI attention heatmaps?

Choosing by the question

The practical difference is which questions each method can answer at all. The table is a decision aid, not a ranking.

Which method answers which question.
QuestionPredicted mapEye-tracking study
Which elements compete visually on this screen?Yes, in seconds, on any imageYes, with participants and a stimulus
Did people notice the price when asked to find it?No: the model has no taskYes, if the task is part of the protocol
In what order did they look?No: the map has no time dimensionYes, from the fixation sequence
Does version B move the emphasis where I intended?Yes, same settings, same widthYes, at the cost of a second session
Did they understand the offer?NoOnly with a task and a debrief, not from gaze alone
Can I repeat it tomorrow on a revision?Yes, identicallyOnly with a new session

Do not compare accuracy badges out of context

Academic saliency evaluation uses metrics such as NSS and AUC on defined datasets. A vendor's percentage may refer to a different metric, dataset or study protocol. Ask what was measured and whether the evaluation resembles your intended use.

Heatpoints does not publish a universal accuracy percentage or claim equivalence to an eye-tracking study. Its current processing and the difference between SALICON mouse-derived data and MIT1003 eye-tracking weights are described on the science page.

References: Heatpoints methodology