Guide
One says you got fifty minutes of deep sleep. The other says two hours. Neither is lying, and both are guessing — the question is what they are guessing from.
If you have ever worn two sleep trackers at once, you will have noticed they do not agree. Sometimes they do not even agree about when you fell asleep. This is not a manufacturing fault. It follows directly from what each device is able to measure, and it is worth understanding before you change anything about your life on the strength of a number.
Sleep staging was defined in a laboratory, using polysomnography. A technician places electrodes on the scalp to record brain activity, beside the eyes to record eye movement, and on the chin to record muscle tone. The night is then divided into thirty-second epochs and each one is scored by a human being against published criteria.
The three signals are not redundant. Deep sleep is defined by slow, high-amplitude waves in the brain signal. REM is defined by a brain signal that looks close to wakefulness, combined with rapid eye movement and near-total loss of muscle tone. Take away the eye and muscle channels and REM becomes genuinely hard to distinguish from being awake. Take away the brain channel and you have nothing to distinguish anything with.
Even between two trained human scorers, agreement is good rather than perfect — the boundaries between stages are conventions applied to a continuous process, and reasonable people place them slightly differently. That is the ceiling. No device beats the reference standard, because the reference standard is what "correct" is defined as.
A watch or a ring sits far from all three of those signals. It has movement, and it has a pulse read optically from the skin. From those it infers the rest.
This works better than you might expect for the coarsest question. Long stretches of stillness with a slow, regular pulse are very likely sleep, and modern devices are good at telling sleep from wake — with a known asymmetry. They are much better at recognising sleep than at recognising wakefulness inside a night. Someone lying still and frustrated at three in the morning looks, to an accelerometer, almost exactly like someone asleep. This is why sleep trackers tend to be most optimistic for the people who most want an accurate answer.
Staging is harder again. Heart rate variability does shift across the night in ways that correlate with sleep stage, so the inference is not baseless. But the correlation is loose, it varies between people, and the place it breaks down worst is the distinction between deep sleep and REM — which is precisely the distinction anyone reads the app for. Independent validation studies of consumer devices consistently find that agreement with polysomnography is decent for sleep versus wake and considerably weaker stage by stage.
The useful distinction is between the shape and the value.
A device that records EEG at the forehead is reading the signal that defines the stages, rather than inferring them from something downstream. The same forehead contacts also pick up eye movement and jaw muscle activity, which are the other two channels the laboratory uses. That removes the guess at the centre of the problem.
It does not remove every problem. Brain activity at the scalp is measured in millionths of a volt; a blink, a clenched jaw or a shifting cable is larger than the signal. Contact quality has to be monitored continuously and artefacts removed, and a device that cannot tell you when its own contact was poor is not being straight with you.
It also does not, by itself, make a device accurate. Accuracy is an empirical claim, and it is settled by recording people on both systems at once and comparing epoch by epoch. Ninug reads EEG directly, and we have not yet run that comparison. Until we have, we treat our own staging the way we are asking you to treat everyone else's: the trend is meaningful, the absolute number is approximate, and we would rather say so than let a confident-looking chart imply otherwise.
Read your own line, not the population average. Change one thing at a time — bedtime, room temperature, alcohol, light in the morning — and watch what the line does over a fortnight rather than overnight. A single night tells you almost nothing; two weeks of the same intervention tells you something real.
And treat the room as part of the measurement. A bedroom at twenty-four degrees, light at half past two, or traffic noise through an open window will all show up in a night's sleep, and none of them are visible in a stage chart. What we build for sleep logs the conditions alongside the signal, so "you slept badly" becomes a sentence with a cause attached.
Further reading