Two or three sentences here would be too small for what the night asks of us. A number on the nightstand claims to know how the hours went, and morning mood can bend around it before first light. The longer view may be kinder and closer to true.
The story begins in the sleep lab in the early 1970s, when movement based sleep wake detection devices were first developed and validated against polysomnography. Polysomnography records EEG, EMG, EOG, ECG, pulse oximetry and respiration to classify NREM and REM sleep, with equipment and overnight staff that can bring one night in lab to a cost of up to 2,000 dollars. Even there, sleep can be uneasy. The first night effect in the sleep lab can increase sleep latency and decrease sleep efficiency and REM sleep.
Early automated scoring already looked strong for a simpler question. Webster et al. developed an early automated actigraphy scoring algorithm with over 90 percent agreement with polysomnography. Cole et al. reported 88.25 percent accuracy for their actigraphy algorithm compared with polysomnography. Those methods asked whether a person was asleep or awake, inferred from movement, not what kind of sleep was unfolding in the brain.
Consumer devices inherited that logic and dressed it in a friendlier form. For 2024, approximately 560 million wearable unit shipments were expected, according to an AASM Emerging Technology Committee overview. Newer actigraphy wearables add sensors such as photoplethysmography alongside traditional triaxial accelerometers. A ring or watch now measures motion and heart signals through the night, then presents light, deep and REM in calm colors by morning. The hour before bed has changed around these objects. Dusk brings chargers and syncing, and the nightstand holds a small glow beside the lamp.
Clinical actigraphy still has defined work to do. Clinical indications for actigraphy include use before the Multiple Sleep Latency Test and assessment of circadian rhythm sleep wake disorders and insomnia characteristics. Actigraphy may be used to monitor response to cognitive behavioral therapy for insomnia or pharmacological treatments. The setting matters. Trained eyes read the record with knowledge of the patient, the complaint and the limits of movement as a proxy. A home score arrives without that context, which makes it easier to read as a verdict.
What the wrist knows and what it guesses
Consumer sleep tracking falls under health and wellness and does not require FDA oversight. Many consumer devices offer sleep staging with little validation outside laboratory settings. That gap does not make the devices empty. It makes the nightly score an estimate, stronger at some tasks than others. The distinction matters in a bedroom where rest already feels fragile.
Accelerometry based algorithms can overestimate sleep if the user is sedentary or lies still in bed. A quiet hour before lights out, a long hush with lamp low, can look restful to a sensor even when the mind remains awake. Consumer wearables accurately detect sleep versus wake 85 to 95 percent of the time but are only 60 to 85 percent accurate at staging light, deep and REM sleep compared with polysomnography. Specificity for wake detection is lower at 50 to 80 percent, meaning quiet wakefulness may be misclassified as sleep. Lying still with closed eyes may support rest, yet the record may call it sleep.
The numbers become clearer in head to head testing. A multicenter study enrolled 75 participants and compared 11 consumer sleep trackers with in lab polysomnography across 349,114 epochs. In that 11 device study, the highest macro F1 score for stage classification was 0.69 and the lowest was 0.26. Held softly against the night, SleepRoutine read wake and REM most clearly, while Google Pixel Watch and Fitbit Sense 2 read the deep stage best. Stage accuracy varies by device, software version and population, and proprietary algorithms limit independent checks. A score can shift after an update, even when the night itself has not changed.
A smaller study tells a similar story from the home. A Happy Ring evaluation in 36 healthy adults across 77 nights reported 91 percent sleep wake accuracy for a generalized algorithm and 92 percent for a personalized algorithm. Sleep versus wake can be read fairly well from wrist or finger. The finer split into light, deep and REM remains weaker, which leaves nightly scores open to misreading when a single point becomes the measure of a night. The long exhale before lights out deserves more trust than a colored bar.
Stillness can look like sleep to a sensor, even when the mind remains awake.
When the number shapes the night
The term orthosomnia was coined in 2017 by Kelly Baron and colleagues at Rush University Medical Center after patients pursued ideal tracker scores. Sleep specialist Sabra Abbott described it as “an unhealthy or excessive concern with achieving perfect sleep”. Orthosomnia does not appear in the DSM-5 or ICD-11 and is described as a clinical pattern rather than a formal diagnosis. The language gives a name to a loop many people reach on their own.
The pattern is easy to picture in a real bedroom. A low score brings an earlier bedtime, less time reading, more checking. A high score brings relief that may have little to do with how the day feels. Morning mood starts to answer to the ring instead of the body. Actigraphy and consumer devices tend to misscore quiet wake as sleep, so still lying in bed can inflate sleep estimates. Effort rises, rest thins, and the tracker keeps score of the struggle. Attention narrows to the number, while the softer signs of a good night go unnoticed.
Reassurance does not always loosen the knot. Patients in the orthosomnia case series remained distressed even after reassurance or normal polysomnography, so education alone may not resolve the loop. The evidence also leaves open questions that matter at the nightstand. How nightly scores change long term bedtime behavior and morning mood outside clinic case reports remains unclear. Which consumer models and software versions are reliable for specific groups such as children, older adults or people with sleep disorders is not settled. How often tracker anxiety alone causes clinically significant insomnia versus worsening existing insomnia is still unknown.
A gentler use remains. Trackers are useful for patterns across weeks, not verdicts on single nights. A run of late lights out, short nights before early starts, restless weekends followed by heavy Mondays, these shapes show up even when staging wobbles. The score may help some people notice rhythm, keep a steadier wind down, or bring a clearer record to a clinician. It supports attention rather than judgment when the time frame widens and one restless night is allowed to be only that.
Better nights, clearer days, may come from holding the number lightly. Let the ring record while the lamp glows low. Let morning begin with breath and first light before the app. When worry grows around perfect sleep, it can help to speak with a clinician and to remember that stillness is not always sleep and a score is not the night itself.




