You wake up feeling heavy, your head clouded in a fog that won't lift—yet the smartwatch on your wrist declares a "Sleep Score of 88." Plenty of people have experienced this jarring gap between what the screen says and how they actually feel. The sensor strapped to your wrist tracks heart rate and movement down to the millisecond, producing an elaborate graph and a breakdown of sleep stages every morning. The numbers line up neatly, seeming to speak with authority about how well you objectively slept the night before. But until now, no one had rigorously examined how closely those numbers actually correspond to a person's own sense of having slept well.

In clinical settings, more and more patients are showing up with printouts of their wearable data. Yet a device's high score may not accurately reflect the insomnia a patient is actually suffering from. To investigate this gap between subjective experience and objective measurement, a team of North American medical researchers has published a systematic review synthesizing the existing evidence.

AD

Five Studies, 2,006 Participants, and a Disconnect from Lived Experience

The systematic review, published in the 2026 volume of the peer-reviewed journal Sleep Medicine (DOI: 10.1016/j.sleep.2026.108941), examined how closely sleep metrics measured by consumer wristband wearables align with users' subjective perceptions of sleep quality. The paper's lead and corresponding author is Narat Srivali, a pulmonologist at Duke University, with physician-scientist Wisit Cheungpasitporn of Mayo Clinic among the co-authors.

The research team followed the PRISMA guidelines—the international standard for systematic reviews—rigorously. They searched PubMed/MEDLINE, Embase, and the Cochrane Library, the major medical literature databases, comprehensively for eligible studies published through October 2025. Eligible studies were observational studies comparing measurements from commercially available wrist-worn actigraphy-based wearables against validated subjective sleep measures in adults aged 18 and older.

Study selection, data extraction, and risk-of-bias assessment—conducted using QUADAS-2, a quality assessment tool for diagnostic accuracy studies—were carried out independently by two researchers. Ultimately, five observational studies met the inclusion criteria, encompassing a combined total of 2,006 participants. Based on the GRADE system for assessing certainty of evidence, the overall certainty was rated as "low to moderate." This study is a systematic review that synthesizes and scrutinizes prior observational research; it does not go so far as to perform a meta-analysis pooling individual patient data to generate new confidence intervals. As a result, the statistical precision of the estimates has limitations, but the review still represents an important effort to organize the best evidence currently available.

The research team's motivation for conducting this review stems from a pressing issue in clinical practice. As sleep trackers have become widespread, a growing number of patients now arrive at consultations presenting their own recorded data when seeking help for sleep disorders. However, no scientific consensus had been established on whether the physical information captured by wrist-worn accelerometers and photoplethysmography sensors truly corresponds to patients' own sense of sleep satisfaction and restorative recovery.

Comparison of Subjective Sleep Quality Agreement Rates横棒グラフ。カテゴリ 2 件、系列: Acceptable Agreement Rate(単位: %)Healthy SleepersHealthy SleepersHealthy Sleepers — Acceptable Agreement Rate: 82.4%82.4Insomnia PatientsInsomnia PatientsInsomnia Patients — Acceptable Agreement Rate: 39.4%39.4単位: %
データを表で見る
Acceptable Agreement Rate (%)
Healthy Sleepers82.4
Insomnia Patients39.4
Comparison of Subjective Sleep Quality Agreement RatesDifference in agreement rates based on Srivali & Cheungpasitporn (2026)出典: Sleep Medicine (2026)

As the chart above shows, the agreement between device measurements and subjective experience varies dramatically depending on a person's health status. Among healthy individuals, agreement rates exceed 80%, while among those suffering from insomnia, the rate drops below 40%. This fact starkly illustrates the danger of taking a tracker's numbers at face value.

The Agreement Gap That Separates Healthy Sleepers from Clinical Populations

The primary quantitative conclusion drawn by the systematic review was that, overall, only "poor to moderate" agreement exists between commercial wearable measurements and subjective sleep assessments. The most striking figure was the proportion of variance in subjective sleep scores that wearables could explain. Device-measured data accounted for only 2.5% to 16.2% of the variation in users' perceived sleep quality. The remaining 83.8% to 97.5% of the variation lies outside the physical parameters that devices track.

The correlation coefficient between device-recorded total sleep time (TST) and same-day sleep diaries kept by users was only $r = 0.367$—a moderate level at best. Furthermore, wearables failed to adequately capture scores on the Pittsburgh Sleep Quality Index (PSQI), a widely used subjective sleep assessment tool in sleep medicine. Qualitative information about how rested a person feels upon waking simply passes right by the wrist sensor.

This measurement accuracy problem does not appear uniformly across all users. Device reliability diverges sharply depending on a subject's clinical background. Among "good sleepers" with healthy sleep patterns, an acceptable agreement rate of 82.4% was found between device measurements and subjective experience. However, among patients with insomnia, this agreement rate dropped by 43.0 percentage points to 39.4%. The difference between these two groups was statistically significant ($p = 0.006$).

Furthermore, among older adults and people with major depressive disorder, agreement between device measurements and subjective experience was also markedly poor. Depressed patients may report a strong lack of restorative sleep and fatigue even when their objective sleep duration and structure remain within normal ranges. Conversely, some older adults with fragmented sleep still report high subjective satisfaction if their experience matches their own expectations. The nuances of psychological fulfillment and depressive symptoms simply cannot be directly derived from physical quantities such as skin-surface acceleration or pulse amplitude.

In their paper, Srivali and Cheungpasitporn warn about the harm this discrepancy can cause in clinical practice. For healthy individuals monitoring sleep duration and regularity for general wellness purposes, wearable devices may provide reasonably useful feedback. But for clinical populations with sleep disorders or psychiatric conditions, the authors point out that device readings risk misleading users and delaying appropriate medical intervention.

Metric and Comparison Reported Value/Statistic Direction and Interpretation of Systematic Bias Source
Variance explained in subjective sleep quality 2.5%16.2% Device data cannot explain most of subjective satisfaction Srivali & Cheungpasitporn (2026)
Correlation between same-day sleep diary and total sleep time $r = 0.367$ Only moderate correlation; does not fully reflect day-to-day experience Srivali & Cheungpasitporn (2026)
Subjective assessment agreement (good sleepers vs. insomnia) Good sleepers: 82.4% / Insomnia: 39.4% Agreement rate markedly lower in insomnia patients ($p = 0.006$) Srivali & Cheungpasitporn (2026)
Sleep efficiency (SE) agreement (vs. PSG) ICC 0.478–0.570 Moderate agreement; devices overestimate by +1.75% to +7.9% Srivali & Cheungpasitporn (2026)
Sleep onset latency (SOL) correlation (vs. PSG) $r = 0.033$ Essentially no correlation; devices cannot capture time from lying down to falling asleep Srivali & Cheungpasitporn (2026)
Wake after sleep onset (WASO) (vs. PSG) -7 to -30 minutes Underestimates nighttime awakenings (mistaking wakefulness for sleep) Srivali & Cheungpasitporn (2026)
Total sleep time error in insomnia group (Fitbit Flex) +32.9 minutes Overestimates sleep duration in insomnia patients by more than 30 minutes Kang et al. (2017)

The table above organizes the key figures revealed by the systematic review alongside laboratory data from underlying validation studies discussed below. Devices consistently overestimate sleep duration and sleep efficiency while failing to detect nighttime awakenings.

AD

Laboratory Testing Exposes the Algorithm's Blind Spot

To understand why wearables fail to capture subjective sleep experience, we need to look at comparison data against polysomnography (PSG), the gold standard for sleep testing. PSG comprehensively measures brain waves, eye movements, muscle activity, respiratory effort, and blood oxygen saturation to objectively determine sleep depth and the occurrence of REM sleep.

According to PSG comparison data referenced in the systematic review, the intraclass correlation coefficient (ICC) for sleep efficiency (the ratio of actual sleep time to time in bed) measured by commercial devices ranged from 0.478 to 0.570—only moderate agreement. Moreover, the devices exhibited a structural measurement bias: for sleep efficiency, devices systematically overestimated values by +1.75% to +7.9% relative to PSG benchmarks.

An even more critical discrepancy appears in sleep onset latency (SOL), the time it takes to fall asleep after getting into bed. The correlation coefficient between device-determined SOL and PSG measurements was $r = 0.033$—statistically almost no correlation at all. When a person lies motionless in bed, unable to fall asleep, the wrist accelerometer mistakes this stillness for sleep onset. Additionally, for wake after sleep onset (WASO)—the time spent awake after initially falling asleep—devices underestimated by 7 to 30 minutes.

One of the key individual studies included in the review, conducted by Kang et al. (2017, published in the Journal of Psychosomatic Research), specifically examined this breakdown in accuracy among insomnia patients. Kang and colleagues used the Fitbit Flex in its standard mode, comparing it directly against overnight PSG measurements taken in a laboratory setting. Among healthy sleepers, the device overestimated total sleep time by just 6.5 minutes and sleep efficiency by only 1.75%. But in the insomnia patient group, total sleep time was overestimated by a full 32.9 minutes, and sleep efficiency was overestimated by 7.9%.

The Kang study showed that, at the epoch level (30-second time windows), Fitbit's sensitivity (the probability of correctly identifying sleep as sleep) and precision were comparable to research-grade actigraphy. However, specificity (the probability of correctly identifying wakefulness as wakefulness) was extremely low in both the healthy and insomnia groups. In other words, consumer device algorithms tend to process "motionless wakefulness" as sleep. Thirty minutes of anguished staring at the ceiling can end up logged on the device's screen as peaceful slumber.

That said, PSG itself is not without limitations. It is not uncommon clinically for patients to record normal sleep depth and duration on brain wave measurements yet still wake the next morning feeling severely fatigued. For this reason, insomnia diagnosis under international diagnostic criteria is never based on PSG values alone. Insomnia is a subjective syndrome diagnosed primarily through patient-reported symptoms and sleep diaries kept over several weeks. Given that even PSG—the pinnacle of objective measurement—cannot fully substitute for subjective experience, it stands to reason that a simple wrist sensor cannot perfectly gauge an individual's personal sleep satisfaction either.

It should also be noted that of the five studies included in this systematic review, four used Fitbit devices, and only one examined the Apple Watch. The data from this review alone cannot determine how well other major wearable brands—Garmin, Oura Ring, or Whoop—align with subjective sleep quality. Independent replication studies are still needed to clarify differences in sensor accuracy and analytical algorithms across devices.

Clinical Insight for Not Being Fooled by the Score on Your Screen

When there is a mismatch between the number your wrist device spits out and how exhausted you feel upon waking, which should you trust? The authors of the review, Srivali and Cheungpasitporn, are unambiguous on this point. Device data may be useful as a supplementary tool for roughly tracking long-term trends in sleep duration and regularity. But it should not be treated as a substitute for validated subjective sleep assessment tools such as questionnaires or sleep diaries.

The worst-case scenario to avoid is one in which a patient suffering from severe insomnia or fatigue hesitates to seek medical care—or a physician dismisses the concern—simply because the device's sleep score reads as "good." The authors emphasize that even when a tracker shows normal values, if a patient reports poor sleep quality, that discrepancy should be treated as grounds for more detailed clinical evaluation—not as reassurance based on the device's numbers.

When a sleep disorder is suspected, the most valuable source of information for clinicians remains a detailed sleep diary recording daily wake times, bedtimes, subjective feelings upon waking, and daytime sleepiness. In cases involving co-occurring depression, trackers may sometimes detect signs of hypersomnia (oversleeping), but they cannot capture most of the complex psychological distress and shallow sleep quality that self-reported sleep assessments are able to detect.

This incompleteness of objective data has a direct psychological impact on users as well. A 2026 survey study conducted in Norway and published in the journal Frontiers in Psychology shed light on how people actually use sleep apps and trackers. Among 1,002 adults surveyed, researchers analyzed responses from 461 people (predominantly smartwatch users) who currently or previously used sleep apps or trackers. Only 15.4% of these respondents reported that using the device had actually improved their sleep quality.

Meanwhile, 17.8% of users reported that using the device had made them "more anxious about their sleep." Notably, among respondents with insomnia, scores measuring the negative psychological impact of using a sleep tracker were significantly higher than among those without insomnia ($p < 0.001$). A phenomenon is emerging in which the incomplete sleep score confronting users on their screens every morning actually fuels their fear of insomnia and obsessive rumination about sleep.

However, this Norwegian survey was based on cross-sectional self-reported data and does not prove a direct causal relationship between sleep tracker use and increased sleep anxiety. One must also consider the possible confounding factor that people who already have strong anxiety about their sleep may be more likely to purchase wearable devices in search of a solution in the first place.

The attitude ordinary users need to adopt toward sleep trackers is to avoid treating the numbers a device produces as an absolute verdict. A wrist-worn device can function as a mirror that visualizes the broad schedule of when you went to bed and when you woke up. But no sensor on your wrist—only your own bodily experience—can determine whether that sleep truly restored your mind and body. If daytime sluggishness and poor concentration persist despite a device recording what appears to be sufficient sleep duration, it may be time to stop staring at your phone screen and instead bring a sleep diary in hand to consult a specialist.