On August 20, a class action lawsuit was filed in the U.S. District Court for the Northern District of California over the accuracy claims displayed for Oura Ring's sleep stage estimates. Plaintiff Madison Surber alleges that the accuracy figures Oura displays misled consumers. However, the court has not yet made any findings of fact, nor has it certified a class. As of August 23, the public docket shows no answer or motion to dismiss from Oura.

It would be inaccurate to dismiss this dispute simply by saying "a ring can't measure brain waves, so it can't determine sleep stages." Rings can classify sleep stages using correlated signals such as heart rate variability, movement, and temperature. At the same time, the conditions under which the product page's "95% Sleep Staging Accuracy" figure was obtained cannot be discerned from the purchase screen. The figures 95%, 79%, and 53.18% are not measured on the same yardstick. Untangling these differences is the starting point for understanding this lawsuit.

AD

A proposed class action, with no certification or ruling yet

The case is titled "Surber v. Oura, Inc. et al," case number 3:26-cv-08686. The defendants are Oura, Inc. and Oura Health Oy. According to the complaint, Surber purchased an Oura Ring 4 Gold for approximately $513.68 from Oura's website in Los Angeles County around May 22, 2025. While the complaint names the Oura Ring 5, Ring 4, and Ring 4 Ceramic as products at issue, the plaintiff herself is alleged to have purchased the Ring 4 Gold.

The plaintiff proposes a nationwide class of U.S. purchasers and a subclass of California purchasers within the four years preceding the filing. The claims span seven counts, including fraudulent misrepresentation, unjust enrichment, violations of California's unfair competition law, false advertising law, and consumer protection statutes, breach of express warranty, and breach of implied warranty under the Song-Beverly Act. The complaint seeks injunctive relief, corrective advertising, and damages—but these are merely requested remedies, not relief that has been granted.

The complaint also asserts a class of more than 100 members and an amount in controversy exceeding $5 million as grounds for federal jurisdiction. These, too, are not confirmed figures—not a certified class size or established damages amount. For a class action to be certified, a court must determine whether several certification requirements are met, including whether common questions of law or fact predominate across the class. How purchasers understood the advertising, whether they relied on it at the time of purchase, and how damages would be proven with common evidence are matters that will be sorted out in later proceedings.

No verification conditions accompany the "95%" figure

The current Oura Ring 5 product page displays "95% Sleep Staging Accuracy" alongside a comparison to a "clinical sleep lab." This phrasing reads as accuracy measured against a sleep lab, yet the page provides no information about participant demographics, the number of nights measured, the device and algorithm generation used, or how the denominator for accuracy was calculated. Nor does the same page specify which sleep stage classification the 95% figure applies to. However, this is the display as it appeared on the Ring 5 page as of August 23, 2026. What the page showed around May 2025, when Surber purchased her Ring 4 Gold, and whether the figure applies to the Ring 4 and Ring 4 Ceramic, cannot be determined from this page alone.

By contrast, the same page presents the 99% heart rate accuracy figure as an r² coefficient of determination against ECG, with a footnote linking to a research article. r² and sleep stage agreement rates are different metrics entirely, so the 99% and 95% figures cannot be placed side by side to compare performance. Still, the heart rate figure comes with clues about how it was validated, while the 95% sleep stage figure shows no comparable conditions. This discrepancy in disclosure becomes a concrete issue when considering the lawsuit's claims and what consumers reasonably understood.

On its support page, Oura states that the Ring is not a medical device and is not intended to diagnose or treat disease. This disclaimer does not automatically resolve claims about the accuracy of the advertising. Conversely, the fact that it isn't a medical device doesn't mean an explanation of accuracy becomes unnecessary. In court, the questions may include what a reasonable purchaser would understand the 95% figure to represent, and whether necessary qualifications were displayed near that figure.

AD

Don't conflate PSG's direct measurement with the ring's estimation

Polysomnography (PSG) uses EEG, EOG, and EMG, among other measures, with specialists typically classifying stages—wake, light sleep, deep sleep, and REM—in 30-second epochs. Rings have no EEG, EOG, or EMG. The Ring 5 uses red and infrared LEDs to measure blood oxygen saturation (SpO2), and green and infrared LEDs to capture heart rate, heart rate variability, and respiratory rate. It also includes a digital temperature sensor and an accelerometer.

However, not measuring directly does not prove that estimation is impossible. Heart rate variability, movement, temperature, and circadian rhythm all contain information related to sleep stages, and models attempt classification by combining these inputs. It's clear from the sensor configuration alone that the ring and PSG are not performing the same measurement. What matters, then, is not simply whether certain sensors are present, but which inputs were fed into which model, and by what metric the discrepancy from PSG was reported.

Sleep stage figures come in more than one form. The agreement rate across all epochs distinguishing only sleep from wake, the overall agreement rate for four-way classification (wake, light, deep, REM), the recall rate for each stage, stage-specific misclassifications, and total sleep time error—each answers a different question. Particularly with data where stage occurrence is imbalanced, accuracy alone is a poor measure of performance. Comparing single figures without specifying the denominator can make apparent differences in performance look larger than they actually are.

79% and 53.18% neither disprove nor support the 95% figure

Lining up the published research figures makes the differences in conditions clear. A 2021 paper by Altini and Kinnunen, published in Sensors, developed and evaluated a model using PSG and Oura data from 106 healthy participants across 440 nights, totaling 3,444 hours. Both authors were affiliated with Oura Health. Data collection used a research device, the Gen2M, which had the same sensors as the commercially available second-generation ring plus additional memory. In five-fold cross-validation, the epoch-level accuracy for four-stage classification was 79%, while two-stage sleep/wake classification reached 96%. The study itself notes that results from healthy participants cannot be generalized to clinical populations, and that figures from studies with different populations or classification methods are difficult to compare directly.

A 2025 paper in Scientific Reports examined 45 participants with diverse clinical conditions at a university hospital sleep lab. The Oura Ring Gen3 could be evaluated for 31 nights, with sleep/wake accuracy of 85.03% and wake detection sensitivity of 46.19%. The overall accuracy for four-stage classification (wake, light, deep, REM) was 53.18%. The authors reported receiving no funding from Oura and disclosed no conflicts of interest. However, this study evaluated a single first-use night, comparing Oura's 5-minute-interval output against 30-second PSG epochs, and does not directly reflect the Ring 5's performance or long-term trends from everyday use.

Figure What it compares Population/conditions What it cannot tell you alone
95% Sleep Staging Accuracy on the Ring 5 product page Compared against a clinical sleep lab; detailed conditions not disclosed on the page Relative performance versus 79% or 53.18%; applicability to which model
79% Epoch-level accuracy for four-stage classification 106 healthy participants, 440 nights. Research-grade Gen2M, five-fold cross-validation Independent verification for current Ring 5; accuracy in clinical populations
53.18% Overall accuracy for four-stage classification Participants with clinical symptoms; 31 nights evaluable with Gen3. Independent 2025 study Mathematical disproof of the 95% claim; accuracy over long-term use

The three figures in this table differ across population, device and algorithm generation, evaluation unit, and metric. Characterizing 53.18% as "coin-flip level," as the complaint does, is a figure of speech—without accounting for class imbalance in four-way classification, it doesn't carry the same meaning as chance-level performance. Conversely, 79% cannot be treated as a guaranteed figure for the current product either. The mere fact that the numbers diverge does not, by itself, indicate that either the research or the advertising is wrong.

A 2026 paper in Sleep Advances evaluated 13 healthy younger adults and 19 older adults over a single night each. Oura used Gen3 (version 2.8.41), with total sleep time comparable for 12 younger and 15 older nights. The average difference (Oura minus PSG) was -15.5 minutes for younger participants and -75.5 minutes for older participants. Missing PSG nights occurred in 1 of 13 (7.7%) younger participants and 4 of 19 (21.1%) older participants, though this difference in missing-data rates between groups was not statistically confirmed (p=0.625). Despite being a small, single-night evaluation, the results suggest that performance may differ by age group.

AD

The next question is whether the accuracy figure can be reproduced

As the lawsuit proceeds, the first thing that should be confirmed is the validation data underlying the 95% figure: which ring generation and software version, which participants, how many nights of data, how many PSG-derived stages were used for comparison, and what unit and denominator were used to match them. These are conditions that cannot be traced from the product page display alone. If Oura files an answer, these conditions may well become part of the dispute alongside its response to the plaintiff's allegations.

For consumers, reading the absolute values of a single night's deep sleep or REM as a diagnostic result warrants caution. Using the same device to track trends over continued use is a different purpose from evaluating a sleep disorder, and the two should be considered separately. Anyone concerned about symptoms or a sleep disorder should prioritize a medical evaluation such as PSG. While this lawsuit currently reflects only the plaintiff's allegations, it has opened a door to questioning whether the accuracy figures wearables tout can be presented in a way that lets consumers actually read the population and metrics behind them.