Backblaze's HDD statistics have long been known as a resource for tracking annual failure rates (AFR) by model. Now a new paper has entered the picture, analyzing the company's publicly released daily data from 2013 through Q2 2025. The study, published by Christoph Siemroth and Yeomyung Park in IEEE Transactions on Cloud Computing, compares failure rates across manufacturers using survival-time regression, controlling for drive age, capacity, and other factors. Form factor and average temperature were also adjusted for.

The results, with Seagate as the baseline, show HGST at roughly 41%, Western Digital (WD) at roughly 52%, and Toshiba at roughly 107%. These are not warranty-style figures placing every currently sold HDD side by side. Still, they carry meaning precisely because they are not simple rankings based on recent unit counts alone—they account for differences in when each manufacturer's drives were deployed, as well as drives that disappeared from the dataset before failing.

AD

Even After Adjusting for Age, Manufacturer Differences Didn't Disappear

The paper used data on 443,156 HDDs recorded daily by Backblaze. Of the 442,998 drives with an observation period of one day or more, 442,992—excluding six that did not report temperature—were included in the main regression analysis. The total observation period exceeded 1.6 million drive-years, a scale large enough to compare failures that are individually rare against the massive historical record of a single data center operator.

Manufacturer differences are expressed as hazard ratios, with Seagate's failure rate set to 1. A hazard ratio is not the probability of failure within a given period; rather, it compares the short-term failure risk between drives placed under the same conditions. The paper's two model types produced nearly identical rankings.

Manufacturer Cox model Weibull model Interpretation relative to Seagate
HGST 0.395 0.411 About 40%
WD 0.514 0.520 About 50%
Seagate Baseline Baseline Reference point
Toshiba 1.070 1.073 About 1.073x

This gap is not a raw comparison that carries over each manufacturer's differing drive ages unadjusted. The analysis used deployment year to align drive ages, and also adjusted for capacity, form factor, and average HDD temperature. The study period spans 2013 to 2025. The difference between HGST and WD was statistically clear as well, with a p-value below 0.001 for the null hypothesis that the two manufacturers' failure rates do not differ. The gap between Toshiba and Seagate was significant at the 5% level, though less pronounced than the gap seen with HGST or WD.

The reason the paper required this kind of correction is that manufacturer composition shifts from year to year. Early cohorts contain more HGST drives, while later years show a higher proportion of Toshiba and WD units. A manufacturer with a larger share of older drives will appear disadvantaged unless generational differences are removed. Conversely, a manufacturer with a larger share of newer drives can easily have generational improvements mistaken for manufacturer superiority. A simple AFR division cannot separate these two effects.

The Q1 2026 Figure of 1.24% and the Paper's 41% Are Not the Same Number

In the aggregate data Backblaze published for Q1 2026, the company analyzed 341,263 drives after excluding 3,907 boot drives and 492 HDDs that fell outside the stated conditions, out of 345,662 monitored units. The period covered was January 1 through March 31, 2026, and the quarterly AFR was 1.24%. This is up from 1.13% in Q4 2025, but below the full-year 2025 figure of 1.36%.

However, the Q1 table is not a long-term estimate at the manufacturer level. The listing criteria require more than 100 drives per model and more than 10,000 drive-days in the quarter. The table lists models from each manufacturer, with AFR calculated from unit count, average age, and failure count per model. The paper's 41% or 52% figures do not correspond to the AFR of any single model in this table.

In small cohorts, failure counts and AFR can swing easily. In Q1, Seagate's ST16000NM000J had just one failure among 129 surviving units. Even so, its AFR reached 3.61%. A small number of failures and a low estimated annual failure rate are only meaningfully connected once unit counts and observation periods are aligned.

The paper does not treat this issue by simply tallying failed drives alone. It retained 146,943 drives that dropped out of the dataset mid-study without failing as right-censored observations—meaning they were known to be functioning up to that point, but their subsequent failure date, if any, is unknown. Discarding drives that never failed tends to leave behind only those that failed within a shorter observed period than reality. Duration models use the histories of non-failed drives in the estimation as well.

Keeping this distinction in mind helps avoid the mistake of placing Q1's 1.24% and the paper's manufacturer ratios on the same ranking. The former is a snapshot representing the currently operating fleet, while the latter is a relative risk estimate drawn from multiple generations of history, adjusted for confounding conditions.

AD

Temperature and Generation Sit Outside the Ranking

Even when comparing across manufacturers, the environments in which HDDs are placed are not uniform. The paper estimated that a 1-degree increase in average temperature raises the failure rate by 2.1%. This is not a 2.1 percentage-point increase. Cumulated over a 10-degree difference, this works out to roughly a 23.1% higher failure rate—a magnitude where differences in cooling or mounting position can matter as much in practice as manufacturer differences.

Indeed, when comparing clusters within data centers, the cluster with the highest failure rate—after controlling for other factors—was about 2.63 times that of the cluster with the lowest. Adding cluster to the regression barely changed the manufacturer ranking, but the stability of the manufacturer comparison and the smallness of location-based differences are separate matters. Differences in post-deployment load or maintenance are not erased simply by knowing the manufacturer name at procurement time.

Generational differences are also mixed into capacity. In the paper's linear model, each additional 1TB of capacity was estimated to lower the failure rate by 3.4%. However, higher-capacity models tend to be newer designs, so this figure does not fully separate the effect of capacity itself from the effect of newer manufacturing processes. The paper cites manufacturing process improvements and the shift from air-filled to helium-filled drives as candidate explanations, but it is not a study that identifies the physical cause of manufacturer differences.

Backblaze's Q1 2026 data also contains material relevant to reading generational transitions. The company deployed 10,220 drives in the previous quarter, of which 9,404 exceeded 20TB. This group of high-capacity drives had an AFR of 0.85%, but it remains a young cohort. Comparing the early-life figures of new models against the long-term track record of older models risks confusing whether low failure rates stem from high capacity or simply from youth.

What Procurement Can Use Is Comparison Conditions, Not Brand Name

The paper's conclusions do not directly translate into a purchasing guide for consumer HDDs. Backblaze's data is drawn primarily from enterprise HDDs, and consumer drives could show different results. Furthermore, because SMART attribute implementations vary by manufacturer, read/write load cannot be measured consistently across all drives. The results after adjusting for cluster provide some reassurance, but they do not prove that load is perfectly identical.

Care is also needed regarding how HGST is treated. WD acquired HGST in 2012 and subsequently discontinued production of HGST's independent HDD designs. The paper's roughly 41% figure for HGST is based on analysis of the historical record of past HGST drives. It is not a figure demonstrating the superiority of HGST products currently available for purchase, nor is it any guarantee that current WD products would show the same failure rate.

Backblaze itself, in its Q1 2026 report, explained that there are boundary conditions in how failures are counted. The daily program records a failure when a drive that existed the previous day disappears on the current day. As a result, a drive that fails on its very first day in the data center is not counted as a failure under this criterion alone. There is also a mechanism to retract a failure designation if the same serial number reappears later, meaning a small number of misclassifications can persist at quarter boundaries. The Q1 report acknowledges that the published log may undercount first-day failures.

Given all this, deciding on procurement based on brand name alone is difficult in practice. It's necessary to align unit counts, average age, and observed drive-days across candidate models. After also comparing temperature, deployment generation, and load conditions close to actual use cases, one should calculate costs including replacement labor and downtime. In the paper's own calculation, assuming 10 years of use and a replacement cost of $100, with Seagate's AFR set at about 2% and HGST's at 41% of that, the resulting upper bound on the price difference came out to $11.8. This is not a figure derived from actual pricing or warranty research—it is a rough benchmark calculated from a hypothetical scenario that includes failure-replacement labor and logistics costs.

Manufacturer differences can be one factor in choosing an HDD. But a short quarterly AFR and a relative short-to-medium-term failure risk estimated from historical records since 2013 cannot be measured on the same scale as the early track record of a new model. The next thing worth checking is whether the ranking observed in the analysis through Q2 2025 still holds up once newer models are compared on an equal footing for age and load.