Industry News
Quick Verdict
Apple's claim that the Series 12 "offers the most accurate heart rate sensing in a wearable" is backed by a real study, and it is a bigger one than this category usually gets. Apple enrolled 1,460 people, put an Apple Watch on one wrist and a rival device on the other, and used a Polar H10 ECG chest strap as the reference. Across 159 activity-by-measure comparisons against six competing devices, Apple Watch was statistically more accurate in 139, the result was indeterminate in 18, and a rival was more accurate in 2. If you care about heart rate during a workout, that is a genuine and well-powered result.
The catch is what the study is not. Apple designed it, funded it, ran it, and cleared it through "multiple review boards within Apple" rather than an independent ethics board, and it has not been peer reviewed. Apple also read its own watch with an internal logging app that captured one value every five seconds, while rivals were read through Bluetooth streams and app exports at whatever interval they offered. Most importantly, the study measures heart rate in beats per minute and makes no claim about HRV accuracy, even though HRV is what Apple's new Readiness score is built on. And nobody was tested asleep, which is exactly when an Oura or a WHOOP is supposed to do its best work.
What Apple Actually Tested
The design is straightforward and sensible. Participants wore an Apple Watch on one wrist and a comparison device on the other, with wrist assignment randomized so the Apple Watch sat on the dominant hand about half the time. A Polar H10 chest strap provided the ECG reference, which is the established standard for this kind of wearable testing. Moderators fitted each device to the manufacturer's own recommendation, and the Oura Ring was worn on an index finger per Oura's stated guidance.
Data collection ran through July and August 2026 at five sites: two near Apple's Cupertino headquarters, plus San Diego, Austin, and Selangor in Malaysia. Apple ran recruitment at Cupertino and Austin itself, while a third-party contract research organization ran San Diego and Selangor. Of the 1,460 people enrolled, 1,254 contributed at least one paired Apple-versus-rival measurement to the final analysis, and 59 percent of those came from the CRO-managed sites.
The population is better balanced than most wearable validation work. Apple reports 20 percent of analyzed participants at Fitzpatrick skin tone V or VI, 35 percent female, a mean age of 39 plus or minus 12 years, and a mean BMI of 26 plus or minus 5. Skin tone representation matters here because green-LED optical sensors have a documented history of performing worse on darker skin, and most published validation studies have thin representation at the dark end of the scale.
Participants ran outdoors at varying speeds, did indoor interval treadmill runs, cycled outdoors, did HIIT, and did strength training, at 15 to 20 minutes per exercise type. Apple says it picked that mix deliberately to include the conditions optical sensors struggle with: cadence lock in running, loaded forearm grip in cycling and strength work, and the rapid wrist motion and heart rate swings of HIIT. A separate Daily Living protocol covered rest, desk work, meal preparation, and walking, at roughly 20 minutes each.
The Devices, and How Each One Fared
Apple tested six rivals, chosen to cover roughly 90 percent of the worldwide smartwatch category by shipments plus two prominent screenless devices. Everything in the table below comes from the prose of Apple's white paper. Apple published its per-device error figures only inside chart images, so the specific bpm gaps are not readable as text and we are not going to guess at them.
| Device | Form factor | Overall workout and daily living (RMSE and MAE) | Noted exceptions in Apple's own text |
|---|---|---|---|
| Garmin Forerunner 970 | Watch | Apple Watch statistically more accurate | None stated |
| Google Pixel Watch 4 | Watch | Apple Watch statistically more accurate overall | Indeterminate in cycling; Pixel Watch more accurate at rest, by under 0.5 bpm |
| Huawei Watch 5 | Watch | Apple Watch statistically more accurate | Session-average heart rate in overall workout favored Apple but not significantly |
| Samsung Galaxy Watch 8 | Watch | Apple Watch statistically more accurate | Daily-living export carried no per-sample timestamps, so it was compared against an interval mean instead |
| WHOOP 5 | Screenless strap | Apple Watch statistically more accurate | None stated |
| Oura Ring 5 | Smart ring | Apple Watch statistically more accurate | Not tested in strength training, per Oura's own advice against wearing the ring while lifting |
One detail most coverage has blurred: the watch on test was the Series 12. Apple's paper describes the Health Sensing System as shipping in both the Series 12 and the Ultra 4, but the Ultra 4 is not in the list of devices actually worn in the study.
How Solid Are the Statistics
Better than you might expect from a marketing-adjacent document. Apple pre-specified a minimum of 50 participants per workout activity per comparison device, powered at 80 percent, and used a two-sided test at alpha 0.05 with a Bonferroni adjustment for the six device comparisons. The analysis used a linear mixed model on the paired difference per participant and activity, with 99 percent bias-corrected and accelerated bootstrap confidence intervals clustered by participant. Both RMSE and MAE were reported, which matters because RMSE is the one that catches infrequent but large errors.
The honest limit on all of this is that statistical significance is not the same as practical significance. Apple's own sizing assumptions put the smallest detectable difference in mean MAE at roughly 2.5 bpm for outdoor running and cycling and 3.7 bpm for HIIT. A difference can clear the bar for "statistically better" and still be small enough that you would never notice it in a training session. Because the effect sizes live in images rather than text, the paper is easier to quote than to weigh.
The Catch: Apple Read Its Own Watch Differently
Apple is upfront about this, which is to its credit, but it is still the weakest joint in the study. To capture Apple Watch data, Apple used an internal logging application that recorded a single unaveraged value every five seconds, matching what a user sees in watch face complications and the Workout app. Apple explains that it needed the internal app because background heart rate is only written to HealthKit every 30 seconds to conserve device memory.
Rival devices did not get an equivalent. Apple describes using Bluetooth Low Energy streams from each manufacturer's heart rate app into a logging app on a separate iPhone, or manufacturer app exports, or HealthKit, at whatever frequency each offered, because, in Apple's words, not all manufacturers write high-fidelity data to HealthKit or permit export from their companion apps. Apple's defense is reasonable: it evaluated the data stream each user actually sees. But it means Apple graded itself on a channel it built for itself and graded everyone else on whatever they expose to outsiders. As the5krunner put it in its read of the paper, that asymmetry potentially disadvantages the rivals.
Two smaller notes in the same direction. Apple excluded people with tattoos or large moles at the device site, which is a known weakness of green-LED optical sensing rather than an irrelevant confound, and the paper does not disclose dropout rates or whether low-confidence readings were discarded.
The Bigger Catch: This Study Never Measured HRV
This is the part that should change how you read the headline. The white paper is about heart rate in beats per minute. It makes no claim about heart rate variability. Apple's Readiness score, the flagship health feature of this generation, is built on Recovery HRV, and Apple's newsroom release says HRV is now sampled up to 24 times more often than before. None of that added sampling density is validated here.
Heart rate accuracy and HRV accuracy are related but not interchangeable. HRV is computed from beat-to-beat interval timing, so it is sensitive to jitter that barely moves an averaged bpm number. A device can land a very good bpm figure and still produce a noisy interval series. Apple has published nothing on the second question.
The same gap applies to sleep. Readiness draws on sleep and overnight vitals, and Apple's protocol contains no overnight arm at all. Every session was 15 to 20 minutes of exercise or roughly 20 minutes of daytime activity. Longer sessions, cold weather, and overnight wear are all untested.
What This Means If You Own an Oura or a WHOOP
Do not throw anything out over this. The study is a fair read on daytime and workout heart rate, and that is a real weakness of rings and screenless straps. A ring on your finger during a heavy set of deadlifts is in close to the worst possible position for optical sensing, which is why Oura tells people not to wear it while lifting and why Apple did not test it there.
But workout heart rate is not the job most Oura and WHOOP owners bought the device for. Those products are built around overnight resting heart rate, overnight HRV, and sleep staging, and this study tested none of the three. The one condition in Apple's protocol closest to how a ring is actually used, the rest condition in Daily Living, is also the single place where a rival beat the Apple Watch: the Pixel Watch 4 was more accurate for both RMSE and MAE at rest, though by less than 0.5 bpm.
If you want to know how these devices do at the thing they are designed for, the useful reading is on sleep and recovery specifically: Oura sleep accuracy, WHOOP recovery accuracy, and which wearable is best for HRV. What would settle the argument is an independent replication under a different sponsor, with an overnight arm and an HRV endpoint.
Where Vora Fits
Vora reads whatever your device writes into Apple Health or Health Connect, so this study is less a buying instruction than a calibration note. If your data comes from a Series 12, the underlying heart rate stream feeding your trends is now better validated than it was. If it comes from a ring or a strap, the daytime and workout numbers deserve a bit more skepticism than the overnight ones, which this study did not examine.
One limitation worth stating plainly, because Apple's paper raises it: background heart rate reaches HealthKit every 30 seconds rather than every five, so a third-party app reading Apple Health is working from a coarser stream than the watch shows on its own face. That is a property of the platform, not of any app on top of it. What each wearable does and does not hand over is broken down in which wearable metrics actually sync to Apple Health, and the integrations page lists the services Vora connects to directly, including Oura, Garmin and WHOOP (in beta).