I build and operate the systems that turn passive smartphone and wearable data into research-grade behavioral signal, and I run the same rigor on myself.
I'm a product leader working at the intersection of digital health, clinical research operations, and mobile/backend engineering. I specialize in developing and maintaining the iOS, Android, and cloud infrastructure behind large-scale digital phenotyping studies.
As Head of Platform for the Beiwe Research Platform at Harvard T.H. Chan School of Public Health, I lead product management for a mobile research tool used in large-scale behavioral health studies, working directly with clients, researchers, and engineers to ship features, debug across the full stack, and manage the pipeline that turns high-throughput sensor data into meaningful behavioral metrics.
Before this, I worked in healthcare data analytics, building workflows to extract insight from complex financial and clinical datasets. My background spans neuroscience, software, and operations, which is mostly what lets me translate what researchers and clinicians actually need into a real development roadmap.
Outside of work: running, skiing, videography, drones, Legos, wearables. I turn all of the above into personal datasets I actually analyze. Reach out if you're working on something at this intersection.
Personal Projects
- Resting heart rate looked like the weakest metric at first (r=0.36) but became one of the strongest (r=0.877) once both devices were matched to the same measurement window.
- Steps agreed well (r=0.78-0.85); Oura consistently read about 1,000 fewer steps per day than Apple.
- SpO2 was the one metric that couldn't be corrected (r=0.26-0.31), reflecting a real difference in what each device measures, not a data quality problem.
Study details
Apple Watch and Oura Ring readings are often treated as roughly interchangeable, but I wear both and rarely checked whether they actually agree when compared directly, night for night and metric for metric. This analysis compares sleep, heart rate, HRV, respiratory rate, SpO2, and step count using about 2 years of daily paired data, correcting two overlooked data quality issues before drawing any conclusions.
Two corrections were applied to every metric before comparison. Source isolation removed steps and sleep records that Oura had written directly into Apple Health, which were contaminating Apple's own exported data. True concurrent-wear filtering restricted the comparison to days with at least 12 hours of confirmed overlap between the two devices, since they weren't always worn together. Agreement was measured with Pearson r, Lin's concordance correlation coefficient (which penalizes scale and offset error that r misses), and Bland-Altman bias and limits of agreement. Year-clustered mixed-effects models confirmed the results weren't artifacts of autocorrelated daily observations.
Agreement varied widely and moved in both directions once corrected. Resting heart rate looked like the weakest metric at first (r=0.36) but turned out to be one of the strongest (r=0.877, CCC=0.544) once each device's measurement window was matched: Apple draws resting heart rate from awake stillness throughout the day, while Oura draws it from sleep only. Full-day continuous heart rate streams from both devices, a window-independent check, agreed closely on their own (r=0.848, CCC=0.847), confirming the discrepancy was a window mismatch rather than a sensor problem. Steps agreed well (r=0.78-0.85), with Oura consistently reading about 1,000 fewer steps per day. Sleep duration and stages agreed moderately (r=0.64-0.69 for duration), with Deep sleep the weakest sleep metric (r=0.34-0.39). HRV followed the same window-mismatch pattern as resting heart rate but left a genuine unexplained gap of about 18 ms even after matching sleep windows. SpO2 was the weakest metric overall (r=0.26-0.31) and, unlike the others, could not be corrected: it reflects a real measurement-window difference by design (all-day for Apple, sleep-only for Oura), not a data quality problem.
Most of what looks like disagreement between the two devices is a measurement-window mismatch or a data contamination artifact, not a true sensor discrepancy. Once corrected, most metrics agree closely. The two exceptions, HRV and SpO2, appear to reflect genuine differences in what each device is actually measuring, not noise that further correction would resolve.
Limitations
Single subject (N=1), personal, non-clinical, exploratory. Not a generalizable device-accuracy claim.
Multiple corrections were layered over about 20 rounds of investigation. The numbers above are the final, most-corrected versions, not the first-pass numbers, which were sometimes very different, most notably resting heart rate's r=0.36 to r=0.877 flip.
No raw personal health data appears here, only aggregate correlation and bias statistics.
- Resting heart rate dropped significantly during Ramadan on both devices (Apple: -7.1 bpm; Oura: -2.6 bpm), the most consistent finding across all nine years.
- Sleep shifted later (bedtime +0.64h, wake +0.61h) and structured exercise fell about 65%.
- The RHR drop has a plausible non-physiological explanation (an activity-linked measurement artifact) that was not ruled out, so the finding is hypothesis-generating, not confirmed.
Study details
Ramadan's dawn-to-sunset fasting shifts eating schedule, sleep timing, and activity substantially, but most physiological evidence on it comes from short clinical trials. This N-of-1 analysis draws on nine consecutive Ramadan cycles (2018-2026) of personal wearable data (Apple Watch, Oura Ring) to characterize within-subject cardiovascular, sleep, and behavioral change relative to matched pre/post-Ramadan baselines.
Values during each Ramadan fasting window were compared against matched 30-day pre- and post-Ramadan baselines using per-cycle Mann-Whitney U tests, pooled across cycles with a sign test on year-level direction and a linear mixed-effects model (random intercept per year). Resting heart rate and HRV were further decomposed into sleep-window and waking-hours components to rule out schedule-shift artifacts.Results Resting heart rate was significantly and robustly lower during Ramadan on both devices (Apple: -7.1 bpm; Oura: -2.6 bpm; p<0.0001), holding under sleep/waking decomposition, the most consistent finding in the study. Waking-hours HRV increased on Apple (+7.4 ms, p<0.0001), concentrated in the afternoon and coinciding with reduced daytime activity, suggesting an activity-mediated effect rather than a direct autonomic one. Sleep shifted later (bedtime +0.64h, wake +0.61h, both p<0.005), though Apple and Oura disagreed on whether total sleep duration and architecture changed. Structured exercise dropped ~65%, and app-tracked meditation minutes fell despite self-reported increases in spiritual practice, likely reflecting a shift toward untracked religious activities rather than less reflective activity overall.
The findings are hypothesis-generating, not confirmatory: Ramadan's lunar drift sweeps the fasting window through nearly every season across nine years, so season and fasting are perfectly confounded in this design, and the RHR finding has a plausible non-physiological explanation (Apple's algorithm may partly derive "resting" values from low-activity periods) that was not ruled out. See Limitations for the full list of caveats.
Limitations
1) No season control: the central unresolved weakness. Ramadan drifted from May-June (2018) to Feb-March (2026), sweeping through nearly every season. A matched same-season control year was attempted and found infeasible with only nine years of data, so it was abandoned rather than forced. Every finding below may be confounded with season, not isolated to fasting per se.
2) The resting heart rate finding, the study's most heavily replicated result, has a plausible non-physiological alternative that was not tested. Apple's RHR algorithm is proprietary; if it partly derives "resting" estimates from low-activity periods, the confirmed drop in daytime activity during Ramadan could mechanically lower the computed value with no true cardiovascular change. Sleep/waking decomposition ruled out a simpler schedule-shift artifact but not this deeper, activity-linked measurement confound.
3) Multiple comparisons were not formally corrected. Several dozen tests were run across metrics, devices, time windows, and sensitivity checks. Year-over-year directional consistency was used as a practical substitute, not a formal correction. Findings that replicate across independent years and devices (resting heart rate) should be weighted far more heavily than single-metric, single-device results (e.g. the Oura-only REM finding, the daytime HRV mechanism).4) The analysis was exploratory, not confirmatory. Sub-analyses (HRV time-of-day binning, activity clock-hour bins, sleep-gap tolerances) were selected after an initial result suggested they were worth running, then tested on the same data that motivated them, a garden-of-forking-paths design, not pre-registered or validated on held-out years.
5) Available, directly relevant covariates were not used. Body mass data exists in the same source records but was never incorporated, despite weight and hydration status being the standard first hypothesis for a fasting-related RHR change. This is a concrete, low-cost extension, not a fundamental limitation of the data.
6) Measurement-construct validity across devices is not fully established. Oura's HRV and RHR algorithms aren't documented against any named methodology (e.g. RMSSD vs. SDNN), so cross-device agreement here reflects agreement in direction, not a confirmed shared physiological construct.
7) Self-reported behavioral context is a single retrospective account covering nine years that show clear internal heterogeneity (e.g. one year with increased rather than decreased workout frequency), useful for interpretation, but may understate real year-to-year variation.
8) Data completeness varies materially by year and metric, so cross-year consistency claims are implicitly weighted toward years with more complete device wear and sync coverage.
Underlying physiological mechanisms (hydration, caloric intake, autonomic tone) were not directly measured in any of the above.
- The same 640-day HRV window, exported seven months apart, correlated at just 0.67, even though nothing about the underlying heart rate data had changed.
- Apple's algorithm had silently reprocessed historical data in between the two exports.
- The finding was covered by The Verge, where JP Onnela called Apple's algorithms black boxes that make wearable data unreliable for research without documented transparency into how they work.
Study details
Consumer wearables are increasingly used as parallel or redundant sources of physiological data, but device-to-device agreement is rarely evaluated under real-world, free-living conditions with two devices worn independently by the same person. This study assesses agreement between an Apple Watch and an Oura Ring across sleep, heart rate, HRV, respiratory rate, SpO2, and step count, using ~2 years of paired daily data from a single subject.
Apple Health and Oura API records were merged over the overlapping window (Aug 2023-Aug 2026, up to 1,068 days). Two corrections were applied throughout: source isolation, since Oura writes into Apple Health and contaminated several Apple record types (63% of raw sleep records, the largest contributor to step counts); and true concurrent-wear filtering, restricting comparisons to days with 12+ hours of overlap between each device's independently-derived worn intervals. Agreement was quantified with Pearson correlation, Lin's concordance correlation coefficient, Bland-Altman bias and limits of agreement, a regression-to-the-mean check, and year-clustered mixed-effects models to guard against pseudoreplication across ~1,000 autocorrelated daily observations.
Agreement varied substantially by metric and was frequently misestimated by naive comparison. Steps showed the strongest raw agreement (r=0.78-0.85), with Oura reading ~1,000 fewer steps/day. Resting heart rate was initially the weakest metric (r=0.36) but this was a category error, not a device discrepancy: Apple's RHR is drawn from awake stillness while Oura's is sleep-only; restricting Apple to the same window Oura uses raised correlation to r=0.877 (a consistent +5.4 bpm offset, not random disagreement), the largest correction identified. HRV showed a comparable circadian effect but a partial, unresolved ~18ms residual gap even within the identical sleep window. Sleep duration agreement was moderate after correction (r=0.64-0.69); Deepsleep was the weakest sleep metric (r=0.34-0.39), indicating genuinely divergent staging algorithms. SpO2 showed weak, uncorrectable agreement (r=0.26-0.31) from a genuine measurement-window mismatch (all-day Apple vs. sleep-only Oura). A separate, workout-specific effect was also confirmed: brief device removal during exercise (more often for running than resistance training) undercounts steps in a way day-level wear-time filtering misses, though it does not measurably affect the other metrics.
Reported device-agreement statistics are highly sensitive to underlying data-pipeline choices: cross-device contamination and non-concurrent wear time can each inflate or deflate apparent agreement as much as genuine device differences. Once controlled for, most apparent Apple-Oura disagreement is explained by differing measurement-window definitions (most clearly for RHR) rather than sensor inaccuracy; HRV's residual gap and Deep sleep-stage disagreement remain genuine and unexplained after exhaustive artifact-checking. See Limitations for the full list of caveats.
Limitations
1) One person, ~1,000 consecutive nights. Every p-value and confidence interval treats those nights as independent samples, but sleep debt, weekly rhythms, and seasons all autocorrelate day to day, so statistical confidence is likely overstated throughout. Flagged everywhere but not corrected for (would need a block bootstrap).
2) Oura's HRV algorithm is unconfirmed. Commonly assumed to be RMSSD-based for ring trackers, but Oura's own API spec never names it: treated as unknown, not RMSSD, throughout.
3) 46MB of per-workout detail files are unused. A side effect of re-ingesting the workouts export, these contain per-second HR/energy during each individual workout: real data, sitting idle.
4) A handful of outlier nights in the sleep comparison (multi-hour disagreements) were never individually root-caused.
5) The Apple wear-time derivation is a judgment-call proxy, not a measured quantity: a 30-minute HR-sampling-gap threshold, not an official output the way Oura's field is. It correctly found days the Watch was clearly off while the Ring was worn, but likely still over-counts some genuinely-worn time as off-wrist, and is reliable at the day level only.
6) Two ingestion scripts silently overwrite the same output file: the CSV-based and export.xml-based pipelines both write to the same workouts file with different schemas, and whichever ran last wins. Found while investigating workouts, not yet fixed.
- Daytime coverage was a near-tie (about 90-91% for both devices), not the lopsided Oura advantage longer battery life would predict.
- Nighttime coverage modestly favored Oura (77.7% vs. 73.9%), driven by how often each device needed charging.
- About 21-22% of nights on both devices produced no sleep record at all even though the device was worn and transmitting, an unexplained quirk affecting both devices about equally.
Study details
The Oura Ring is often assumed to capture more continuous data than an Apple Watch, driven by longer battery life (days vs. about one day) and a form factor that's easier to keep on continuously. This analysis tests that assumption directly using ~2 years of data from both devices worn together as a daily habit, comparing daytime and nighttime coverage separately.
Daily data presence was compared for both devices across the full paired window, split into daytime coverage and nighttime (sleep-record) coverage. Gap episodes (runs of missing data) were characterized by frequency and duration for each device, then decomposed into two distinct mechanisms: physical absence (device genuinely not worn) versus algorithmic miss (device worn and transmitting, but no sleep record produced). Two multi-day full outages were cross-checked against calendar events to confirm cause.
Daytime coverage was a near-tie, not the lopsided Oura advantage the battery-life hypothesis predicted: both devices had data on about 90-91% of days, with similarly short, similarly frequent gaps. Nighttime coverage did favor Oura, but modestly (77.7% of nights vs. 73.9% for Apple Watch, n=1,085 dual-coverage nights). A raw look at gap-episode counts initially looked backwards: Oura had more individual missing-night episodes (186) than Apple (133), even though each Oura episode was shorter, until checking whether the Watch was actually worn during those "missing sleep" stretches revealed it usually was: 17 of 20 long gaps were nights the Watch was worn and transmitting fine, it simply didn't log a sleep record. Once physical absence was separated from algorithmic miss, the original battery-life theory held up cleanly: the Watch had about 1.8x more genuine physical overnight gaps than the Ring (50 episodes/68 nights vs. 28 episodes/37 nights), consistent with more frequent and longer charging needs. The earlier "Oura has more gaps" number was mostly counting thealgorithmic-miss category, which affects both devices about equally. Two week-long full outages (one per device) both lined up almost exactly with confirmed travel dates when a charger wasn't available.
The naive "the Ring wins on coverage" framing isn't quite right. Daytime coverage is a wash, nighttime coverage modestly favors the Ring, and the mechanism really is charging frequency and duration, but only once separated from a larger, equally-shared, and still-unexplained sleep-detection quirk affecting roughly a fifth of all nights on both devices. That quirk, not the original wear-time question, is now the largest open question in the dataset.
Limitations
1) Two full-device outages are generalized in any published account to "a trip" with no further identifying detail (specific people, relationships, venues, or addresses). The date-level confirmation against calendar events is the interesting part, not who or where specifically.
2) The physical-vs-algorithmic gap split is a derived classification (based on whether heart-rate data was being transmitted during a "missing sleep" window), not a directly measured or officially labeled distinction from either device.
3) N=1, single subject, observational, not a generalizable device-reliability claim.
4) The ~21-22% algorithmic sleep-detection-miss rate remains unexplained; this analysis confirms it's mostly independent per-device (82% of miss-nights affect only one device, weak correlation) but does not identify a root cause.
- The original theory (more communication lowers stress) did not hold up. An early result that seemed to confirm it turned out to be a data-coverage bug, and disappeared once fixed.
- What actually held up on both devices independently: late-night texting (10pm-5am), not message volume overall, associates with a higher resting heart rate, lower HRV, and less sleep.
- Texting late delays bedtime by about a minute per message, which explains most of the sleep-duration effect, but only partly explains the resting-heart-rate effect on one of the two devices.
Study details
Two earlier studies on this same dataset had each found something during Ramadan: resting heart rate drops, HRV rises, and separately, message and call volume both go up. The theory this piece tests is whether the increased communication, talking to family, feeling connected, is actually driving the lower physiological stress, not just moving alongside it. To test that instead of assuming it, the comparison was run across the full multi-year history, not just inside Ramadan, since Ramadan changes diet, sleep timing, and exercise all at once and could never isolate communication as the cause on its own.
Communication metrics (call count, message count, character count, unique contacts reached) were correlated against resting heart rate, HRV, and sleep duration from both Apple Watch and Oura Ring, independently. Ramadan was kept in as a statistical control rather than the whole test. A second pass split messages into late-night (10pm-5am) and daytime, once volume alone stopped looking meaningful, and controlled for shared long-term drift and bedtime delay to isolate what was actually driving the pattern.
The first pass looked like it confirmed the theory: more calls that day associated with meaningfully lower resting heart rate and higher HRV, on both devices. It was a coverage bug. The call-data source only reliably logs calls from mid-2025 onward, and earlier "zero calls" days were mostly missing data mislabeled as zero, which manufactured the contrast. Restricting to the reliable window erased almost the entire effect, and one correlation flipped sign entirely. On the honest data, none of the communication metrics tested showed more communication leading to lower stress markers. If anything, higher message volume associated with slightly higher resting heart rate, the opposite of the theory. Chasing down why led to the one finding that held up: late-night texting specifically, not volume overall, associates with a higher resting heart rate, lower HRV, and less sleep the same or next night, on both devices independently, even after removing a shared multi-year drift that had been inflating the first version of this result too. Late-night messages predict a measurably later bedtime, about a minute of delay per message, which fully explains the sleep-duration effect. It only partly explains the resting-heart-rate effect: the effect survives controlling for bedtime on one device (Oura) but mostly disappears on the other (Apple Watch). A separate device disagreement, on whether a later bedtime associates with higher or lower resting heart rate, turned out to have a known cause: Apple's daily resting-heart-rate figure excludes the sleep window by design, while Oura's is sleep-only by design, a roughly 10 bpm gap between those two windows for this person. Restricting Apple's data to the same sleep window it was missing made the two devices agree.
The instinct that connecting with people more should lower stress did not hold up when tested. Communication volume is not the story. What actually held up, on two independent devices, is more mundane and more actionable: texting late at night specifically delays bedtime and shows up in worse next-morning readings. If there is a behavior change here, it is stop texting late, not text more.
Limitations
N=1, observational, no causal claim, correlation only, and not adjusted for every possible confound. The piece itself demonstrates two confounds found and fixed along the way, and the process is not exhaustive.
Multiple comparisons were run across many metrics and lags without formal statistical correction. Direction and magnitude consistency across independent devices was the evidence bar used throughout, not any single p-value in isolation.
No message content or contact identity was used anywhere, only timestamps and counts.
Call-data reliability window: the call analysis is only trustworthy from mid-2025 onward; none of the calls-related numbers generalize to the full multi-year history.
- Unique contacts reached per day were significantly higher during Ramadan (+1.14 contacts/day), consistent across all 5 tested years.
- Message character volume was also significantly higher (+778 characters/day), while message count trended higher without reaching conventional significance.
- This runs counter to the general sense that Ramadan feels quieter and more focused, but that impression is about daily life broadly, not communication behavior specifically, so the two aren't in direct conflict.
Study details
Ramadan is culturally a highly social, family- and community-oriented period, but its effect on measured communication behavior, as opposed to self-reported impressions of daily life, hasn't been checked directly. This N-of-1 analysis uses a local call/message export to test whether Ramadan changes how much I reach out to and hear from other people, split out as its own project from a broader physiological Ramadan study since it never touched a physiological metric directly.
Daily call and text counts, unique contacts reached, and message character volume were compared between the Ramadan window and matched baseline periods, per cycle, then pooled across years with a sign test on year-level direction and a linear mixed-effects model. Messages have continuous coverage from 2022-2026 (5 Ramadan cycles); calls are usable only from 2025 onward (the source database is a rolling cache), so only the 2026 cycle falls inside the coverage window and call findings are reported descriptively, without a statistical test. Scope was deliberately restricted to aggregate daily counts: no message content or contact identifiers were read or used at any point.Results Unique contacts reached per day were significantly higher during Ramadan (+1.14 contacts/day, mixed-effects p=0.0009, unanimous across all 5 tested years). Message character volume was also significantly higher (+778 characters/day, p=0.022, 4 of 5 years positive), while message count trended higher (+9/day, 4/5 years positive) without reaching conventional significance (p=0.106). The single available call-only year (2026) showed roughly flat call counts but higher total minutes, though this is descriptive only and not independently verifiable further back.
This runs counter to the general self-reported sense that Ramadan feels quieter and more focused, but that impression was about daily life broadly, not communication behavior specifically, so the two aren't in direct conflict. The most plausible explanation is that Ramadan's social rituals (iftar coordination, checking in on family, Ramadan/Eid greetings) increase how many people I reach even on days that otherwise feel calmer. This mirrors a parallel finding in the companion physiological study, where app-tracked meditation minutes fell during Ramadan despite self-reported increases in spiritual engagement. In both cases, a measured proxy points a different direction than the general self-reported impression, because it's measuring something narrower than the impression describes.
Limitations
1) No season control, for the same underlying reason as the companion physiological project: Ramadan's lunar drift moves the fasting window through every season over the study period, so every finding here is potentially confounded with season, not isolated to Ramadan or fasting specifically.
2) Calls are evaluable for exactly one year (2026). No claim about call behavior should be generalized beyond that single cycle.
3) Breadth vs. depth is not distinguished by this data. A higher count of unique contacts reached could mean many brief, low-effort greetings to a wide circle, or genuinely deeper engagement with a similar-sized circle. This analysis can't tell those apart, and per-contact data that could (available in de-identified, hashed form) was deliberately not used.
4) N=1, observational, no causal claim supported.
5) No message content or contact identity is available in this analysis's source data: nothing to redact, but also nothing to add for texture or specificity if asked for examples.