I build and operate the systems that turn passive smartphone and wearable data into research-grade behavioral signal, and I run the same rigor on myself.
I'm a product leader working at the intersection of digital health, clinical research operations, and mobile/backend engineering. I specialize in developing and maintaining the iOS, Android, and cloud infrastructure behind large-scale digital phenotyping studies.
As Head of Platform for the Beiwe Research Platform at Harvard T.H. Chan School of Public Health, I lead product management for a mobile research tool used in large-scale behavioral health studies, working directly with clients, researchers, and engineers to ship features, debug across the full stack, and manage the pipeline that turns high-throughput sensor data into meaningful behavioral metrics.
Before this, I worked in healthcare data analytics, building workflows to extract insight from complex financial and clinical datasets. My background spans neuroscience, software, and operations, which is mostly what lets me translate what researchers and clinicians actually need into a real development roadmap.
Outside of work: running, skiing, videography, drones, Legos, wearables. I turn all of the above into personal datasets I actually analyze. Reach out if you're working on something at this intersection.
Personal Projects
- Resting heart rate dropped significantly during Ramadan on both devices (Apple: -7.1 bpm; Oura: -2.6 bpm), the most consistent finding across all nine years.
- Sleep shifted later (bedtime +0.64h, wake +0.61h) and structured exercise fell about 65%.
- The RHR drop has a plausible non-physiological explanation (an activity-linked measurement artifact) that was not ruled out, so the finding is hypothesis-generating, not confirmed.
Study details
Ramadan's dawn-to-sunset fasting shifts eating schedule, sleep timing, and activity substantially, but most physiological evidence on it comes from short clinical trials. This N-of-1 analysis draws on nine consecutive Ramadan cycles (2018-2026) of personal wearable data (Apple Watch, Oura Ring) to characterize within-subject cardiovascular, sleep, and behavioral change relative to matched pre/post-Ramadan baselines.
Values during each Ramadan fasting window were compared against matched 30-day pre- and post-Ramadan baselines using per-cycle Mann-Whitney U tests, pooled across cycles with a sign test on year-level direction and a linear mixed-effects model (random intercept per year). Resting heart rate and HRV were further decomposed into sleep-window and waking-hours components to rule out schedule-shift artifacts.Results Resting heart rate was significantly and robustly lower during Ramadan on both devices (Apple: -7.1 bpm; Oura: -2.6 bpm; p<0.0001), holding under sleep/waking decomposition, the most consistent finding in the study. Waking-hours HRV increased on Apple (+7.4 ms, p<0.0001), concentrated in the afternoon and coinciding with reduced daytime activity, suggesting an activity-mediated effect rather than a direct autonomic one. Sleep shifted later (bedtime +0.64h, wake +0.61h, both p<0.005), though Apple and Oura disagreed on whether total sleep duration and architecture changed. Structured exercise dropped ~65%, and app-tracked meditation minutes fell despite self-reported increases in spiritual practice, likely reflecting a shift toward untracked religious activities rather than less reflective activity overall.
The findings are hypothesis-generating, not confirmatory: Ramadan's lunar drift sweeps the fasting window through nearly every season across nine years, so season and fasting are perfectly confounded in this design, and the RHR finding has a plausible non-physiological explanation (Apple's algorithm may partly derive "resting" values from low-activity periods) that was not ruled out. See Limitations for the full list of caveats.
Limitations
1) No season control: the central unresolved weakness. Ramadan drifted from May-June (2018) to Feb-March (2026), sweeping through nearly every season. A matched same-season control year was attempted and found infeasible with only nine years of data, so it was abandoned rather than forced. Every finding below may be confounded with season, not isolated to fasting per se.
2) The resting heart rate finding, the study's most heavily replicated result, has a plausible non-physiological alternative that was not tested. Apple's RHR algorithm is proprietary; if it partly derives "resting" estimates from low-activity periods, the confirmed drop in daytime activity during Ramadan could mechanically lower the computed value with no true cardiovascular change. Sleep/waking decomposition ruled out a simpler schedule-shift artifact but not this deeper, activity-linked measurement confound.
3) Multiple comparisons were not formally corrected. Several dozen tests were run across metrics, devices, time windows, and sensitivity checks. Year-over-year directional consistency was used as a practical substitute, not a formal correction. Findings that replicate across independent years and devices (resting heart rate) should be weighted far more heavily than single-metric, single-device results (e.g. the Oura-only REM finding, the daytime HRV mechanism).4) The analysis was exploratory, not confirmatory. Sub-analyses (HRV time-of-day binning, activity clock-hour bins, sleep-gap tolerances) were selected after an initial result suggested they were worth running, then tested on the same data that motivated them, a garden-of-forking-paths design, not pre-registered or validated on held-out years.
5) Available, directly relevant covariates were not used. Body mass data exists in the same source records but was never incorporated, despite weight and hydration status being the standard first hypothesis for a fasting-related RHR change. This is a concrete, low-cost extension, not a fundamental limitation of the data.
6) Measurement-construct validity across devices is not fully established. Oura's HRV and RHR algorithms aren't documented against any named methodology (e.g. RMSSD vs. SDNN), so cross-device agreement here reflects agreement in direction, not a confirmed shared physiological construct.
7) Self-reported behavioral context is a single retrospective account covering nine years that show clear internal heterogeneity (e.g. one year with increased rather than decreased workout frequency), useful for interpretation, but may understate real year-to-year variation.
8) Data completeness varies materially by year and metric, so cross-year consistency claims are implicitly weighted toward years with more complete device wear and sync coverage.
Underlying physiological mechanisms (hydration, caloric intake, autonomic tone) were not directly measured in any of the above.
- The same 640-day HRV window, exported seven months apart, correlated at just 0.67, even though nothing about the underlying heart rate data had changed.
- Apple's algorithm had silently reprocessed historical data in between the two exports.
- The finding was covered by The Verge and led JP Onnela's research team to drop the Apple Watch from a planned study.
Study details
Consumer wearables are increasingly used as parallel or redundant sources of physiological data, but device-to-device agreement is rarely evaluated under real-world, free-living conditions with two devices worn independently by the same person. This study assesses agreement between an Apple Watch and an Oura Ring across sleep, heart rate, HRV, respiratory rate, SpO2, and step count, using ~2 years of paired daily data from a single subject.
Apple Health and Oura API records were merged over the overlapping window (Aug 2023-Aug 2026, up to 1,068 days). Two corrections were applied throughout: source isolation, since Oura writes into Apple Health and contaminated several Apple record types (63% of raw sleep records, the largest contributor to step counts); and true concurrent-wear filtering, restricting comparisons to days with 12+ hours of overlap between each device's independently-derived worn intervals. Agreement was quantified with Pearson correlation, Lin's concordance correlation coefficient, Bland-Altman bias and limits of agreement, a regression-to-the-mean check, and year-clustered mixed-effects models to guard against pseudoreplication across ~1,000 autocorrelated daily observations.
Agreement varied substantially by metric and was frequently misestimated by naive comparison. Steps showed the strongest raw agreement (r=0.78-0.85), with Oura reading ~1,000 fewer steps/day. Resting heart rate was initially the weakest metric (r=0.36) but this was a category error, not a device discrepancy: Apple's RHR is drawn from awake stillness while Oura's is sleep-only; restricting Apple to the same window Oura uses raised correlation to r=0.877 (a consistent +5.4 bpm offset, not random disagreement), the largest correction identified. HRV showed a comparable circadian effect but a partial, unresolved ~18ms residual gap even within the identical sleep window. Sleep duration agreement was moderate after correction (r=0.64-0.69); Deepsleep was the weakest sleep metric (r=0.34-0.39), indicating genuinely divergent staging algorithms. SpO2 showed weak, uncorrectable agreement (r=0.26-0.31) from a genuine measurement-window mismatch (all-day Apple vs. sleep-only Oura). A separate, workout-specific effect was also confirmed: brief device removal during exercise (more often for running than resistance training) undercounts steps in a way day-level wear-time filtering misses, though it does not measurably affect the other metrics.
Reported device-agreement statistics are highly sensitive to underlying data-pipeline choices: cross-device contamination and non-concurrent wear time can each inflate or deflate apparent agreement as much as genuine device differences. Once controlled for, most apparent Apple-Oura disagreement is explained by differing measurement-window definitions (most clearly for RHR) rather than sensor inaccuracy; HRV's residual gap and Deep sleep-stage disagreement remain genuine and unexplained after exhaustive artifact-checking. See Limitations for the full list of caveats.
Limitations
1) One person, ~1,000 consecutive nights. Every p-value and confidence interval treats those nights as independent samples, but sleep debt, weekly rhythms, and seasons all autocorrelate day to day, so statistical confidence is likely overstated throughout. Flagged everywhere but not corrected for (would need a block bootstrap).
2) Oura's HRV algorithm is unconfirmed. Commonly assumed to be RMSSD-based for ring trackers, but Oura's own API spec never names it: treated as unknown, not RMSSD, throughout.
3) 46MB of per-workout detail files are unused. A side effect of re-ingesting the workouts export, these contain per-second HR/energy during each individual workout: real data, sitting idle.
4) A handful of outlier nights in the sleep comparison (multi-hour disagreements) were never individually root-caused.
5) The Apple wear-time derivation is a judgment-call proxy, not a measured quantity: a 30-minute HR-sampling-gap threshold, not an official output the way Oura's field is. It correctly found days the Watch was clearly off while the Ring was worn, but likely still over-counts some genuinely-worn time as off-wrist, and is reliable at the day level only.
6) Two ingestion scripts silently overwrite the same output file: the CSV-based and export.xml-based pipelines both write to the same workouts file with different schemas, and whichever ran last wins. Found while investigating workouts, not yet fixed.
- Daytime coverage was a near-tie (about 90-91% for both devices), not the lopsided Oura advantage longer battery life would predict.
- Nighttime coverage modestly favored Oura (77.7% vs. 73.9%), driven by how often each device needed charging.
- About 21-22% of nights on both devices produced no sleep record at all even though the device was worn and transmitting, an unexplained quirk affecting both devices about equally.
Study details
The Oura Ring is often assumed to capture more continuous data than an Apple Watch, driven by longer battery life (days vs. about one day) and a form factor that's easier to keep on continuously. This analysis tests that assumption directly using ~2 years of data from both devices worn together as a daily habit, comparing daytime and nighttime coverage separately.
Daily data presence was compared for both devices across the full paired window, split into daytime coverage and nighttime (sleep-record) coverage. Gap episodes (runs of missing data) were characterized by frequency and duration for each device, then decomposed into two distinct mechanisms: physical absence (device genuinely not worn) versus algorithmic miss (device worn and transmitting, but no sleep record produced). Two multi-day full outages were cross-checked against calendar events to confirm cause.
Daytime coverage was a near-tie, not the lopsided Oura advantage the battery-life hypothesis predicted: both devices had data on about 90-91% of days, with similarly short, similarly frequent gaps. Nighttime coverage did favor Oura, but modestly (77.7% of nights vs. 73.9% for Apple Watch, n=1,085 dual-coverage nights). A raw look at gap-episode counts initially looked backwards: Oura had more individual missing-night episodes (186) than Apple (133), even though each Oura episode was shorter, until checking whether the Watch was actually worn during those "missing sleep" stretches revealed it usually was: 17 of 20 long gaps were nights the Watch was worn and transmitting fine, it simply didn't log a sleep record. Once physical absence was separated from algorithmic miss, the original battery-life theory held up cleanly: the Watch had about 1.8x more genuine physical overnight gaps than the Ring (50 episodes/68 nights vs. 28 episodes/37 nights), consistent with more frequent and longer charging needs. The earlier "Oura has more gaps" number was mostly counting thealgorithmic-miss category, which affects both devices about equally. Two week-long full outages (one per device) both lined up almost exactly with confirmed travel dates when a charger wasn't available.
The naive "the Ring wins on coverage" framing isn't quite right. Daytime coverage is a wash, nighttime coverage modestly favors the Ring, and the mechanism really is charging frequency and duration, but only once separated from a larger, equally-shared, and still-unexplained sleep-detection quirk affecting roughly a fifth of all nights on both devices. That quirk, not the original wear-time question, is now the largest open question in the dataset.
Limitations
1) Two full-device outages are generalized in any published account to "a trip" with no further identifying detail (specific people, relationships, venues, or addresses). The date-level confirmation against calendar events is the interesting part, not who or where specifically.
2) The physical-vs-algorithmic gap split is a derived classification (based on whether heart-rate data was being transmitted during a "missing sleep" window), not a directly measured or officially labeled distinction from either device.
3) N=1, single subject, observational, not a generalizable device-reliability claim.
4) The ~21-22% algorithmic sleep-detection-miss rate remains unexplained; this analysis confirms it's mostly independent per-device (82% of miss-nights affect only one device, weak correlation) but does not identify a root cause.
Study details
Ramadan is culturally a highly social, family- and community-oriented period, but its effect on measured communication behavior, as opposed to self-reported impressions of daily life, hasn't been checked directly. This N-of-1 analysis uses a local call/message export to test whether Ramadan changes how much I reach out to and hear from other people, split out as its own project from a broader physiological Ramadan study since it never touched a physiological metric directly.
Daily call and text counts, unique contacts reached, and message character volume were compared between the Ramadan window and matched baseline periods, per cycle, then pooled across years with a sign test on year-level direction and a linear mixed-effects model. Messages have continuous coverage from 2022-2026 (5 Ramadan cycles); calls are usable only from 2025 onward (the source database is a rolling cache), so only the 2026 cycle falls inside the coverage window and call findings are reported descriptively, without a statistical test. Scope was deliberately restricted to aggregate daily counts: no message content or contact identifiers were read or used at any point.Results Unique contacts reached per day were significantly higher during Ramadan (+1.14 contacts/day, mixed-effects p=0.0009, unanimous across all 5 tested years). Message character volume was also significantly higher (+778 characters/day, p=0.022, 4 of 5 years positive), while message count trended higher (+9/day, 4/5 years positive) without reaching conventional significance (p=0.106). The single available call-only year (2026) showed roughly flat call counts but higher total minutes, though this is descriptive only and not independently verifiable further back.
This runs counter to the general self-reported sense that Ramadan feels quieter and more focused, but that impression was about daily life broadly, not communication behavior specifically, so the two aren't in direct conflict. The most plausible explanation is that Ramadan's social rituals (iftar coordination, checking in on family, Ramadan/Eid greetings) increase how many people I reach even on days that otherwise feel calmer. This mirrors a parallel finding in the companion physiological study, where app-tracked meditation minutes fell during Ramadan despite self-reported increases in spiritual engagement. In both cases, a measured proxy points a different direction than the general self-reported impression, because it's measuring something narrower than the impression describes.
Limitations
1) No season control, for the same underlying reason as the companion physiological project: Ramadan's lunar drift moves the fasting window through every season over the study period, so every finding here is potentially confounded with season, not isolated to Ramadan or fasting specifically.
2) Calls are evaluable for exactly one year (2026). No claim about call behavior should be generalized beyond that single cycle.
3) Breadth vs. depth is not distinguished by this data. A higher count of unique contacts reached could mean many brief, low-effort greetings to a wide circle, or genuinely deeper engagement with a similar-sized circle. This analysis can't tell those apart, and per-contact data that could (available in de-identified, hashed form) was deliberately not used.
4) N=1, observational, no causal claim supported.
5) No message content or contact identity is available in this analysis's source data: nothing to redact, but also nothing to add for texture or specificity if asked for examples.