Hassan Dawood
Product Manager · Digital Health · Digital Phenotyping

I build and operate the systems that turn passive smartphone and wearable data into research-grade behavioral signal — and I run the same rigor on myself.

I'm a product leader working at the intersection of digital health, clinical research operations, and mobile/backend engineering. I specialize in developing and maintaining the iOS, Android, and cloud infrastructure behind large-scale digital phenotyping studies.

As Head of Platform for the Beiwe Research Platform at Harvard T.H. Chan School of Public Health, I lead product management for a mobile research tool used in large-scale behavioral health studies — working directly with clients, researchers, and engineers to ship features, debug across the full stack, and manage the pipeline that turns high-throughput sensor data into meaningful behavioral metrics.

Before this, I worked in healthcare data analytics, building workflows to extract insight from complex financial and clinical datasets. My background spans neuroscience, software, and operations — which is mostly what lets me translate what researchers and clinicians actually need into a real development roadmap.

Outside of work: running, skiing, videography, drones, Legos, wearables — and turning all of the above into personal datasets I actually analyze. Reach out if you're working on something at this intersection.

Health Data Studies

N-of-1 Study
Ramadan Fasting: 9-Year Wearable Analysis

How does Ramadan dawn-to-sunset fasting affect resting heart rate, HRV, sleep timing, and activity, compared to matched pre/post-Ramadan baselines across nine annual cycles?

Study details
Background

Ramadan's dawn-to-sunset fasting shifts eating schedule, sleep timing, and activity substantially, but most physiological evidence on it comes from short clinical trials. This N-of-1 analysis draws on nine consecutive Ramadan cycles (2018-2026) of personal wearable data (Apple Watch, Oura Ring) to characterize within-subject cardiovascular, sleep, and behavioral change relative to matched pre/post-Ramadan baselines.

Methods

Values during each Ramadan fasting window were compared against matched 30-day pre- and post-Ramadan baselines using per-cycle Mann-Whitney U tests, pooled across cycles with a sign test on year-level direction and a linear mixed-effects model (random intercept per year). Resting heart rate and HRV were further decomposed into sleep-window and waking-hours components to rule out schedule-shift artifacts.

Results

Resting heart rate was significantly and robustly lower during Ramadan on both devices (Apple: -7.1 bpm; Oura: -2.6 bpm; p<0.0001), holding under sleep/waking decomposition -- the most consistent finding in the study. Waking-hours HRV increased on Apple (+7.4 ms, p<0.0001), concentrated in the afternoon and coinciding with reduced daytime activity, suggesting an activity-mediated effect rather than a direct autonomic one. Sleep shifted later (bedtime +0.64h, wake +0.61h, both p<0.005), though Apple and Oura disagreed on whether total sleep duration and architecture changed. Structured exercise dropped ~65%, and app-tracked meditation minutes fell despite self-reported increases in spiritual practice -- likely reflecting a shift toward untracked religious activities rather than less reflective activity overall.

Conclusions

The findings are hypothesis-generating, not confirmatory: Ramadan's lunar drift sweeps the fasting window through nearly every season across nine years, so season and fasting are perfectly confounded in this design, and the RHR finding has a plausible non-physiological explanation (Apple's algorithm may partly derive "resting" values from low-activity periods) that was not ruled out. See Limitations for the full list of caveats.

Design
N-of-1, 9 Ramadan cycles (2018–2026) vs. matched 30-day pre/post baselines — per-cycle tests pooled via mixed-effects model
Streams
Resting HR, HRV, sleep timing & architecture, workouts, active energy (Apple Watch 2018–2026, Oura Gen 3 2023–2026), self-reported context
Analysis
Mann-Whitney Ulinear mixed effectssign testsleep/waking decompositioncross-device replication
Open questions
Is the RHR drop confounded by season or by activity-linked measurement construction? Why do Apple and Oura disagree on sleep duration and architecture during Ramadan?
Limitations

1) No season control -- the central unresolved weakness. Ramadan drifted from May-June (2018) to Feb-March (2026), sweeping through nearly every season. A matched same-season control year was attempted and found infeasible with only nine years of data, so it was abandoned rather than forced. Every finding below may be confounded with season -- not isolated to fasting per se.

2) The resting heart rate finding -- the study's most heavily replicated result -- has a plausible non-physiological alternative that was not tested. Apple's RHR algorithm is proprietary; if it partly derives "resting" estimates from low-activity periods, the confirmed drop in daytime activity during Ramadan could mechanically lower the computed value with no true cardiovascular change. Sleep/waking decomposition ruled out a simpler schedule-shift artifact but not this deeper, activity-linked measurement confound.

3) Multiple comparisons were not formally corrected. Several dozen tests were run across metrics, devices, time windows, and sensitivity checks. Year-over-year directional consistency was used as a practical substitute, not a formal correction. Findings that replicate across independent years and devices (resting heart rate) should be weighted far more heavily than single-metric, single-device results (e.g. the Oura-only REM finding, the daytime HRV mechanism).

4) The analysis was exploratory, not confirmatory. Sub-analyses (HRV time-of-day binning, activity clock-hour bins, sleep-gap tolerances) were selected after an initial result suggested they were worth running, then tested on the same data that motivated them -- a garden-of-forking-paths design, not pre-registered or validated on held-out years.

5) Available, directly relevant covariates were not used. Body mass data exists in the same source records but was never incorporated, despite weight and hydration status being the standard first hypothesis for a fasting-related RHR change. This is a concrete, low-cost extension, not a fundamental limitation of the data.

6) Measurement-construct validity across devices is not fully established. Oura's HRV and RHR algorithms aren't documented against any named methodology (e.g. RMSSD vs. SDNN), so cross-device agreement here reflects agreement in direction, not a confirmed shared physiological construct.

7) Self-reported behavioral context is a single retrospective account covering nine years that show clear internal heterogeneity (e.g. one year with increased rather than decreased workout frequency) -- useful for interpretation, but may understate real year-to-year variation.

8) Data completeness varies materially by year and metric, so cross-year consistency claims are implicitly weighted toward years with more complete device wear and sync coverage.

Underlying physiological mechanisms (hydration, caloric intake, autonomic tone) were not directly measured in any of the above.

N-of-1 Study
Apple Watch vs. Oura Ring Comparison

How do summary metrics from the Apple Watch Ultra (Gen 1) and Oura Ring (Gen 3) compare when measured over the same year-long period?

Study details
Background

Consumer wearables are increasingly used as parallel or redundant sources of physiological data, but device-to-device agreement is rarely evaluated under real-world, free-living conditions with two devices worn independently by the same person. This study assesses agreement between an Apple Watch and an Oura Ring across sleep, heart rate, HRV, respiratory rate, SpO2, and step count, using ~2 years of paired daily data from a single subject.

Methods

Apple Health and Oura API records were merged over the overlapping window (Aug 2023-Aug 2026, up to 1,068 days). Two corrections were applied throughout: source isolation, since Oura writes into Apple Health and contaminated several Apple record types (63% of raw sleep records, the largest contributor to step counts); and true concurrent-wear filtering, restricting comparisons to days with 12+ hours of overlap between each device's independently-derived worn intervals. Agreement was quantified with Pearson correlation, Lin's concordance correlation coefficient, Bland-Altman bias and limits of agreement, a regression-to-the-mean check, and year-clustered mixed-effects models to guard against pseudoreplication across ~1,000 autocorrelated daily observations.

Results

Agreement varied substantially by metric and was frequently misestimated by naive comparison. Steps showed the strongest raw agreement (r=0.78-0.85), with Oura reading ~1,000 fewer steps/day. Resting heart rate was initially the weakest metric (r=0.36) but this was a category error, not a device discrepancy: Apple's RHR is drawn from awake stillness while Oura's is sleep-only; restricting Apple to the same window Oura uses raised correlation to r=0.877 (a consistent +5.4 bpm offset, not random disagreement) -- the largest correction identified. HRV showed a comparable circadian effect but a partial, unresolved ~18ms residual gap even within the identical sleep window. Sleep duration agreement was moderate after correction (r=0.64-0.69); Deep sleep was the weakest sleep metric (r=0.34-0.39), indicating genuinely divergent staging algorithms. SpO2 showed weak, uncorrectable agreement (r=0.26-0.31) from a genuine measurement-window mismatch (all-day Apple vs. sleep-only Oura).

Conclusions

Reported device-agreement statistics are highly sensitive to underlying data-pipeline choices: cross-device contamination and non-concurrent wear time can each inflate or deflate apparent agreement as much as genuine device differences. Once controlled for, most apparent Apple-Oura disagreement is explained by differing measurement-window definitions (most clearly for RHR) rather than sensor inaccuracy; HRV's residual gap and Deep sleep-stage disagreement remain genuine and unexplained after exhaustive artifact-checking. See Limitations for the full list of caveats.

Design
Observational, ~1 year (May 2024–May 2025), no control period — daily & weekly comparison
Streams
Resting heart rate, HRV, sleep duration, and other overlapping summary metrics from both devices
Analysis
descriptive statsBland-AltmanPearson / Spearmanpaired t-testcross-correlationmixed effects
Open questions
Do the devices agree on directionality of change? Are discrepancies larger during travel, illness, or poor sleep?
Limitations

1) One person, ~1,000 consecutive nights. Every p-value and confidence interval treats those nights as independent samples, but sleep debt, weekly rhythms, and seasons all autocorrelate day to day, so statistical confidence is likely overstated throughout. Flagged everywhere but not corrected for (would need a block bootstrap).

2) Oura's HRV algorithm is unconfirmed. Commonly assumed to be RMSSD-based for ring trackers, but Oura's own API spec never names it -- treated as unknown, not RMSSD, throughout.

3) 46MB of per-workout detail files are unused. A side effect of re-ingesting the workouts export, these contain per-second HR/energy during each individual workout -- real data, sitting idle.

4) A handful of outlier nights in the sleep comparison (multi-hour disagreements) were never individually root-caused.

5) The Apple wear-time derivation is a judgment-call proxy, not a measured quantity -- a 30-minute HR-sampling-gap threshold, not an official output the way Oura's field is. It correctly found days the Watch was clearly off while the Ring was worn, but likely still over-counts some genuinely-worn time as off-wrist, and is reliable at the day level only.

6) Two ingestion scripts silently overwrite the same output file -- the CSV-based and export.xml-based pipelines both write to the same workouts file with different schemas, and whichever ran last wins. Found while investigating workouts, not yet fixed.

N-of-1 Study
Screen Time Effects (Work in Progress)

How does daily screen time impact HRV, sleep quality, and Oura readiness score?

Study details
Design
All available historical data, high- vs. low-screen-time days; travel days excluded
Streams
Daily screen time (iPhone/Watch/third-party logs), HRV, readiness score, sleep metrics, travel indicators
Analysis
quantile stratificationunpaired t-testlinear regressionmixed effectslagged regression
Open questions
Does screen time closer to bedtime hit harder? Are some app categories more disruptive than others? Does high physical activity buffer the effect?

Personal Projects

Essay · Personal Site Architecture
I Built a Personal Website That Publishes Itself

Most personal sites either go stale after one build or turn every update into a small chore of editing code and redeploying. This one is backed entirely by Notion instead of a CMS dashboard I'd have to learn: adding a project means filling out a database row, and the site reflects what I'm actually working on because updating it is that easy.

Tool · Personal Behavioral Analytics
Instagram Activity Analyzer

I analyze my own personal data as a habit, not because anyone asked, so when Instagram made a full data export available I built a tool to actually look at what was in mine. It's a single self-contained HTML file that reads the export ZIP entirely in the browser: nothing gets uploaded anywhere, and the tool is built to ignore message and comment content outright, extracting only timestamps and anonymized contact IDs.

Essay · Research Funding Automation
I Built a Grant-Finding Tool for My Lab

Part of running operations for the Beiwe Research Platform at Harvard's Onnela Lab is keeping it funded, which used to mean checking Grants.gov, the NIH Guide, and half a dozen foundation sites by hand, hoping I didn't miss a deadline three weeks out. I built a Python scraper that pulls open solicitations from public grant sources and has Claude score each one against the platform's actual research profile.

Essay · Digital Health Research
I Exported My Apple Watch Data Twice. It Didn't Match.

In 2018 I started tracking my own Apple Watch HRV data, just personal curiosity. When I exported the same historical stretch twice, once in 2020 and again in 2021, the numbers didn't match. Nothing about my actual heart rate history had changed, but Apple's algorithm had silently reprocessed it in between.

Essay · Secure Data Engineering
Building a Privacy-First Pipeline for My Own Call and Text History

Beiwe, the research platform I run at work, used to be able to collect call and text metadata directly from a participant's phone, until Apple changed what third-party iOS apps are allowed to see. I got curious whether there was still a way to get at equivalent data some other way, starting purely as a personal project on my own Mac, with every identifier hashed from the very first line of code.

Essay · Personal Productivity
Systems Not Silos — A Productivity Framework

I work at the intersection of academic research, software development, platform operations, technical support, and client account management — which means constant context-switching across very different kinds of work. Managing that manually, keeping it all in my head and improvising responses to recurring situations, was unsustainable. So I built what I think of as a professional operating system: a defined set of inputs, transformation rules, storage locations, and outputs that runs recurring work without me reinventing it every time.

Custom GPT · Product Case Study
JARVIS — Job Analysis & Role Vetting Intel System

Job searches are overwhelming — juggling resumes, tailoring applications, figuring out which roles actually align with a long-term goal. I wanted a tool that could stay organized, give strategic feedback, and support the whole process. So I treated it like launching a real product: built, tested, refined, and iterated on prompts to fix hallucinations — with my sister, who was job-hunting herself, giving real-time feedback on which features actually mattered.

Custom GPT
FRIDAY — Fantasy Research, Data & Analysis for You

Fantasy football is part skill, part luck, and a whole lot of research — juggling analyst sites, injury updates, and endless flex debates. I built a persistent, stateful GPT that acts as a personal GM: it remembers my roster across the season, builds an ideal lineup each week from matchup-based projections, produces sourced player reports (ESPN, FantasyPros, Yahoo, RotoWire, Draft Sharks), evaluates waiver and trade moves with tier-based value, and proactively flags injury news on my own roster with replacement suggestions.