Hassan Dawood
Product Manager · Digital Health · Digital Phenotyping

I build and operate the systems that turn passive smartphone and wearable data into research-grade behavioral signal, and I run the same rigor on myself.

I'm a product leader working at the intersection of digital health, clinical research operations, and mobile/backend engineering. I specialize in developing and maintaining the iOS, Android, and cloud infrastructure behind large-scale digital phenotyping studies.

As Head of Platform for the Beiwe Research Platform at Harvard T.H. Chan School of Public Health, I lead product management for a mobile research tool used in large-scale behavioral health studies, working directly with clients, researchers, and engineers to ship features, debug across the full stack, and manage the pipeline that turns high-throughput sensor data into meaningful behavioral metrics.

Before this, I worked in healthcare data analytics, building workflows to extract insight from complex financial and clinical datasets. My background spans neuroscience, software, and operations, which is mostly what lets me translate what researchers and clinicians actually need into a real development roadmap.

Outside of work: running, skiing, videography, drones, Legos, wearables. I turn all of the above into personal datasets I actually analyze. Reach out if you're working on something at this intersection.

Personal Projects

Health Data Analytics
Long-form
Ramadan Fasting: 9-Year Wearable Analysis

How does Ramadan dawn-to-sunset fasting affect resting heart rate, HRV, sleep timing, and activity, compared to matched pre/post-Ramadan baselines across nine annual cycles?

  • Resting heart rate dropped significantly during Ramadan on both devices (Apple: -7.1 bpm; Oura: -2.6 bpm), the most consistent finding across all nine years.
  • Sleep shifted later (bedtime +0.64h, wake +0.61h) and structured exercise fell about 65%.
Study details
Key Findings
  • Resting heart rate dropped significantly during Ramadan on both devices (Apple: -7.1 bpm; Oura: -2.6 bpm), the most consistent finding across all nine years.
  • Sleep shifted later (bedtime +0.64h, wake +0.61h) and structured exercise fell about 65%.
  • The RHR drop has a plausible non-physiological explanation (an activity-linked measurement artifact) that was not ruled out, so the finding is hypothesis-generating, not confirmed.
Background

Ramadan's dawn-to-sunset fasting shifts eating schedule, sleep timing, and activity substantially, but most physiological evidence on it comes from short clinical trials. This N-of-1 analysis draws on nine consecutive Ramadan cycles (2018-2026) of personal wearable data (Apple Watch, Oura Ring) to characterize within-subject cardiovascular, sleep, and behavioral change relative to matched pre/post-Ramadan baselines.

Methods

Values during each Ramadan fasting window were compared against matched 30-day pre- and post-Ramadan baselines using per-cycle Mann-Whitney U tests, pooled across cycles with a sign test on year-level direction and a linear mixed-effects model (random intercept per year). Resting heart rate and HRV were further decomposed into sleep-window and waking-hours components to rule out schedule-shift artifacts.Results Resting heart rate was significantly and robustly lower during Ramadan on both devices (Apple: -7.1 bpm; Oura: -2.6 bpm; p<0.0001), holding under sleep/waking decomposition, the most consistent finding in the study. Waking-hours HRV increased on Apple (+7.4 ms, p<0.0001), concentrated in the afternoon and coinciding with reduced daytime activity, suggesting an activity-mediated effect rather than a direct autonomic one. Sleep shifted later (bedtime +0.64h, wake +0.61h, both p<0.005), though Apple and Oura disagreed on whether total sleep duration and architecture changed. Structured exercise dropped ~65%, and app-tracked meditation minutes fell despite self-reported increases in spiritual practice, likely reflecting a shift toward untracked religious activities rather than less reflective activity overall.

Conclusions

The findings are hypothesis-generating, not confirmatory: Ramadan's lunar drift sweeps the fasting window through nearly every season across nine years, so season and fasting are perfectly confounded in this design, and the RHR finding has a plausible non-physiological explanation (Apple's algorithm may partly derive "resting" values from low-activity periods) that was not ruled out. See Limitations for the full list of caveats.

Design
N-of-1, 9 Ramadan cycles (2018-2026) vs. matched 30-day pre/post baselines, per-cycle tests pooled via mixed-effects model
Streams
Resting HR, HRV, sleep timing & architecture, workouts, active energy (Apple Watch 2018-2026, Oura Gen 3 2023-2026), self-reported context
Analysis
Mann-Whitney Ulinear mixed effectssign testsleep/waking decompositioncross-device replication
Open questions
Is the RHR drop confounded by season or by activity-linked measurement construction? Why do Apple and Oura disagree on sleep duration and architecture during Ramadan?
Limitations

1) No season control: the central unresolved weakness. Ramadan drifted from May-June (2018) to Feb-March (2026), sweeping through nearly every season. A matched same-season control year was attempted and found infeasible with only nine years of data, so it was abandoned rather than forced. Every finding below may be confounded with season, not isolated to fasting per se.

2) The resting heart rate finding, the study's most heavily replicated result, has a plausible non-physiological alternative that was not tested. Apple's RHR algorithm is proprietary; if it partly derives "resting" estimates from low-activity periods, the confirmed drop in daytime activity during Ramadan could mechanically lower the computed value with no true cardiovascular change. Sleep/waking decomposition ruled out a simpler schedule-shift artifact but not this deeper, activity-linked measurement confound.

3) Multiple comparisons were not formally corrected. Several dozen tests were run across metrics, devices, time windows, and sensitivity checks. Year-over-year directional consistency was used as a practical substitute, not a formal correction. Findings that replicate across independent years and devices (resting heart rate) should be weighted far more heavily than single-metric, single-device results (e.g. the Oura-only REM finding, the daytime HRV mechanism).4) The analysis was exploratory, not confirmatory. Sub-analyses (HRV time-of-day binning, activity clock-hour bins, sleep-gap tolerances) were selected after an initial result suggested they were worth running, then tested on the same data that motivated them, a garden-of-forking-paths design, not pre-registered or validated on held-out years.

5) Available, directly relevant covariates were not used. Body mass data exists in the same source records but was never incorporated, despite weight and hydration status being the standard first hypothesis for a fasting-related RHR change. This is a concrete, low-cost extension, not a fundamental limitation of the data.

6) Measurement-construct validity across devices is not fully established. Oura's HRV and RHR algorithms aren't documented against any named methodology (e.g. RMSSD vs. SDNN), so cross-device agreement here reflects agreement in direction, not a confirmed shared physiological construct.

7) Self-reported behavioral context is a single retrospective account covering nine years that show clear internal heterogeneity (e.g. one year with increased rather than decreased workout frequency), useful for interpretation, but may understate real year-to-year variation.

8) Data completeness varies materially by year and metric, so cross-year consistency claims are implicitly weighted toward years with more complete device wear and sync coverage.

Underlying physiological mechanisms (hydration, caloric intake, autonomic tone) were not directly measured in any of the above.

Health Data Analytics
Long-form
I Exported My Apple Watch Data Twice. It Didn't Match.

In 2018 I started tracking my own Apple Watch HRV data, just personal curiosity. When I exported the same historical stretch twice, once in 2020 and again in 2021, the numbers didn't match. Nothing about my actual heart rate history had changed, but Apple's algorithm had silently reprocessed it in between.

  • The same 640-day HRV window, exported seven months apart, correlated at just 0.67, even though nothing about the underlying heart rate data had changed.
  • Apple's algorithm had silently reprocessed historical data in between the two exports.
Study details
Key Findings
  • The same 640-day HRV window, exported seven months apart, correlated at just 0.67, even though nothing about the underlying heart rate data had changed.
  • Apple's algorithm had silently reprocessed historical data in between the two exports.
  • The finding was covered by The Verge and led JP Onnela's research team to drop the Apple Watch from a planned study.
Background

Consumer wearables are increasingly used as parallel or redundant sources of physiological data, but device-to-device agreement is rarely evaluated under real-world, free-living conditions with two devices worn independently by the same person. This study assesses agreement between an Apple Watch and an Oura Ring across sleep, heart rate, HRV, respiratory rate, SpO2, and step count, using ~2 years of paired daily data from a single subject.

Methods

Apple Health and Oura API records were merged over the overlapping window (Aug 2023-Aug 2026, up to 1,068 days). Two corrections were applied throughout: source isolation, since Oura writes into Apple Health and contaminated several Apple record types (63% of raw sleep records, the largest contributor to step counts); and true concurrent-wear filtering, restricting comparisons to days with 12+ hours of overlap between each device's independently-derived worn intervals. Agreement was quantified with Pearson correlation, Lin's concordance correlation coefficient, Bland-Altman bias and limits of agreement, a regression-to-the-mean check, and year-clustered mixed-effects models to guard against pseudoreplication across ~1,000 autocorrelated daily observations.

Results

Agreement varied substantially by metric and was frequently misestimated by naive comparison. Steps showed the strongest raw agreement (r=0.78-0.85), with Oura reading ~1,000 fewer steps/day. Resting heart rate was initially the weakest metric (r=0.36) but this was a category error, not a device discrepancy: Apple's RHR is drawn from awake stillness while Oura's is sleep-only; restricting Apple to the same window Oura uses raised correlation to r=0.877 (a consistent +5.4 bpm offset, not random disagreement), the largest correction identified. HRV showed a comparable circadian effect but a partial, unresolved ~18ms residual gap even within the identical sleep window. Sleep duration agreement was moderate after correction (r=0.64-0.69); Deepsleep was the weakest sleep metric (r=0.34-0.39), indicating genuinely divergent staging algorithms. SpO2 showed weak, uncorrectable agreement (r=0.26-0.31) from a genuine measurement-window mismatch (all-day Apple vs. sleep-only Oura). A separate, workout-specific effect was also confirmed: brief device removal during exercise (more often for running than resistance training) undercounts steps in a way day-level wear-time filtering misses, though it does not measurably affect the other metrics.

Conclusions

Reported device-agreement statistics are highly sensitive to underlying data-pipeline choices: cross-device contamination and non-concurrent wear time can each inflate or deflate apparent agreement as much as genuine device differences. Once controlled for, most apparent Apple-Oura disagreement is explained by differing measurement-window definitions (most clearly for RHR) rather than sensor inaccuracy; HRV's residual gap and Deep sleep-stage disagreement remain genuine and unexplained after exhaustive artifact-checking. See Limitations for the full list of caveats.

Design
Observational, single-subject, ~2 years of paired daily data (Aug 2023-Aug 2026, up to 1,068 days), source-isolated and concurrent-wear-filtered comparison, no control period
Streams
Resting heart rate, HRV, sleep duration & stages, respiratory rate, SpO2, and step count from both devices
Analysis
Pearson rLin's CCCBland-Altmanregression-to-the-mean checkyear-clustered mixed-effects models
Open questions
Why does the ~18ms HRV gap persist even after matching sleep windows? What drives the divergence in Deep sleep staging between devices? Does the SpO2 measurement-window mismatch explain all of its weak agreement, or is there a residual sensor difference too?
Limitations

1) One person, ~1,000 consecutive nights. Every p-value and confidence interval treats those nights as independent samples, but sleep debt, weekly rhythms, and seasons all autocorrelate day to day, so statistical confidence is likely overstated throughout. Flagged everywhere but not corrected for (would need a block bootstrap).

2) Oura's HRV algorithm is unconfirmed. Commonly assumed to be RMSSD-based for ring trackers, but Oura's own API spec never names it: treated as unknown, not RMSSD, throughout.

3) 46MB of per-workout detail files are unused. A side effect of re-ingesting the workouts export, these contain per-second HR/energy during each individual workout: real data, sitting idle.

4) A handful of outlier nights in the sleep comparison (multi-hour disagreements) were never individually root-caused.

5) The Apple wear-time derivation is a judgment-call proxy, not a measured quantity: a 30-minute HR-sampling-gap threshold, not an official output the way Oura's field is. It correctly found days the Watch was clearly off while the Ring was worn, but likely still over-counts some genuinely-worn time as off-wrist, and is reliable at the day level only.

6) Two ingestion scripts silently overwrite the same output file: the CSV-based and export.xml-based pipelines both write to the same workouts file with different schemas, and whichever ran last wins. Found while investigating workouts, not yet fixed.

Health Data Analytics
Which Wearable Actually Captures More Data?

Does the Oura Ring actually capture more data than the Apple Watch, as its longer battery life and form factor would predict?

  • Daytime coverage was a near-tie (about 90-91% for both devices), not the lopsided Oura advantage longer battery life would predict.
  • Nighttime coverage modestly favored Oura (77.7% vs. 73.9%), driven by how often each device needed charging.
Study details
Key Findings
  • Daytime coverage was a near-tie (about 90-91% for both devices), not the lopsided Oura advantage longer battery life would predict.
  • Nighttime coverage modestly favored Oura (77.7% vs. 73.9%), driven by how often each device needed charging.
  • About 21-22% of nights on both devices produced no sleep record at all even though the device was worn and transmitting, an unexplained quirk affecting both devices about equally.
Background

The Oura Ring is often assumed to capture more continuous data than an Apple Watch, driven by longer battery life (days vs. about one day) and a form factor that's easier to keep on continuously. This analysis tests that assumption directly using ~2 years of data from both devices worn together as a daily habit, comparing daytime and nighttime coverage separately.

Method

Daily data presence was compared for both devices across the full paired window, split into daytime coverage and nighttime (sleep-record) coverage. Gap episodes (runs of missing data) were characterized by frequency and duration for each device, then decomposed into two distinct mechanisms: physical absence (device genuinely not worn) versus algorithmic miss (device worn and transmitting, but no sleep record produced). Two multi-day full outages were cross-checked against calendar events to confirm cause.

Results

Daytime coverage was a near-tie, not the lopsided Oura advantage the battery-life hypothesis predicted: both devices had data on about 90-91% of days, with similarly short, similarly frequent gaps. Nighttime coverage did favor Oura, but modestly (77.7% of nights vs. 73.9% for Apple Watch, n=1,085 dual-coverage nights). A raw look at gap-episode counts initially looked backwards: Oura had more individual missing-night episodes (186) than Apple (133), even though each Oura episode was shorter, until checking whether the Watch was actually worn during those "missing sleep" stretches revealed it usually was: 17 of 20 long gaps were nights the Watch was worn and transmitting fine, it simply didn't log a sleep record. Once physical absence was separated from algorithmic miss, the original battery-life theory held up cleanly: the Watch had about 1.8x more genuine physical overnight gaps than the Ring (50 episodes/68 nights vs. 28 episodes/37 nights), consistent with more frequent and longer charging needs. The earlier "Oura has more gaps" number was mostly counting thealgorithmic-miss category, which affects both devices about equally. Two week-long full outages (one per device) both lined up almost exactly with confirmed travel dates when a charger wasn't available.

Conclusions

The naive "the Ring wins on coverage" framing isn't quite right. Daytime coverage is a wash, nighttime coverage modestly favors the Ring, and the mechanism really is charging frequency and duration, but only once separated from a larger, equally-shared, and still-unexplained sleep-detection quirk affecting roughly a fifth of all nights on both devices. That quirk, not the original wear-time question, is now the largest open question in the dataset.

Design
Observational, ~2 years of paired daily data, both devices worn simultaneously as a habit, daytime and nighttime coverage compared separately
Streams
Daily data presence/absence and sleep-record presence for both devices, cross-checked against workout logs and calendar events for outage attribution
Analysis
coverage-rate comparisongap-episode frequency/duration analysisphysical-vs-algorithmic gap decomposition
Open questions
What explains the ~21-22% of nights, on both devices, where the device was worn and transmitting heart-rate data but produced no sleep record at all? Is it a shared hard-to-classify-night phenomenon or independent per-device failures?
Limitations

1) Two full-device outages are generalized in any published account to "a trip" with no further identifying detail (specific people, relationships, venues, or addresses). The date-level confirmation against calendar events is the interesting part, not who or where specifically.

2) The physical-vs-algorithmic gap split is a derived classification (based on whether heart-rate data was being transmitted during a "missing sleep" window), not a directly measured or officially labeled distinction from either device.

3) N=1, single subject, observational, not a generalizable device-reliability claim.

4) The ~21-22% algorithmic sleep-detection-miss rate remains unexplained; this analysis confirms it's mostly independent per-device (82% of miss-nights affect only one device, weak correlation) but does not identify a root cause.

Tools for Research
Long-form
I Built a Grant-Finding Tool for My Lab

Part of running operations for the Beiwe Research Platform at Harvard's Onnela Lab is keeping it funded, which used to mean checking Grants.gov, the NIH Guide, and half a dozen foundation sites by hand, hoping I didn't miss a deadline three weeks out. I built a Python scraper that pulls open solicitations from public grant sources and has Claude score each one against the platform's actual research profile.

Tools for Research
Long-form
Building a Privacy-First Pipeline for My Own Call and Text History

Beiwe, the research platform I run at work, used to be able to collect call and text metadata directly from a participant's phone, until Apple changed what third-party iOS apps are allowed to see. I got curious whether there was still a way to get at equivalent data some other way, starting purely as a personal project on my own Mac, with every identifier hashed from the very first line of code.

Tools for Research
Long-form
JARVIS: Job Analysis & Role Vetting Intel System

Job searches are overwhelming: juggling resumes, tailoring applications, figuring out which roles actually align with a long-term goal. I wanted a tool that stayed organized and gave strategic feedback across the whole process. So I treated it like launching a real product: built, tested, refined, and iterated on prompts to fix hallucinations, with my sister, who was job-hunting herself, giving real-time feedback on which features actually mattered.

Health Data Analytics
Ramadan and Social Reach: A Communication Log Analysis

Does Ramadan change how much I reach out to and hear from other people (call/text volume, unique contacts, message depth)?

Study details
Background

Ramadan is culturally a highly social, family- and community-oriented period, but its effect on measured communication behavior, as opposed to self-reported impressions of daily life, hasn't been checked directly. This N-of-1 analysis uses a local call/message export to test whether Ramadan changes how much I reach out to and hear from other people, split out as its own project from a broader physiological Ramadan study since it never touched a physiological metric directly.

Methods

Daily call and text counts, unique contacts reached, and message character volume were compared between the Ramadan window and matched baseline periods, per cycle, then pooled across years with a sign test on year-level direction and a linear mixed-effects model. Messages have continuous coverage from 2022-2026 (5 Ramadan cycles); calls are usable only from 2025 onward (the source database is a rolling cache), so only the 2026 cycle falls inside the coverage window and call findings are reported descriptively, without a statistical test. Scope was deliberately restricted to aggregate daily counts: no message content or contact identifiers were read or used at any point.Results Unique contacts reached per day were significantly higher during Ramadan (+1.14 contacts/day, mixed-effects p=0.0009, unanimous across all 5 tested years). Message character volume was also significantly higher (+778 characters/day, p=0.022, 4 of 5 years positive), while message count trended higher (+9/day, 4/5 years positive) without reaching conventional significance (p=0.106). The single available call-only year (2026) showed roughly flat call counts but higher total minutes, though this is descriptive only and not independently verifiable further back.

Conclusions

This runs counter to the general self-reported sense that Ramadan feels quieter and more focused, but that impression was about daily life broadly, not communication behavior specifically, so the two aren't in direct conflict. The most plausible explanation is that Ramadan's social rituals (iftar coordination, checking in on family, Ramadan/Eid greetings) increase how many people I reach even on days that otherwise feel calmer. This mirrors a parallel finding in the companion physiological study, where app-tracked meditation minutes fell during Ramadan despite self-reported increases in spiritual engagement. In both cases, a measured proxy points a different direction than the general self-reported impression, because it's measuring something narrower than the impression describes.

Design
N-of-1, per-cycle Ramadan-window vs. baseline comparison pooled across years via mixed-effects model and sign test; messages cover 5 Ramadan cycles (2022-2026), calls only 1 (2026, descriptive only)
Streams
Daily call/text counts, unique contacts reached, message character volume, call minutes: aggregate counts only, no message content or contact identity
Analysis
mixed-effects modelsign test
Open questions
Does the increased reach reflect many brief greetings to a wide circle, or deeper engagement with the same circle? Would call behavior replicate once more years of call data exist?
Limitations

1) No season control, for the same underlying reason as the companion physiological project: Ramadan's lunar drift moves the fasting window through every season over the study period, so every finding here is potentially confounded with season, not isolated to Ramadan or fasting specifically.

2) Calls are evaluable for exactly one year (2026). No claim about call behavior should be generalized beyond that single cycle.

3) Breadth vs. depth is not distinguished by this data. A higher count of unique contacts reached could mean many brief, low-effort greetings to a wide circle, or genuinely deeper engagement with a similar-sized circle. This analysis can't tell those apart, and per-contact data that could (available in de-identified, hashed form) was deliberately not used.

4) N=1, observational, no causal claim supported.

5) No message content or contact identity is available in this analysis's source data: nothing to redact, but also nothing to add for texture or specificity if asked for examples.

Thought Piece
Long-form
Systems Not Silos: A Productivity Framework

I work at the intersection of academic research, software development, platform operations, technical support, and client account management, which means constant context-switching across very different kinds of work. Managing that manually, keeping it all in my head and improvising responses to recurring situations, was unsustainable. So I built what I think of as a professional operating system: a defined set of inputs, transformation rules, storage locations, and outputs that runs recurring work without me reinventing it every time.

Tools for Research
Long-form
I Built a Personal Website That Publishes Itself

Most personal sites either go stale after one build or turn every update into a small chore of editing code and redeploying. This one is backed entirely by Notion instead of a CMS dashboard I'd have to learn: adding a project means filling out a database row, and the site reflects what I'm actually working on because updating it is that easy.

Tools for Research
Instagram Activity Analyzer

I analyze my own personal data as a habit, not because anyone asked, so when Instagram made a full data export available I built a tool to actually look at what was in mine. It's a single self-contained HTML file that reads the export ZIP entirely in the browser: nothing gets uploaded anywhere, and the tool is built to ignore message and comment content outright, extracting only timestamps and anonymized contact IDs.

Tools for Research
FRIDAY: Fantasy Research, Data & Analysis for You

Fantasy football is part skill, part luck, and a whole lot of research: juggling analyst sites, injury updates, and endless flex debates. I built a persistent, stateful GPT that acts as a personal GM: it remembers my roster across the season, builds an ideal lineup each week from matchup-based projections, produces sourced player reports (ESPN, FantasyPros, Yahoo, RotoWire, Draft Sharks), evaluates waiver and trade moves with tier-based value, and proactively flags injury news on my own roster with replacement suggestions.

Health Data Analytics
Screen Time Effects (Work in Progress)

How does daily screen time impact HRV, sleep quality, and Oura readiness score?

Study details
Design
All available historical data, high- vs. low-screen-time days; travel days excluded
Streams
Daily screen time (iPhone/Watch/third-party logs), HRV, readiness score, sleep metrics, travel indicators
Analysis
quantile stratificationunpaired t-testlinear regressionmixed effectslagged regression
Open questions
Does screen time closer to bedtime hit harder? Are some app categories more disruptive than others? Does high physical activity buffer the effect?