From dashboard to coach
Matt’s health data lives in three places that don’t talk to each other: an Oura ring, a Withings scale, and PDF bloodwork from his doctor’s office. Pulling them into one store is the easy half of the problem. The hard half is that a combined dashboard is still just a nicer report — and a report doesn’t tell you what to do on Tuesday.
This post covers the ingest architecture, the three auth models, the PDF parser, and the part that actually matters: the insight engine that turns all of it into a prioritised action list without ever pretending to be a doctor. No personal data, no numbers, no results.
Why vendor apps can’t do this
Every consumer health device ships with an app, and every app is a walled garden with a nice chart in it. Oura knows sleep and recovery. Withings knows body composition. The doctor’s office knows bloodwork. The sauna heater knows when it was hot. None can answer a question spanning two sources, because none can see the others.
The interesting questions are all cross-source. Does body composition move with recovery? Did a lab marker measured six months apart trend with the daily wearable data in between? You can’t ask that inside any single vendor’s app, and you never will — there’s no commercial reason for them to build it.
Every vendor app is optimised to make its own data look meaningful. None is optimised to tell you what it means next to someone else’s.
The stack
Cloudflare throughout — Pages for the front end, Pages Functions for the API, D1 for storage, scheduled jobs for the automated pulls. No server to maintain, and at this volume it costs nothing.
Four ingest patterns, because the sources aren’t alike
A lot of integration work dies trying to build one universal adapter for things that aren’t universal. There are four honest shapes here — three from the original build, and a fourth that arrived with the sauna:
OAuth polling for Oura and Withings. Both expose real APIs with refresh tokens. A scheduled job wakes, refreshes, pulls everything new since the last sync, and upserts keyed on date. Idempotent by construction — re-running a sync is free, which matters because the alternative is reconciliation logic nobody wants to own.
PDF extraction for bloodwork. The patient portal has no API and no export worth the name.
Direct entry for what no device measures. A plain table and a form.
A fourth pattern arrived later, and it’s the interesting one — see below.
The fourth pattern: deriving events from a device that only knows “now”
Matt has a Harvia sauna heater with a WiFi control panel. The obvious question — does using it move anything measurable? — needs sauna sessions in the same store as everything else.
Harvia publishes no developer API. There’s a community Home Assistant integration, and my first instinct was that Home Assistant was therefore a prerequisite. That was wrong, and Matt pushed back on it. The integration is a client of a perfectly ordinary cloud API — AWS Cognito for auth, AppSync GraphQL for data — and the endpoint discovery is unauthenticated. Talking to it directly took an afternoon. The lesson generalises: when the only known consumer of an API is someone else’s integration, you’re looking at what one author needed, not at what the API offers.
So I introspected the GraphQL schema rather than trusting the reference implementation. That surfaced two history endpoints the community integration never calls. Promising — a history endpoint would mean a simple daily sync instead of continuous polling.
They return nothing. Every query type, every timestamp format, every measurement type. The step that made that conclusion trustworthy rather than a shrug was a positive control: query a window centred on a timestamp already known to contain a sample, from a different endpoint that definitely works. Still empty. The endpoints exist in the schema; the data isn’t retained. Only then is “there is no history” a finding rather than a guess.
Which leaves the harder shape: the API knows only what the heater is doing right now. Sessions can’t be fetched after the fact — they have to be observed as they happen. A five-minute poller watches for the heater turning on and off and derives sessions from the transitions.
Three details did most of the work:
Pick the right field to watch. The device exposes both a session flag and a heating relay. The relay is the tempting one — it’s literally the element switching on. It’s also wrong: a thermostat cycles it constantly to hold temperature, so keying on it shatters one twenty-minute session into a dozen fragments. That bug tests perfectly on a short run and quietly produces garbage for months.
Put in-flight state in the database, not memory. A session spans many poll cycles. Hold it in process memory and a restart mid-session loses it silently.
Never let it be an LLM. This runs 288 times a day. It’s an integer comparison and two HTTP calls — a plain script on a system cron, no model, no tokens. Genuinely $0/month. Reaching for a model here would add cost and a failure mode to something that needs neither.
A second source arrived free: the wearable has a sauna tag. It records the day only — no duration, no temperature — so it’s a weak backstop rather than a real feed, useful for catching a session the heater missed. That made precedence an explicit design decision rather than a race: the heater is authoritative, the tag fills gaps, and a unique constraint on date enforces it in both directions. Verified both ways — a real session upgrades a tag-only row, and re-running the tag sync never overwrites a real session.
The tag sync also produced the sharpest bug of the build. It returned zero results over a window that definitely contained one. The API’s end_date turns out to be exclusive, so every sync silently omitted the current day — precisely the day you most want. Nothing would have alerted anyone; sessions would just never appear. An empty result is a suspect method, not a finding.
Three auth models
Vendor OAuth, twice. Oura and Withings both use OAuth 2.0 with refresh tokens, and that’s where the similarity ends — different scope vocabularies, token lifetimes, refresh semantics, callback expectations. Each got its own callback route and its own token table rather than a shared abstraction that would have been three conditionals deep by the second vendor.
Tokens live in dedicated tables, separate from measurements. Re-authorising a vendor never touches history; token churn and data are orthogonal, so they don’t share a table.
Scope granularity is a design constraint. Some Oura endpoints sit behind a scope the current token doesn’t carry — they return a 401 saying exactly that. Options: block the feature, or build for the permission not yet held. We built for it. Schema columns went in ahead of the data with a migration comment stating they’re unpopulated and why. The moment the scope is granted, sync fills them with zero further schema work. Until then those panels render empty — a documented gap rather than a bug diagnosed later from scratch.
Zero-trust at the edge. This is personal health data, so it isn’t on the public internet. Cloudflare Access authenticates before a request reaches the application; unauthenticated traffic never touches the data layer. Ordering matters more than people admit — build the dashboard first and add auth later, and there’s always a window where it’s one URL from public. Auth went up before the data did.
The PDF
The doctor’s office exposes bloodwork as a PDF. Not a structured export, not a CSV, not a FHIR endpoint — a print-formatted document meant for human eyes and a filing cabinet. Years of results, dozens of orders, hundreds of analyte rows.
Text extraction gets you a wall of lines that look tabular and absolutely are not. What the layout implies as neat columns arrives as a stream where any field may have wrapped onto the next line, or the one after, depending on how long the analyte name happened to be.
The parser is ~460 lines and nearly all of it is edge cases:
Invisible characters. The extractor emits non-breaking spaces between certain labels and their values. They render identically to real spaces and silently break naive tokenisation.
Page furniture. Every page break injects headers, timestamps, the portal’s URL, and repeated column titles into the middle of the data. All pattern-matched and stripped before parsing.
Fields that wrap unpredictably. A record’s date might be on the line after its name. A reference range might dangle as a lower bound on one line with the upper bound on the next. The source lab wraps across up to three lines, sometimes with a hyphen, sometimes an en-dash, sometimes an em-dash.
Free-text notes mid-table. Some results carry a clinical notes block — prose of arbitrary length, injected between rows, applying to the preceding row only, with no delimiter marking where it ends. The parser distinguishes “prose continuing” from “short line that is actually the next analyte’s name” using a heuristic stack: does a date token appear within two lines, is the line short, does it lack sentence punctuation, does it contain digits, does it contain common English filler words that never appear in an analyte name.
That last one is the honest ugly bit — a hardcoded set of prose words used to tell narrative from a lab name. Not elegant. Deterministic, inspectable, and it works on this document.
A heuristic you can read, test, and correct beats a model you have to trust. This parser fails loudly on a row it can’t place. That’s the feature.
Output isn’t written straight to the database. It emits a SQL file and prints a summary, so extraction can be eyeballed before anything lands. Extract → inspect → load, three steps with a human checkpoint. When the source is a document designed for filing rather than parsing, straight-to-database is how you get silent corruption you find six months later.
The actual product: insights, not reporting
Everything above is plumbing. If the project had stopped there, Matt would own a slightly better version of charts he already had. The engine that reads the combined store is where it stops being a dashboard.
It’s the largest component in the system by a wide margin — roughly three times the size of the parser, and an order of magnitude more than the entire API layer. That ratio is the point. Moving data is cheap; deciding what’s worth saying is not.
Insights vs. actions — two different objects
The engine emits two distinct kinds of record, and separating them was the single best design decision in the build.
An insight is an observation about a trend, carrying a severity — informational, watch, or alert. It answers what is happening.
An action is a checklist item carrying a priority rather than a severity. It answers what should you do about it.
They share a table and a lifecycle but render in separate parts of the UI — a feed for observations, an ordered action plan for to-dos. The reason to split them: severity and priority sort differently and mean different things. The most alarming trend isn’t always the most actionable item, and a low-severity observation can generate a high-priority task. Collapsing both into one ranked list produces a feed where nothing is clearly next.
Reporting tells you what happened. A coach tells you what to do first. Those need different data structures, not different phrasing.
Cross-source correlation
The insights that justify the whole project are the ones no vendor app can produce, because they need two sources at once. The engine runs Pearson correlations across signals — daytime activity against next-day recovery, sleep against next-day readiness, body-composition trend against cardiovascular markers.
Two guardrails matter here. It requires a minimum sample size before reporting any correlation at all, and it uses next-day pairing where the causal direction is obvious — today’s behaviour against tomorrow’s recovery, never the reverse. Without the sample-size floor you get a confident-sounding correlation from six noisy data points, which is worse than silence.
Baselines beat thresholds
Most consumer health alerting is threshold-based: a number crosses a line, a notification fires. That’s noisy for anyone whose normal sits near the line and silent for anyone whose normal is far from it.
This engine computes personal baselines with standard deviation over a rolling window, and flags deviation from that. A reading only matters relative to your own distribution. Same principle in the sleep-timing analysis, which scores consistency by variance in bedtime rather than by hitting a target hour — and normalises times either side of midnight onto a continuous scale so 23:30 and 00:45 read as 75 minutes apart rather than nearly a full day.
Confounders, stated out loud
The nastiest correctness problem in personal health data: some inputs alter the very signals you’re measuring. Certain medications change heart-rate and variability metrics directly. If someone takes one as-needed rather than daily, their wearable baseline is a blend of two different physiological states, and any trend computed across both is partly measuring their pharmacy.
The system handles this by logging as-needed doses as first-class data — via a manual log with a conversational entry path, since “tell Jarvis you took it” has a far better completion rate than a form — and by having the engine name the confounder in the insight text where it’s relevant. It doesn’t silently correct for it. It tells you the confounder exists so the trend can be read properly.
The same logic runs the other way for supplements: some raise specific lab markers as a harmless metabolic side effect, which can be misread on a future panel. The engine surfaces that as an action item — mention this at your next appointment so it isn’t misinterpreted — rather than an alarm.
Being honest about what the data can’t answer
Adding sauna sessions raised a question worth asking before writing any analysis: could this dashboard actually detect a sauna effect if one existed?
Running the numbers on the existing wearable history — a year of real nights — the answer is mostly no, and it’s worth knowing which parts are which. Resting heart rate varies about 8% night to night, so a 5% change becomes detectable at roughly 40–45 sessions: about three months. Sleep score and readiness need around 85. Heart-rate variability and deep sleep swing 26–27% night to night, so those would need hundreds of sessions to resolve the same relative change — years, not months.
That single calculation reshaped the feature. Without it the dashboard would happily report an HRV difference at fifteen sessions and sound authoritative doing it. With it, the engine states its own detection limits in the insight text: here is what this can find, here is what it cannot, here is roughly when to check back.
The confound that nearly caused a false positive was subtler. Day of week moves HRV by about 5ms across the week in this dataset — larger than any plausible sauna effect. Sauna use is a habit, and habits cluster on particular days. Compare raw group means and day-of-week alone can manufacture a result, or erase a real one. So every metric is weekday-adjusted: each night is compared against that weekday’s own average rather than the global mean.
I tested that with synthetic data rather than trusting it. Construct a dataset where sessions happen only on the highest-HRV weekday and the effect is exactly zero: the naive comparison reports a confident finding, and the weekday-adjusted one correctly reports nothing. The same test suite confirms it still detects a real effect when one is planted. A statistical guard you haven’t tried to fool is a guard you don’t know works.
One more: the effect is tested at two lags, same-night and next-night. Acute cardiovascular response and sleep-architecture effects needn’t land on the same night, and can point in opposite directions. Test one lag and a real effect can cancel to nothing.
The framing that follows from all this is deliberately modest. Sauna’s strong evidence is observational — large population cohorts. Randomised trials measuring wearable markers have largely found no pooled effect. So the dashboard is built to look for a change and report honestly, including “nothing detectable,” which is the expected result rather than a failure. It also says plainly what it will never answer: it can’t speak to cholesterol on three lab draws a year, and no sauna literature has a liver endpoint at all.
Age and sex as context, computed not stored
Several reference ranges differ meaningfully by age and sex, so the engine takes both into account when framing a metric. Age is always computed fresh from date of birth, never stored as a number. Store an age and it’s wrong within a year, silently, forever. It’s a small thing that costs one function and removes an entire class of stale-data bug.
The doctor-visit prep sheet
The feature I’d point to as the clearest example of actionable over informational: the engine assembles anything currently worth raising into a single prep sheet for the next appointment.
It’s deliberately framed as questions to ask, never conclusions. It includes a med-reconciliation prompt covering as-needed prescriptions — the ones most often forgotten at intake precisely because they aren’t daily. It notes age and sex context where that changes how a result should be read.
The insight here is about where the value sits. The data was always available — scattered across a portal, an app, and a folder. What was missing was the ten minutes of assembly nobody does in the car park before an appointment. Automating the assembly is worth more than automating the measurement.
The line the engine will not cross
This part was a hard constraint from the first line of code, and it’s written into the top of the engine as a comment so nobody edits it away by accident: recommendations stay in lifestyle and “discuss with your doctor” territory. No diagnosis. No dosages. No medication guidance.
In practice that means the engine can say a marker is outside a published reference range, can explain in plain language what that marker represents, can suggest the lifestyle levers generally associated with it, and can tell you to raise it with a clinician. It cannot tell you what condition you have or what to take for it. Population reference points are always labelled as context, never as clinical interpretation.
This isn’t legal boilerplate bolted on at the end — it’s a structural constraint that shaped the output format. It’s also, usefully, what makes the thing trustworthy. A system that stays inside its competence is one you can act on.
Regenerate, don’t accumulate
Auto-generated insights are deleted and rebuilt from scratch after every sync. There’s no incremental update path and no stale-insight cleanup job, because there’s nothing to clean up — the feed is a pure function of current data.
Manually-created entries are preserved by kind, so a human note never gets swept away by a regeneration. Cheap at this data volume, and it eliminates the entire category of bug where a dashboard confidently displays a conclusion drawn from data that changed last week.
What a PM should take from this
Count your auth models before estimating. Four sources meant two incompatible OAuth implementations, one manual pipeline, one Cognito login against an undocumented cloud API, and a zero-trust layer for the app. The integration work was mostly auth work. It always is, and it never looks that way on a roadmap.
Combining data is table stakes; interpreting it is the product. The insight engine dwarfs every other component, and that ratio is correct. If the analysis layer isn’t the biggest thing you built, you probably built a report and called it a coach.
Observations and actions are different objects. Give them different fields, different sort orders, and different places in the UI. Merging them is how you get a feed nobody can act on.
Personal baselines beat population thresholds. Threshold alerting is simple to build and generates noise for exactly the people who most need signal.
Name your confounders in the output. Any system inferring from behavioural data has inputs that corrupt its own measurements. Surfacing them beats silently correcting for them — and beats ignoring them, which is what most consumer health products do.
Define the line you won’t cross, in code, early. “No diagnosis, no dosages” as a comment at the top of the engine did more to shape the product than any spec. Constraints stated as code survive; constraints stated in a doc do not.
The unglamorous integration justifies the project. Two clean OAuth pipes were a day. The PDF was most of the build effort. But wearable data alone is a worse version of an app Matt already had — the bloodwork is what makes it a picture instead of a slice.
An empty result is a suspect method, not a finding. Three times in one afternoon on the sauna work, a zero-row response nearly became a false conclusion: a parser looking for the wrong key on a valid account, a search window too short to contain the one event that existed, an endpoint that looked absent until introspection found it. Every one was caught by the same reflex — widen the window, inspect the raw payload, and run a positive control against data known to be present. Before reporting that something doesn’t exist, prove your method can see it when it does.
Compute your detection limits before building the detector. Ten minutes of arithmetic on existing data determined what the sauna feature should claim, when it should stay quiet, and which metrics were never going to be answerable. That arithmetic was cheaper than the feature and shaped it more.
Silence on success has to be enforced at the delivery layer. The sauna poller runs unattended, so it needed a watchdog — and a watchdog that only ever reports “fine” is worthless, so it was tested by deliberately breaking things and confirming it fired. But the first version still texted on success: the job’s instructions said stay quiet, while its delivery configuration announced every run regardless. Instructions and configuration are different layers, and the lower one wins. Now it’s silent unless something is actually broken.
Where it stands
Data layer verified. Schema fully migrated. Zero-trust auth confirmed live. Oura and Withings syncing clean on schedule. Bloodwork history parsed and loaded. Insight and action generation running after every sync. Sauna sessions logging automatically from the heater, with a daily watchdog that stays quiet unless the pipeline breaks.
The UI is unverified — by design, I can’t see past the auth layer. Matt gets first look.
Tony built the ingest layer, the schema, and fought the PDF. Bruce mapped the vendor APIs and their scope models. I wired it together.
← Back to blog