Measure What Matters: Biomarkers Without Obsession
A number deserves attention only when it measures the intended thing well and changes a worthwhile decision.
The dashboard contains forty-seven numbers.
Sleep score. Readiness. Resting heart rate. Heart-rate variability. Respiratory rate. Glucose variability. Biological age. Recovery age. Inflammation index. Metabolic score. Stress score. Fitness age.
Three values are green. Two are yellow. One has turned red.
The owner of the dashboard is healthy enough to work, travel, train, and care for his family. He is also now preoccupied with a number generated by an algorithm he cannot inspect, measuring a construct no clinician has diagnosed, compared with a population the company does not describe.
Data has produced information.
It has not produced clarity.
Why accomplished people love measurement
Measurement feels like control.
In business, a useful metric can expose a failing process before revenue declines. A dashboard converts complexity into a visible system. The same instinct naturally migrates into health.
But health data differs from many business metrics.
The body varies. Tests contain error. Reference ranges serve different purposes. An unusual result may be temporary, clinically important, meaningless, or the first step in a long series of investigations. A wearable may measure a signal accurately while the score built from it remains opaque.
More measurement can improve decisions.
It can also create false precision, incidental findings, anxiety, unnecessary cost, and interventions aimed at making a number look better without showing that the person becomes healthier.
The question is not whether data is good.
It is whether this data deserves to change what you do.
First, define what kind of measure it is
The word biomarker covers several roles.
A measure might be used to:
- Diagnose or help identify a condition
- Estimate prognosis
- Predict response to a treatment
- Monitor disease or treatment
- Detect a safety problem
- Show biological response to an intervention
- Act as a surrogate for a clinical outcome
- Stratify risk
These roles are not interchangeable.
A marker useful for monitoring a known disease may be poor for screening healthy people. A value associated with risk may not be a useful treatment target. A change showing that a drug reached a pathway may not prove that patients feel or function better.
Before interpreting the number, ask what job it was designed to perform.
Four gates before a metric earns authority
1. Analytical validity
Does the test or device measure the signal accurately and reliably?
Questions include:
- Precision
- Calibration
- Sample handling
- Laboratory method
- Device performance
- Repeatability
- Interference
- Conditions of measurement
A beautiful interpretation cannot rescue a poor measurement.
2. Clinical validity
Does the measure reliably relate to the condition, outcome, or state being claimed?
A sensor can accurately measure a physical signal while the company’s “stress age” built from it lacks strong validation.
Analytical accuracy and clinical meaning are separate gates.
3. Clinical utility
Does using the result improve a worthwhile decision or outcome?
A test may identify more abnormalities without helping people. It may lead to additional imaging, procedures, medication, expense, or anxiety. Utility asks whether acting on the information creates more benefit than harm.
4. Personal relevance
Does this measure answer a question that matters for this person now?
A valid test can still be unnecessary. The person may be at low risk, already receiving appropriate management, or unable to act differently based on the result.
The premium decision is sometimes not to measure.
A reference range is not an optimization target
Laboratory reports often place a value inside or outside a reference interval.
A reference interval commonly reflects the distribution of results in a defined reference population under specific methods. It does not automatically define the boundary between health and disease, nor the ideal target for every individual.
Several concepts are often confused:
Reference interval A statistical description of results in a reference population.
Clinical decision limit A threshold used to guide evaluation or treatment in a defined context.
Treatment target A goal selected according to evidence, condition, and individual risk.
Commercial “optimal range” A term that may or may not have validated clinical meaning.
A value can be within a reference interval and still matter in context. A value can be slightly outside and reflect normal variation, method differences, timing, or a finding that needs confirmation rather than immediate treatment.
“Optimal” should not be accepted merely because it appears between narrower green lines.
One draw is one moment
Laboratory values and physiological signals vary with:
- Time of day
- Fasting or feeding
- Hydration
- Exercise
- Illness
- Sleep
- Stress
- Medication
- Alcohol
- Menstrual status where relevant
- Sample handling
- Laboratory method
- Device placement
- Random variation
This is why trends can be useful—but trends can also mislead if conditions change from one measurement to the next.
A person who tests after a hard workout, poor sleep, travel, and dehydration may not be observing a stable baseline.
Repeatability requires enough consistency that the comparison means something.
More tests create more abnormal results
Suppose twenty independent tests each use a reference interval expected to contain 95 percent of results from a reference population.
Even in an idealized healthy person, the chance that at least one result falls outside its interval rises as more tests are ordered.
Real panels are not independent, and clinical interpretation is more complex. The principle remains: broad testing increases the probability of finding something.
That finding may be important.
It may also lead to:
- Repeat testing
- Imaging
- Specialist visits
- Biopsy or procedure
- Medication
- Anxiety
- Insurance consequences
- Attention diverted from a more meaningful risk
Screening is not harmless because the needle was small.
The benefits and downstream costs belong in the same analysis.
Biomarkers are not outcomes by default
A biomarker may sit somewhere along a pathway between an intervention and an outcome.
Changing it can mean:
- The intervention reached a target
- A biological process shifted
- Risk may have changed
- Nothing meaningful for the person changed
- An unintended tradeoff occurred
FDA distinguishes biomarkers and surrogate endpoints from direct outcomes involving how people feel, function, or survive. Some surrogate endpoints are validated for specific uses. Others are only reasonably likely to predict benefit, and some exploratory markers are far earlier.
Validation is not transferable by analogy.
If one marker predicts outcomes in one disease and treatment context, that does not make every commercially similar measure a longevity endpoint.
The question is:
Has changing this marker been shown to improve an outcome that matters in this context?
Biological age is a model, not a birthday
“Biological age” products compress multiple measures into a single number intended to represent something about aging, risk, or physiological state.
The appeal is obvious. One number seems capable of answering the largest question.
But biological-age models differ in:
- Inputs
- Population
- Statistical method
- Outcome used to train the model
- Technical reproducibility
- Sensitivity to short-term conditions
- Interpretation
- Evidence that changing the score changes health
Two clocks can assign different ages to the same person because they are measuring different patterns.
A lower score after an intervention may reflect a true biological change, technical variation, regression toward the mean, a short-term response, or the behavior of the model itself.
The number can be useful in research.
It should not be allowed to become a consumer verdict without understanding what it predicts and how well.
Wearables are excellent at some jobs
Consumer devices can be valuable for:
- Counting or estimating activity
- Revealing sleep opportunity and schedule
- Observing resting trends
- Supporting habit formation
- Noticing changes during travel or illness
- Providing feedback that motivates movement
- Recording symptoms or events
Their limitations include:
- Proprietary algorithms
- Updates that change scores
- Variable performance across devices and populations
- Error during movement or poor contact
- Incomplete validation for clinical decisions
- Overinterpretation of sleep stages
- Confusion between wellness and medical functions
A wearable can prompt a useful question.
It should not answer a diagnostic question it was not validated to answer.
Write the action plan before ordering the test
Before measuring, complete these sentences:
1. The question is: What exactly are we trying to learn?
2. The test is appropriate because: What evidence supports this use?
3. If the result is high, we will: Name the next step.
4. If the result is low, we will: Name the next step.
5. If the result is ambiguous, we will: Decide whether to repeat, confirm, or stop.
6. The result may vary because: Identify timing, behavior, method, and biological variability.
7. The potential downstream harm is: Include anxiety, procedures, expense, and overdiagnosis.
8. The qualified interpreter is: Who places the number in context?
9. The review date is: When does the result change a decision?
10. The meaningful outcome is: How will the person’s health, function, risk management, or care improve?
If no action changes, the test may be curiosity rather than decision support.
Curiosity is allowed. It should be priced and interpreted honestly.
Build a hierarchy of measurement
A practical measurement system can be organized into layers.
Layer one: measures tied to established care
Examples may include blood pressure, relevant laboratory risk factors, age- and risk-appropriate screening, and monitoring of known conditions under qualified care.
Layer two: functional measures
Strength, aerobic capacity, movement, sleep opportunity, activity, symptoms, and the ability to perform valued tasks.
Layer three: behavior and adherence
Did the plan occur? Was training completed? Was sleep protected? Did alcohol use change? Did the person attend the appointment?
Layer four: emerging or exploratory measures
Novel clocks, composite scores, advanced panels, and experimental biomarkers.
The fourth layer may be interesting. It should not displace the first three.
Data should return you to life
A health dashboard is successful when it improves a decision and then becomes quiet.
It helps a person notice a change, evaluate risk, reinforce a behavior, or review an intervention. It does not require daily emotional allegiance.
The strongest metric may be one that never appears in an app:
- You can walk the city again.
- You no longer fall asleep in meetings.
- Your blood pressure is appropriately managed.
- Training has been consistent for a year.
- You stopped using alcohol as the only way to end the day.
- Your physician has the information needed to care for you.
- You feel sufficiently well to forget about optimization for an afternoon.
Measure what matters.
Then make sure the measurement leaves enough attention for the thing it was meant to protect.
Sources & references
- FDA-NIH BEST Resource: Biomarkers, EndpointS, and other Tools glossary
- FDA-NIH BEST Resource: Evidentiary criteria and validation
- U.S. Food and Drug Administration: Digital health technologies for remote data acquisition
- U.S. Food and Drug Administration: Direct-to-consumer tests
- U.S. Food and Drug Administration: Surrogate endpoint resources
Get the weekly digest
One email a week. Unsubscribe anytime.
More in Sustainable Optimization
Recovery Is Part of the Work
August 17, 2026 — Effort makes a demand. Recovery determines whether the body can answer it again.
Optimization Without Extremism
August 17, 2026 — The purpose of optimization is to make more of life available—not to make health management consume the life it was meant to improve.
Your Next 30 Years: Building a Personal Operating System for Longevity
August 17, 2026 — A durable longevity strategy is not a shopping list. It is a system for making better health decisions repeatedly as the person and the evidence change.
Discussion
Loading comments…