Why no existing tech or methods can reliably tell you why
On the limits of health AI, what honest architecture looks like, and why female hormones and perimenopause is one of the hardest possible domains to get this right
An internal OpenAI LLM just disproved an 80-year-old geometry conjecture that had eluded mathematicians. The headlines called it a breakthrough. Proof that AI can now tackle hard scientific problems that humans couldn’t crack.
It’s worth spending a paragraph on what actually happened, because the story being told about it matters for how we understand what AI can and cannot do in biology and health.
The system that solved the problem wasn’t a language model reasoning its way to a proof. The actual mechanistic explanation is the LLM’s “solution” is a search advantage under exhaustion-free conditions, not a reasoning superiority. It recombined existing tools ( the mathematical tools used were not novel, although their application in this domain appears to be) and explored a search space that humans had cognitively pre-excluded due to prior belief the conjecture was true.
This occurrence does not surprise us, nor does it change our belief that genuinely new, groundbreaking ideas is beyond the reach of current LLMs. LLMs are great at mining the literature for rare gems where humans missed a relatively simple approach, but they cannot generate new mathematical objects. Next-token prediction has no internal representation of logical entailment. It has learned statistical co-occurrence patterns of symbols that humans used when reasoning correctly.
The geometry result was possible because geometry has a closed formal structure: a proof can be verified internally, without appealing to anything outside the system. You don’t need a petri dish to confirm that two angles are equal. The loop closes inside the mathematics.
Biology doesn’t work that way. A hypothesis about what is driving your night sweats cannot be verified inside a model. It has to be tested against a real body, in real time, with real data. The loop closes in an experiment, in a clinic, in the lived experience of a specific woman over weeks and months. Not in the weights of a neural network.
The literature problem nobody talks about
The AI systems being deployed in women’s health are trained on a research literature that is, to put it plainly, not fit for the purpose asked of it.
Perimenopausal biology is chronically understudied. The foundational trial data, as we’ve written about before, was derived largely from older postmenopausal women using formulations and routes of administration that don’t represent the options available today. Much of what passes for evidence about hormonal mechanisms in women has been extrapolated from studies on men or younger women with intact reproductive cycles. The correlation-versus-causation problem is endemic: a literature full of observational associations, confounded by lifestyle, geography, socioeconomic status, and the profound individual variability of the perimenopausal transition itself.
When an AI system pattern-matches over that literature to generate “personalized insights,” it is producing recombinations of those priors. If the priors are biased, context-collapsed, or correlation-heavy, the outputs inherit those properties. A confident-sounding inference delivered in elegant prose is not evidence that the underlying claim is causally valid. It’s evidence that the model has learned to write in the register of confident-sounding medical inference.
This is not a criticism of AI. It is a structural description of what AI can and cannot do at this layer of biological complexity. The question is whether we build a product that obscures this limitation or one that is honest about it and architects around it.
What honest architecture looks like today
Drop has the most distinctive and recognizable symptom signature of all five states:
-
Heavy or prolonged bleeding: the uterine lining built up during Surge now sheds without progesterone to regulate it; often the heaviest period arrives here
-
Hot flashes: intense, sudden, often accompanied by flushing and sweating
-
Night sweats: soaking through clothing and sheets, sometimes multiple times per night
-
Heart palpitations: a racing, fluttering, or skipping heartbeat, particularly at night
-
Difficulty falling asleep: different from Dominance’s 3am waking; Drop insomnia is at sleep onset
-
Sudden low mood or depression: particularly striking after the high-irritability of Surge
-
Brain fog and fatigue: sharp cognitive dulling that can feel alarming
-
Headaches continuing from the Surge→Drop transition
The heart palpitations deserve specific attention because they can be frightening. They are caused by three mechanisms working together: autonomic nervous system dysregulation as estrogen withdraws, direct estrogen receptor effects on cardiac tissue, and the adrenaline surge that sometimes accompanies hot flashes. Women in hormonal transition without underlying cardiac conditions can experience palpitations during Drop. They are commonly benign. However, if they are severe, frequent, or accompanied by chest pain, always discuss them with your doctor.
What honest architecture looks like today
The key structural insight, which almost no consumer health company has operationalized correctly, is that there are two separate functions that must not be conflated.
The first is hypothesis generation. This is where a well-designed AI layer earns its place. Given a user’s symptom profile, behavioral patterns, hormonal state, and goals, an LLM with a well-curated knowledge base can reduce a combinatorial space of thousands of candidate mechanisms to a manageable ranked set of plausible hypotheses worth testing. That is valuable. It does for the individual what a menopause-literate clinician with deep literature knowledge does in a consultation: narrows the search space using prior knowledge so you’re not testing everything at random.
But generating a hypothesis is not the same as validating one. And this is where the second function comes in.
The only valid source of causal signal for a specific woman is her own longitudinal data, structured as a natural experiment. Perimenopause is not a population-level phenomenon that individual women instantiate uniformly. It is an individual-level trajectory that population statistics approximate poorly and sometimes not at all. Hormonal profiles, symptom expression, and the coupling between behavioral inputs and biological outcomes are idiosyncratic enough that a cohort average is often the wrong prior for any specific person.
What produces valid individual signals is: what you did, what you consumed, how you slept, what followed, measured with enough temporal resolution to detect real lead-lag relationships between inputs and outcomes.
As that data accumulates, the system’s hypothesis ranking gets updated by evidence that is actually about you. The LLM proposed. Your data arbitrates. That is a fundamentally different and more realistic process than a model delivering “personalized insights” in week one based entirely on population priors.
What this means for how Waves Women works
In your early days on Waves Women, what it can tell you is what’s plausible for women with your profile. That’s the population knowledge layer: what the literature supports, what women with similar symptom clusters and goals have found helpful, what the evidence says is worth testing first. This is not nothing. It is a meaningful starting point that is better than random and better than generic.
What Waves Women cannot tell you early on is what is specifically causing your symptoms. Because we don’t have your data yet. We have priors. Priors are a starting point, not a conclusion.
As you complete experiments, as you log inputs and outcomes, as the system accumulates temporal data with enough density to detect real signal, the inferences become progressively more individual and progressively more defensible. The transition from “here’s what tends to work for women like you” to “here’s what is actually working for you” is the core product journey. It takes time. It requires sustained engagement. It cannot be shortcut by a model that generates confident outputs from population statistics alone.
Why we think this is the only version worth building
There is a version of this product that delivers confident AI-generated insights from day one, draws on population correlations dressed as personalized analysis, and feels sophisticated on the surface. That product is easier to build, and easier to explain in a pitch or marketing copy.
The architecture we are building is slower to deliver individual signals. It requires the user to trust a process before she sees results. It asks more of the engagement, and it is honest about what it doesn’t yet know.
But it produces something the fast version never can: inferences that are valid for her, derived from her data, updated by her outcomes. Not a statistical average wearing a personalized label.
That is the version we are building. Not because it is the easiest path. Because it is honest and useful.
Waves does not provide medical diagnosis or treatment. Always consult your healthcare provider for personalized medical advice.
