NeuroAudit®
Aivoauditointi™
NeuroAudit® Oy

Generative AI Moves Fast. Science Must Lead.

[AI image] A dark abstract information landscape in which fragmented research signals resolve from left to right into a luminous layered structure in deep navy, gold and turquoise.
The Evidence Engine is the layer where fragmented research signal becomes a traceable, bounded claim. AI-generated image.

Why NeuroAudit® Ltd. is building an Evidence Engine outside the generative model, for brain health, cognitive ergonomics and a more precise reading of the research evidence on women.

Generative AI has changed the speed of health communication. The question is no longer whether a person gets an answer. The question is what the answer rests on and who it actually applies to. In brain health that difference matters, because the same sentence can be a harmless generalisation or an interpretation that steers someone's choices.

NeuroAudit® Ltd. is building an Evidence Engine: a governance layer that sits outside the language model. It determines what kind of scientific evidence may support a claim, which population the claim applies to, and what cannot be inferred from it. The language model may help formulate wording. It must not independently decide scientific truth, source eligibility, population applicability or clinical meaning. The NeuroAudit™ Method is how that principle shows up in practice.

1. A fluent answer is not the same as a verified claim

A language model produces plausible language. Plausibility is not the same property as traceability. In a comparative analysis where language models were asked to produce references under systematic review criteria, the share of hallucinated references, meaning references with at least two of three core details wrong, was 39.6 percent for GPT-3.5, 28.6 percent for GPT-4 and 91.4 percent for Bard (Chelli et al. 2024).

This does not mean a language model is inherently incapable of recognising study designs. It means that language generation on its own is not sufficient scientific governance when the consequences of the communication concern a person's health. An umbrella review of research on language models in healthcare identified promising applications in patient care and decision support, but rated 12 of the 17 included reviews as methodologically low quality and stressed that adoption requires caution alongside ethical and legal regulation (Iqbal et al. 2025).

Wording and verification are two different jobs. A language model optimises how a sentence sounds. Scientific verification asks whether there is an identified publication behind the sentence, how strong its design is, who its data covered, and what limitations the authors recorded themselves. That second job needs rules that sit outside the model and stay the same no matter how the question was phrased.

The real risk is not a visible error. The real risk is fluent overreach: a causal link the data does not carry, a result moved from one population to another, or individual guidance derived from a group-level observation.

2. Scientific evidence requires a governance layer outside the model

Governance applies equally to what a source does not carry. For every accepted source we record the information that constrains interpretation: the design, the population, and the key limitation the authors named themselves. If no eligible source supports a proposed sentence, the sentence is removed or softened. Being bounded is a feature of the product, not a shortfall in it.

In practice, governance means keeping six things apart even though generative output merges them into one piece of text:

3. The evidence gap is both historical and continuing

Accounting for sex in biomedical research has improved, but unevenly. A bibliometric follow-up across nine biological disciplines found that more papers published in 2019 included both sexes than a decade earlier, yet in most disciplines the proportion of studies that analysed data by sex did not increase, and single-sex designs or the absence of sex-based analysis were rarely justified (Woitowich et al. 2020).

In clinical research the picture is field-specific. A cross-sectional analysis of US interventional trials registered over twenty years found female representation lowest relative to burden of disease in fields including neurology, where women carried 56 percent of disability-adjusted life-years and made up 53 percent of participants. The same analysis found men underrepresented in eight disease categories (Steinberg et al. 2021). So this is not one simple bias, but representation that varies by field and by study design.

The gap is not only historical. It continues for as long as results go unreported by sex and life stage in samples large enough to support that analysis. Until then, whoever interprets a finding has to know what the data did not measure. This is the most common way health communication exceeds its evidence: the overgeneralisation does not come from a faulty source, but from a sound source applied to a population it never covered.

A practical rule follows. Types of evidence have to be kept apart: animal studies, mixed-sex cohorts in which results were not disaggregated by sex, female-specific cohorts, and perimenopausal and postmenopausal populations treated separately. Merging these into a single body of research evidence is exactly where interpretation fails. The Evidence Engine records the distinction in machine-readable form instead of leaving it to the writer's memory.

4. Women's brain health as the first high-resolution evidence vertical

The first high-resolution vertical is women's brain health. Not because the question is more interesting than others, but because the margin for interpretive error is widest there: hormonal transitions, sleep, mood, subjective cognitive symptoms and work demands all converge in the same life stage.

In a systematic review and meta-analysis of menopausal stage and cognition, postmenopausal women performed worse on delayed verbal memory tasks than premenopausal and perimenopausal women, and peri- and postmenopausal women were at higher risk of depression than premenopausal women. The authors themselves noted that the results cannot necessarily be generalised beyond the studies included (Weber et al. 2014).

A white paper commissioned by the International Menopause Society describes cognitive changes at midlife as common and frames the practitioner's task as normalising the experience and counselling carefully, rather than reading a symptom as a sign of serious disease (Maki and Jaff 2022).

For sleep, the mechanism is unsettled. A systematic review of perimenopausal sleep disturbances found that good-quality studies pointed to a contribution from the postmenopausal decline in estrogen and progesterone, while the authors stated that the pathophysiology and causes remain poorly understood and that studies of the neurobiological pathways are still needed (Haufe et al. 2022).

Life-stage-aware interpretation does not mean that life stage explains the experience. It means a result is read against the right comparison group. A self-report based result describes how a person characterises their own load and functioning at this moment. It is not a measurement of brain function and not a diagnosis, and it can still be a useful starting point for discussion and follow-up.

Three constraints follow, and in the Evidence Engine they are rules rather than stylistic choices. Age does not indicate hormonal status. Estrogen or progesterone does not explain an individual's cognitive experience. Observational, preclinical or methodologically heterogeneous evidence does not license mechanistic certainty. NeuroAudit® Ltd. does not diagnose symptoms and does not infer anyone's hormonal stage from their age.

5. Cognitive ergonomics moves the question from the individual to the work system

Interruption load, friction in retrieving information, decision load and the absence of structural recovery are properties of a work system. They are hypotheses that organisation-level analysis can examine. They are not measured quantities of brain function, and NeuroAudit® Ltd. does not clinically measure brain function.

There is credible longitudinal evidence on psychosocial work exposure and mental health. In a meta-analysis of eight prospective cohorts covering 84,963 employees and 2,897 new cases of depressive disorder, effort-reward imbalance at work predicted the risk of depressive disorders (pooled estimate 1.49; 95 percent confidence interval 1.23 to 1.80) (Rugulies et al. 2017). That is a group-level risk ratio based largely on self-reported exposure, not a prediction about an individual.

Recovery, in turn, is modifiable. In a meta-analysis of 30 studies and 34 interventions with 3,725 participants in total, interventions designed to improve psychological detachment from work produced a small but significant average effect (d = 0.36), and the effect was larger in longer and higher-dosage interventions (Karabinski et al. 2021).

Where the evidence is thin, that has to be said out loud. In a systematic review of ten observational studies, menopausal status was not consistently related to work ability, the association between symptoms and work performance was mixed, and every included study carried a high risk of bias in at least one assessed domain (Taylor et al. 2025). In the Evidence Engine an area like this is explicitly flagged as uncertain, and guidance formulated from it is not presented as settled.

Looking at the work system also changes which way responsibility points. If load is produced by interruptions, competing priorities and scattered information, it is not fixed by advising an individual to concentrate harder. The things to change are meeting structure, the fragmentation of tasks, expectations of availability, and whether recovery has any protected place inside the structure of work rather than only in off-job time.

Organisational analysis therefore produces group-level observations and development hypotheses that can be used to reshape the structures of work. An individual employee's results do not reach the employer. We do not promise to prevent burnout, dementia, sickness absence or employee attrition.

6. What NeuroAudit® Ltd. is building: current capability

The following list describes the current state, meaning what already exists in the product and in the code:

6b. North star: what is not built yet

The following are on the roadmap. None of them is complete, and none is presented as a current capability:

7. The principle

Population-level evidence is valuable and, at the same time, easy to overread. The 2024 report of the Lancet standing Commission on dementia estimated from modelling that around 45 percent of dementia cases could potentially be prevented or delayed if fourteen modifiable risk factors were addressed across different stages of life (Livingston et al. 2024). That is a population-level estimate, not a promise to an individual, and it must not be converted into a product claim.

This demands restraint from the company. Some of the questions a customer would like to ask us are questions the current evidence does not answer. In those cases the right output is a bounded answer with the uncertainty named, not a more convincing sentence. We have not resolved scientific uncertainty. We are building the layer that makes uncertainty visible and workable.

The goal is not to generate more health content. The goal is to make every consequential claim traceable, bounded and appropriate for the person or population it addresses.

Science must lead.

Sources and research basis

  1. Chelli M, Descamps J, Lavoue V, Trojani C, Azar M, Deckert M, Raynier JL, Clowez G, Boileau P, Ruetsch-Chelli C. Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis. Journal of Medical Internet Research 2024;26:e53164. DOI: 10.2196/53164. PMID: 38776130.
  2. Iqbal U, Tanweer A, Rahmanti AR, Greenfield D, Lee LT, Li YJ. Impact of large language model (ChatGPT) in healthcare: an umbrella review and evidence synthesis. Journal of Biomedical Science 2025;32(1):45. DOI: 10.1186/s12929-025-01131-z. PMID: 40335969.
  3. Woitowich NC, Beery A, Woodruff T. A 10-year follow-up study of sex inclusion in the biological sciences. eLife 2020;9:e56344. DOI: 10.7554/eLife.56344. PMID: 32513386.
  4. Steinberg JR, Turner BE, Weeks BT, Magnani CJ, Wong BO, Rodriguez F, Yee LM, Cullen MR. Analysis of Female Enrollment and Participant Sex by Burden of Disease in US Clinical Trials Between 2000 and 2020. JAMA Network Open 2021;4(6):e2113749. DOI: 10.1001/jamanetworkopen.2021.13749. PMID: 34143192.
  5. Weber MT, Maki PM, McDermott MP. Cognition and mood in perimenopause: a systematic review and meta-analysis. Journal of Steroid Biochemistry and Molecular Biology 2014;142:90-98. DOI: 10.1016/j.jsbmb.2013.06.001. PMID: 23770320.
  6. Maki PM, Jaff NG. Brain fog in menopause: a health-care professional's guide for decision-making and counseling on cognition. Climacteric 2022;25(6):570-578. DOI: 10.1080/13697137.2022.2122792. PMID: 36178170.
  7. Haufe A, Baker FC, Leeners B. The role of ovarian hormones in the pathophysiology of perimenopausal sleep disturbances: A systematic review. Sleep Medicine Reviews 2022;66:101710. DOI: 10.1016/j.smrv.2022.101710. PMID: 36356400.
  8. Taylor S, Callahan B, Grant J, Islam RM, Davis SR. Menopause and work performance: a systematic review of observational studies. Menopause 2025;32(8):769-778. DOI: 10.1097/GME.0000000000002557. PMID: 40460347.
  9. Rugulies R, Aust B, Madsen IE. Effort-reward imbalance at work and risk of depressive disorders. A systematic review and meta-analysis of prospective cohort studies. Scandinavian Journal of Work, Environment and Health 2017;43(4):294-306. DOI: 10.5271/sjweh.3632. PMID: 28306759.
  10. Karabinski T, Haun VC, Nubold A, Wendsche J, Wegge J. Interventions for improving psychological detachment from work: A meta-analysis. Journal of Occupational Health Psychology 2021;26(3):224-242. DOI: 10.1037/ocp0000280. PMID: 34096763.
  11. Livingston G, Huntley J, Liu KY, Costafreda SG, Selbaek G, Alladi S, et al. Dementia prevention, intervention, and care: 2024 report of the Lancet standing Commission. The Lancet 2024;404(10452):572-628. DOI: 10.1016/S0140-6736(24)01296-0. PMID: 39096926.