An encouraging treatment result can lose statistical significance at follow-up even when its estimated effect remains much the same. Calling that a fading benefit turns uncertainty into a clinical conclusion.
A randomized study published September 19 in Molecular Psychiatry makes the distinction concrete. Laurel Morris and colleagues tested a single session of 7-Tesla MRI biofeedback in depression. The active group showed greater improvement than sham on a derived mood-and-affect measure immediately after training. At 30 days, the standardized effect estimate was nearly identical to the immediate one, despite a lower estimate at seven days. Yet the 30-day comparison no longer crossed the conventional significance threshold.
That is a reason to examine precision and measurement, rather than declare that the treatment wore off. It is also only half the story. The primary brain-activation outcome was nonsignificant, and active feedback was associated with greater reported engagement. The experiment raises questions about both persistence and what produced the early clinical signal.
What participants actually did
The analyzed sample included 32 unmedicated participants with major depressive disorder and 30 healthy controls. Within the depression group, 17 received active feedback and 15 received sham feedback. The total of 62 is not the size of the depression treatment comparison. Here, “unmedicated” included avoiding CNS-active medications for one week before MRI; that does not by itself indicate whether participants were drug-naive or off treatment long term.
Participants completed a single scanning session using a 7-Tesla MRI system. During training, they tried mental strategies intended to increase motivation while watching a thermometer-like display. The active group saw feedback derived from their own ventral tegmental area, or VTA. The sham group saw replayed feedback from an active participant. Both groups performed the task; the comparator was not simply waiting without an intervention.
A final run removed the feedback so researchers could examine regulation after training. Follow-up assessments extended to 30 days.
This particular setup matters. The study tested a defined MRI procedure with a defined comparison. Its findings cannot establish the effectiveness of an EEG headset, a home brain-training product, or biofeedback in general. Those would be different interventions requiring their own evidence.
The positive result needs its proper name
The principal clinical analysis used a latent improvement factor derived from mood and affect measures. In plain language, the investigators combined patterns of change across questionnaire subscales into a statistical measure of improvement. The primary factor model used the Profile of Mood States and the Positive and Negative Affect Schedule.
That approach can summarize related changes and reduce the need to test each subscale separately. It is still a derived outcome. Improvement on it should not be translated into a remission rate, a percentage of patients who recovered, or a demonstrated return to everyday functioning.
The active depression group improved more than the sham depression group immediately after training. The reported difference was also significant at the exploratory 24-hour follow-up. The immediate group-by-intervention interaction supported a larger active-versus-sham difference in the depression group than in healthy controls (p = 0.033). That strengthens the early clinical signal, although it does not identify its mechanism. Exploratory analyses found greater improvement with active feedback in depressed mood and negative affect, while positive affect did not show an intervention effect. Anticipatory, but not consummatory, anhedonia also improved at 24 hours.
The Montgomery-Åsberg Depression Rating Scale (MADRS), administered by blinded clinical raters, was not included in the primary factor model. The authors cite its longer look-back period as a limitation for interpreting early change. MADRS was collected at early assessments and included in a supplementary 24-hour factor model; the standalone MADRS follow-up analyses began at seven days.
At seven and 30 days, the supplement reports improvement over time across the depression participants. That paragraph supplies no active-versus-sham contrast or corresponding interaction estimates. It therefore supports overall improvement, without resolving the added benefit of active feedback on clinician-rated depression severity.
The endpoint history adds an important qualification. The earliest and latest public registry versions list the same single primary outcome involving VTA activation. The publication additionally describes the derived factor as a primary clinical outcome, and the supplement distinguishes that clinical analysis from the registered neural outcome. Readers should not assume that the positive clinical factor was the registered primary endpoint.
This does not erase the clinical signal. It changes how much confirmatory weight that signal can carry.
A changed p value is not a measured loss of benefit
The reported standardized between-group effect was d = 0.88 immediately after training, 0.95 at 24 hours, 0.63 at seven days, and 0.87 at 30 days. At the last visit, the estimate was nearly the same size as immediately after training.
The uncertainty was different. The immediate model coefficient was 15.85, with a 95% confidence interval of 1.68 to 30.02. At 30 days, it was 17.34, with a wider interval of -1.48 to 36.15. The seven-day and 30-day comparisons had p values of 0.14 and 0.069. Both intervals included zero.
Fewer depression participants attended later assessments: the supplementary visit table lists 24 at seven days and 24 at 30 days, compared with 32 in the scanning sample. These visit counts are not necessarily the analytic sample for every model. Attrition is relevant to uncertainty and potential bias, but the available summaries do not isolate its contribution to the changed p values.
There is also a measurement qualification. The factor was derived separately at each timepoint. A one-factor solution appeared in 98% to 100% of iterations at the earlier assessments and 73% at 30 days. The authors report additional evidence of factor stability, but these estimates should not be treated as repeated readings from a fixed clinical ruler. Similar effect sizes alone cannot establish unchanged benefit.
A mixed-effects model combining all post-training assessments (immediate, 24 hours, seven days, and 30 days) found an overall intervention effect favoring active treatment in the depression group (p = 0.043). That supports an average difference across the observation period; it does not resolve the uncertainty at the final visit.
The later comparisons therefore neither demonstrate fading benefit nor establish durable efficacy. Their point estimates, intervals, and changing measurement structure support a more precise conclusion: persistence remains uncertain.
What produced the early signal?
The intervention was designed around the VTA, a region involved in motivation and reward-related circuitry. Yet the primary comparison of VTA activation change from baseline to the final transfer run did not show significant effects or interactions.
The investigators did find exploratory associations between aspects of VTA activity and clinical improvement, as well as findings involving connectivity. These observations can help develop hypotheses. They cannot substitute for the missing group-level effect on the primary neural outcome.
The paper itself makes this limitation explicit. Its proposed mechanistic link rests on preliminary correlational evidence. An exploratory mediation analysis was nonsignificant, and an analysis relating within-person change in regulation to clinical change did not confirm the association.
The study's power calculation drew on a large VTA-regulation effect previously observed in healthy volunteers. The authors acknowledge that effects in depression may differ.
The feedback experience offers another plausible explanation for some of the clinical change. Both groups watched a display and tried motivational strategies, but only active feedback tracked the participant's own signal. Replayed feedback controls for many features of the session without necessarily matching the experience of receiving responsive feedback.
Participants in the active groups reported higher effort/importance on the Intrinsic Motivation Inventory (p < 0.001). An engaging, responsive task could affect immediate mood ratings through routes other than the proposed VTA-training mechanism. Engagement was assessed after training, so the analysis leaves open whether it caused improvement or accompanied it.
Perceived ability to self-regulate did not differ significantly by group or intervention. The investigators report that participants and clinical raters were blinded, while study coordinators and the principal investigator were not; participants were not asked to guess their assignment, so masking was not directly tested in that way.
Finally, the feedback used an fMRI blood-oxygen-level-dependent signal. It did not directly measure correction of a dopamine deficiency. A clinical signal, a neural mechanism, and a neurotransmitter explanation remain separate claims.
What to say when a patient asks
The early difference on a derived mood-and-affect measure is encouraging, but it is not a demonstrated advantage in clinician-rated depression severity.
The treatment conversation should stay anchored to the actual intervention and outcome. A claim about “neurofeedback for depression” needs enough detail to identify the equipment, training schedule, comparison group, symptom measure, and follow-up period. Without those details, the name can make distinct procedures sound interchangeable.
Repeated training is a reasonable research question. The authors discuss additional sessions and booster protocols as future directions. This single-session trial did not establish an effective maintenance schedule, however, and it cannot tell a patient how many sessions would produce a lasting benefit.
The study gives the next experiment two tasks: determine whether the clinical difference persists and distinguish the proposed neural mechanism from the experience of responsive feedback. Losing significance answers neither question. The early signal remains worth investigating, with its size, uncertainty, and alternative explanations kept in view.
Sources
Morris LS, Beltrán JM, Kvamme TL, et al. Targeting the dopaminergic midbrain with precision 7-Tesla biofeedback training in depression: A proof-of-principle randomized controlled trial. Molecular Psychiatry. Published September 19, 2026. doi:10.1038/s41380-026-03911-x. Open-access primary article; full main text reviewed.
Morris LS, et al. Supplementary information accompanying the trial. Participant flow, visit counts, design and primary-outcome description, and supplementary clinical analyses reviewed.
ClinicalTrials.gov. NCT04138680: Real-time Biofeedback With 7-Tesla MRI for Treatment of Depression. Public registry record, last update posted September 4, 2024; current record retrieved September 29, 2026. Earliest (version 1, dated October 23, 2019) and latest (version 10, dated September 1, 2024) registry versions also compared. Used for the registered primary outcome, not as independent replication of the paper.
Sources reviewed September 29, 2026. Clinical interpretation and suggested questions are editorial synthesis. This article concerns one experimental MRI protocol and does not establish the effectiveness of other neurofeedback approaches.
Educational Disclaimer: The Psychiatric Record provides general educational information for psychiatric and mental-health professionals. Content does not constitute medical, legal, regulatory, compliance, billing, or other professional advice; does not establish a standard of care; and is not a substitute for independent professional judgment. Appropriate assessment and treatment depend on the individual, setting, and current authoritative guidance. This article does not recommend replacing established depression care with an experimental intervention. Patients should discuss treatment decisions with their treating clinician.
