Essay 8 min read

Most anxiety apps say they work. Almost none show their work.

App stores are full of apps promising to reduce anxiety and improve mood. Audits keep finding almost none of those claims are backed by cited research. Our fix is unglamorous, label every technique's evidence like a food label lists ingredients: strong, promising, or traditional. Some of our labels are modest, on purpose.

The claims gap

In 2019, researchers audited 73 of the top-rated mental-health apps on the two major app stores, the ones users would find first, searching for help. Sixty-four percent claimed to be effective at diagnosing, treating, or improving a mental-health condition.[1] Forty-four percent leaned on scientific-sounding language as their main way of backing that claim, rather than pointing to a specific study.[1]

The language didn't point at much. Of the apps that claimed a scientific basis, 33% described techniques with no published research behind them at all.[1] Only one of the 73 apps cited actual published research anywhere in its store listing or marketing. Zero cited any certification or accreditation.[1]

This wasn't a one-off finding from a single research team. One Mind PsyberGuide, a nonprofit that exists specifically to rate mental-health apps for credibility, has now scored 161 of them using its own methodology. The average score: 2.51 out of 5.[2] Most apps in the category, by the rating body built to judge them, land closer to unproven than to trustworthy.

Seventy-three top-rated mental-health apps. One citation of published research between them.[1]

Claims are cheap. Clearance is rare.

Claiming effectiveness costs nothing. Proving it is a different exercise, and regulators already have rules for exactly that gap. The Federal Trade Commission requires health claims to rest on "competent and reliable scientific evidence", in practice, randomized, controlled testing on humans, not testimonials or an app's own internal data.[3] Lumosity found that out the hard way, paying $2 million in 2015 after the FTC found its brain-training claims weren't backed by the science it cited.[3]

Very few mental-health apps clear a higher bar than a store listing: FDA clearance as a digital therapeutic, the same kind of review a medical device goes through. DaylightRx, a CBT-based program for generalized anxiety disorder, was cleared in September 2024, after a trial in which 71% of participants reached remission by week ten.[4] Rejoyn, a working-memory and CBT program for depression, was authorized in April 2024, among the first FDA-authorized digital therapeutics for the condition.[5] Both had to show their trial data to a regulator before they could say what they say.

Thousands of apps in this category claim to work. A small handful have actually cleared a regulator on that claim. The American Psychiatric Association built a five-step framework to sort the difference, accessibility, privacy, clinical evidence, engagement, and interoperability, each step harder to pass than the last. When it ran 92 apps through all five, only 8% passed the step that checks for a clinical evidence base. Just 1% passed every step.[6] Most apps never get past the easy parts.

The techniques are not the problem

None of this means the techniques inside these apps are useless. That's the part worth sitting with: several of the practices most apps build around have real evidence behind them. The apps just rarely say so accurately, or say so at all.

A meta-analysis of 12 randomized trials, 785 participants total, found that slow-paced breathing techniques reduced stress (g = −0.35), anxiety (g = −0.32), and depression symptoms (g = −0.40) relative to controls.[7] That is the evidence a technique like box breathing draws on: not a trial of box breathing in isolation, but a body of randomized research on slow-paced breathing as a category, of which box breathing is one well-known form.

Cyclic sighing, a specific slow-exhale pattern also known as the physiological sigh, has its own dedicated randomized trial: 111 participants, five minutes a day, for a month. It outperformed a mindfulness-meditation control group on positive affect and lowered participants' resting respiratory rate, a physiological result, not just a self-reported one.[8]

None of that adds up to a cure, and no honest label should imply one. A 2024 meta-analysis of 176 randomized trials on mental-health apps generally found small-to-medium effects on anxiety and depression, with an effect size between g = −0.35 and −0.40.[9] That is a real, useful effect, comparable in size to the breathing research above. It is not transformative. Apps don't need to inflate what they offer; the honest ceiling is already worth having.

What honest labels look like

Reground rates every technique on a three-tier scale, and the numbers above are exactly what that scale is built on.

Strong means meta-analyses or large randomized trials tested this exact practice. Not the family it belongs to. Not a longer programme that contains it as one step. Two techniques in the app clear that bar: the body scan, which a large multi-site trial trained people to do on their own, exactly the way the app teaches it, and progressive muscle relaxation, behind which sits a systematic review of forty-six studies.

Promising means real trials support it, but they are few, small, or tested something close rather than exactly this. Box breathing and the physiological sigh live here. The breathing research above is genuinely good, and it is research on slow-paced breathing as a category, of which these are two named forms. That distinction is the whole difference between the two labels.

Traditional means long use in clinical or contemplative settings, with limited direct study either way. The 5-4-3-2-1 grounding sequence sits here: recommended by the WHO, familiar to any clinician, and never isolated in a trial of its own. So does the STOP practice, one skill from a dialectical behavior therapy curriculum whose group programmes are well studied while the four steps alone are not. It is not a demotion, and it doesn't mean a technique doesn't work. It's a description of what does and doesn't exist yet in the published literature.

We tightened that first definition in July 2026, and it cost us. Under the older wording, a technique could inherit "strong" from the category it belonged to, and eight techniques did. Under this one, six of them moved down, including the two named above. Nothing about the research changed. The label had simply been describing the family rather than the exercise. A scale that never moves against you isn't a scale.

The label sits on the technique's card inside the app, and again at the top of every guide on this site, not buried in a footnote but next to the instructions themselves. In the app it now carries a "Sources" link beside it: the papers, what each one measured, and where it stops short of the exercise you are about to do. The point isn't to make everything look strong. It's that a label willing to say "promising," or even "traditional," is the only kind whose "strong" actually means something.

How to read any app's claims

None of this requires a research background to check yourself. A few plain questions travel well across any app in the category, ours included.

Look for a technique with a real, searchable name, box breathing, progressive muscle relaxation, cognitive restructuring, instead of a proprietary "method" that exists only inside one app's marketing. A named technique can be checked against actual published research by anyone with a search engine. A trademarked one usually can't, because there's nothing outside the app to check it against.

Look for citations to journals, not to the phrase "studies show." A real citation names a paper, a year, a sample size, something you or anyone else could go find and read for yourself. "Studies show," on its own, names nothing and commits to nothing.

Notice whether the claim is about the technique or about the app. A citation showing slow breathing reduces anxiety says nothing about whether one particular app's implementation of it works, whether people stick with it, or whether the app resembles the version that was actually studied.

And remember the regulatory bar, because it's a useful yardstick even outside a courtroom: the FTC's own standard for a health claim is randomized, controlled testing on humans.[3] Most of what fills an app store description doesn't clear that bar, and most of the time, doesn't even claim to.

The honest limits

Our own labels deserve the same scrutiny we're asking you to apply to everyone else's. They are our editorial judgment, applied as consistently as we know how, not a certification from any outside body, and not equivalent to FDA clearance or a peer-reviewed rating.

Evidence moves, and labels have to move with it. A technique labeled "promising" today could earn "strong" after the next trial is published, or new research could complicate a label we currently list as settled. Nothing here is static, and a label that stops updating stops being honest.

Nobody audits us the way PsyberGuide audits the industry.[2] That absence is exactly why every guide on this site links its primary sources at the bottom, the same way this essay does, so anyone can check a label against the paper it's supposedly built on, instead of taking our word for it.

A disclosure, kept to the end on purpose: we make Reground. The three-tier label sits on every technique card in the app, not just in this essay. Read the technique guides on this site, where each label carries its own sources, and judge them the way you'd judge any claim here, by checking, not by trusting.

Sources reviewed · July 2026

The labels ship with the app, not just the essay.

Every technique card carries its tier, Strong, Traditional or Promising, and the sources behind it. Judge them the way you judged this page.

Free to try. $5.99 once for all 22 techniques, no subscription.

Download on theApp Store