Does dual n-back training work?

Last updated 1 August 2026

It makes you better at dual n-back. The claim that it raises fluid intelligence rests on one 2008 study, and the effect largely disappears in the studies designed to test it properly. The pattern in the evidence is unusually clean, which makes this a good case for understanding how a cognitive training claim rises and falls.

Dual n-back deserves its own page rather than being folded into the general brain training question, because it is the one version of the claim that came from a serious laboratory, appeared in a serious journal, and had a plausible mechanism behind it. It is not a marketing invention. That is exactly what makes its trajectory instructive.

What the task actually is

A sequence of squares appears at eight positions on screen, one every three seconds. Simultaneously, one of eight consonants plays through headphones. You respond whenever the current position matches the one shown n steps back, and separately whenever the current letter matches the one heard n steps back.

Two independent streams, held and updated at once. The difficulty adapts: perform well and n increases, struggle and it drops. It is genuinely demanding in a way most brain training games are not, and that difficulty is central to the argument made for it — the reasoning being that a task loading working memory this heavily ought to engage the machinery that fluid reasoning also depends on.

The 2008 result

Susanne Jaeggi, Martin Buschkuehl, John Jonides and Walter Perrig published Improving fluid intelligence with training on working memory in the Proceedings of the National Academy of Sciences in 2008. Healthy young adults trained on adaptive dual n-back and improved on tests of fluid intelligence, with more training sessions producing larger gains.

Two features made it persuasive. The trained task looked nothing like the outcome measure, so the improvement appeared to be genuine far transfer rather than practice. And the dose-response relationship — more training, more gain — is the kind of pattern that is hard to produce by accident.

The paper's own framing was notably careful. It opened by acknowledging the long history of cognitive training research showing that trained-task performance rises dramatically while transfer does not follow. The authors knew what they were claiming to have overturned.

The finding spread far beyond the paper. A generation of free n-back apps and paid brain training products traces its marketing directly to this single result.

The replication that was designed to settle it

Thomas Redick and colleagues published the strongest single rebuttal in 2013, under the title No Evidence of Intelligence Improvement After Working Memory Training: A Randomized, Placebo-Controlled Study.

The design closed the gaps the original left open. Twenty sessions of adaptive dual n-back. An active placebo control group doing adaptive visual search — a task of comparable demand and duration that has no theoretical reason to raise intelligence. And a no-contact control group on top of that.

The dose-response signal did not appear. More dual n-back practice did not produce more fluid intelligence gain. The specific feature that had made the 2008 result convincing was the feature that failed to replicate.

The meta-analysis, and the number underneath the number

Individual studies went both ways after 2008, which is what meta-analysis exists to resolve. The most cited one comes from Au and colleagues in 2015, and it is generally read as sympathetic to training.

Its headline figure is an overall effect on fluid intelligence of g = 0.24. Small, but statistically real.

The breakdown is where it matters. Studies using passive, do-nothing control groups produced g = 0.44. Studies using active control groups — participants doing something else of similar intensity — produced g = 0.06.

Effect of n-back training on fluid intelligence, by control group type (Au et al., 2015).
Control groupEffect size (g)Reading
Passive (does nothing)0.44Moderate
Active (does something else)0.06Essentially nothing
Pooled0.24An average of the two

An effect that is moderate against people who did nothing and absent against people who did something else is not an effect of the training. It is an effect of being in a study: expectation, motivation, the practice of being tested, the simple fact of turning up.

The broader meta-analysis by Melby-Lervag, Redick and Hulme in 2016, covering 87 publications and 145 experimental comparisons across working memory training generally, reached the same conclusion by the same route. Reliable transfer to working memory tasks; no convincing far transfer once treated controls were used.

What this case teaches about reading training claims

The dual n-back story is the clearest available worked example of a pattern that recurs constantly in this field.

When you next encounter a cognitive training product citing a study, the single most useful question is what the control group did. If they did nothing, the study cannot separate the training from the experience of being trained.

Should you do it anyway?

If you enjoy it, there is no reason not to. It is a demanding task, it is free, and getting better at it is a legitimate small achievement.

What the evidence does not support is doing it in the expectation of becoming generally smarter, or of scoring higher on an unrelated reasoning test. Twenty sessions is roughly ten hours. The same ten hours spent on something you actually want to know — a language, an instrument, a technical subject — produces knowledge that is useful in itself, which is a better return than a working memory span that transfers nowhere.

Frequently asked questions

Was the 2008 Jaeggi study wrong?
Not necessarily wrong, but not replicated in the form that mattered. The dose-response relationship between training amount and intelligence gain, which was the study's most persuasive feature, did not appear in Redick's placebo-controlled follow-up.
Why do some meta-analyses find a positive effect?
Because they pool studies with different control conditions. Au and colleagues found g = 0.44 against passive controls and g = 0.06 against active ones. The pooled figure of 0.24 is an average of a real-looking effect and no effect at all.
Does dual n-back improve working memory?
Yes, on working memory tasks resembling the training. That is near transfer and it is not disputed. The contested claim is transfer to fluid intelligence, which is a different thing.
How many sessions would I need?
The studies typically run 17 to 20 sessions of around 30 minutes. That is the protocol under which no reliable intelligence gain was found, so there is no session count that the evidence supports as sufficient.

More in this series