Clinical Workflows

What a Positive Phalen's Test Is Actually Worth

Published accuracy for Phalen's test ranges from 0.50 to 0.70 sensitivity depending on which paper you read. Why the spread is that wide, what the 2024 AAOS guideline changed, and what to write in the chart.

ZD

Zdrovia Editorial

7 August 202610 min read

Ask ten clinicians how accurate Phalen’s test is and you’ll get ten confident answers that don’t agree. That isn’t a knowledge gap. The published figures genuinely are that far apart, and the reasons why turn out to be more useful than the numbers themselves.

Two credible papers, very different answers

The largest synthesis available is a 2023 meta-analysis in Cureus by Ozdag and colleagues, pooling 67 articles covering 8,924 hands. Median values across those studies:

Test Sensitivity Specificity
Phalen’s 0.70 0.80
Durkan (carpal compression) 0.67 0.74
Tinel’s 0.59 0.80
Two-point discrimination 0.51 0.90
Semmes-Weinstein monofilament 0.39 0.85

On those numbers Phalen’s looks respectable, and better than Tinel’s on sensitivity while matching it on specificity.

Now compare a prospective study in the Journal of Hand Surgery Global Online by Zhang and colleagues, which ran 85 hands from 55 patients against electrodiagnostic testing. Phalen’s came out at 0.50 sensitivity and 0.33 specificity. Tinel’s at 0.47 and 0.56. Durkan’s at 0.71 and 0.22.

A specificity of 0.33 means that in that cohort, two-thirds of the patients without electrodiagnostically confirmed CTS still had a positive Phalen’s. Against the pooled median of 0.80, that’s not a small disagreement.

Why the spread is so wide

Very little of it is about the test being performed badly.

Start with what counts as positive, which is not standardised at all. Some protocols require numbness in a median distribution; others accept any paraesthesia in the hand. Some hold the position 60 seconds, some 30. Move that threshold and you slide along the same tradeoff curve, which makes comparing one paper to another treacherous before you’ve read a single result.

Then there’s the reference standard, where the circularity gets awkward. Zhang’s group used nerve conduction studies as the truth. But CTS is a clinical diagnosis that electrodiagnostics support rather than define, so a patient with convincing symptoms and normal conduction studies gets counted as a false positive instead of as early disease. Measure a provocative test against an imperfect gold standard and it will read worse than it performs.

And who’s in the room matters more than either. A hand surgery clinic sees patients already filtered by a referrer, which narrows the range of severity and does unpredictable things to both numbers. A physiotherapy caseload looks nothing like that.

Zhang’s authors landed on the observation that ties this together: as the manoeuvre becomes more provocative, the test gets more sensitive and less specific. Durkan’s applies direct pressure and picked up the most cases at 0.71, while flagging almost everyone else too, at 0.22 specificity. Durkan’s isn’t performing badly there. You’re just watching the tradeoff happen in a single row of a table.

What changed in 2024

The more important development isn’t about any individual test.

The AAOS clinical practice guideline for carpal tunnel syndrome was updated in 2024, replacing the 2016 edition, and it carries a strong-evidence recommendation that CTS-6 can be used to diagnose carpal tunnel syndrome in lieu of routine ultrasonography or NCV/EMG. On moderate evidence it also recommends against MRI and upper limb neurodynamic testing for diagnosis.

That’s a substantial reframing. A validated clinical instrument now stands in for electrodiagnostics as routine practice, which is good news for anyone working where nerve conduction studies mean a three-month wait.

CTS-6 comes from Graham and colleagues in the Journal of Hand Surgery, 2006, built by Delphi consensus: 57 candidate findings ranked by expert clinicians, the top eight combined into 256 case histories, then scored by separate expert panels and fitted with logistic regression. The model’s predicted probability correlated with the clinicians’ at 0.71.

Its six weighted items are numbness predominantly or exclusively in the median nerve distribution, nocturnal numbness, thenar atrophy, loss of two-point discrimination, a positive Phalen’s test, and a positive Tinel’s sign. Scoring above 5 puts the probability of CTS around 0.25; above 12, around 0.80.

Phalen’s didn’t get discarded in that framework. It got put back in proportion, as one weighted input among six, sitting alongside the history rather than carrying the diagnosis by itself. Running Phalen’s, getting a positive, and stopping there means you’ve completed one item of a six-item instrument.

Our Phalen’s test reference sets out the technique, the accuracy ranges across the literature, and how it compares with Tinel’s and Durkan’s, which is the comparison the numbers above make unavoidable.

The timing detail almost nobody records

A Brazilian study on Phalen test positivation time graded 33 patients by how fast symptoms appeared: severe under 10 seconds, moderate between 10 and 30, mild beyond 30. Compared against electroneuromyography, 26 of 33 agreed, or 78.8 percent.

Break that down and it gets more interesting. Agreement was 100 percent at both extremes, for the severe and the mild groups. In the moderate band it collapsed to 46.2 percent.

So a Phalen’s that goes positive in eight seconds is telling you something you can lean on, and so is one that takes forty. The middle band is close to a coin flip against electrodiagnosis and shouldn’t be driving a decision by itself.

Which makes time to onset the single most valuable thing you can add to the record, and it’s the thing that almost never gets written down. “Phalen’s positive” throws away the one variable that separates a finding you can act on from one you can’t. It’s a small sample and it needs replication, but the direction is clear enough to change what you chart tomorrow.

Rule out the things that mimic it

Median distribution is doing real work in the CTS-6 criteria, and it’s the part that gets eyeballed rather than tested.

Numbness across the whole hand isn’t median. Numbness that includes the little finger isn’t median. Symptoms in the thumb and index alongside neck involvement may well be a C6 radiculopathy, and a positive Phalen’s in that patient is a distraction rather than a finding. Our dermatome map is a quick way to check a reported distribution against a segmental level before you commit to a working diagnosis.

The same reasoning applies to any provocative test with modest standalone accuracy. Kernig’s sign has the same shape of problem in a completely different context: low sensitivity means a negative tells you very little, and treating it as a rule-out is a well-documented way to be wrong.

What belongs in the note

If Phalen’s is one item inside a six-item instrument, the chart entry has to preserve enough for the instrument to be reconstructed later. In practice that means:

  • Side. Bilateral is common and asymmetry matters.
  • Time to onset, in seconds, not just positive or negative.
  • Symptom distribution, specifically whether it was confined to the median territory.
  • Which position was used, since wrist flexion, reverse Phalen’s, and direct compression are different tests with different operating characteristics.
  • The other CTS-6 items you checked, including the negatives. A documented absence of thenar atrophy is a finding.

The reason to be strict about this is repetition. These tests get repeated at follow-up, and a second “Phalen’s positive” laid next to a first “Phalen’s positive” tells you nothing about whether the patient is worse. Two entries reading positive at 30 seconds and then positive at 8 seconds tell you a great deal.

Making it survive contact with the chart

None of the above happens reliably in a free-text box. On a full day, “Phalen’s positive” is what gets typed, because typing the full finding as a sentence is a paragraph of work.

We’ve written separately about why assessment scores disappear into prose, and provocative tests are the sharpest version of that problem, since there’s no headline number to anchor the entry. Zdrovia’s smart blocks give side, onset time, distribution, and position their own fields inside charting that still reads normally, so a follow-up sits next to the baseline without anyone reconstructing it. For a clinic deciding what its documentation should look like before the caseload gets big, our guide to software for a new Ontario clinic covers the surrounding decisions.

The short version

  • Published accuracy for Phalen’s runs from 0.50 to 0.70 sensitivity depending on the positivity threshold, reference standard, and case mix. Pick a protocol and stay on it.
  • More provocative manoeuvres buy sensitivity and pay for it in specificity. Durkan’s at 0.71 and 0.22 is the clearest illustration.
  • Since 2024, AAOS recommends CTS-6 on strong evidence in place of routine NCV/EMG or ultrasound, and recommends against MRI and neurodynamic testing.
  • Phalen’s is one of six weighted CTS-6 items, so a positive on its own is a fraction of an answer rather than a diagnosis.
  • Time to onset separates the trustworthy positives from the ambiguous ones. Under 10 seconds and over 30 agreed with electrodiagnosis completely in one small series; the middle band agreed less than half the time.
  • Check the distribution against a dermatome before you settle on CTS.
  • Record side, seconds, distribution, position, and the other items you checked. Otherwise the follow-up has nothing to compare against.

The instinct to treat a provocative test as a yes-or-no is what makes it weak. Recorded properly, with a time and a territory attached, it’s one of the more useful things you can do in ninety seconds without equipment.

← Back to Blog