The PRTEE adds up five pain items and ten function items, but the sum of the ten is divided by two before entering the total. The calculator below performs that division. It also recalls the only thing that matters afterwards: the change the patient judges important is worth 9 points, when the instrument only tells a change from noise from 17 upwards.
Two validated French-language versions exist, one from Liège and one from Quebec, and neither publishes a numerical threshold. The detail is below.
Figures taken from the studies listed at the end of this page, every value carrying its source where it is written. The questionnaire itself is not reproduced.
What the PRTEE measures, and why function is halved
The PRTEE has fifteen items, all scored from 0 to 10, split into two subscales of equal weight. Five items for pain, ten for function, and a total out of 100 where pain and disability weigh exactly the same. The direction of the score is that of a complaint: 0 means no pain and no disability, 100 the worst state.
Here is where the scoring most often goes wrong. The five pain items are simply added and give a score out of 50. The ten function items give a raw sum running from 0 to 100, which must be divided by two before being added to pain. Without that division, function would weigh twice as much as pain and the total could reach 150.
The author's manual gives the worked example: pain at 31 out of 50, a function sum of 28 becoming 14 out of 50, and a total of 45 out of 100. That is exactly the operation the calculator at the top of this page performs, and why it asks for the two sums rather than the total.
Two administration details are worth knowing. The ten function items split into six specific activities and four usual activities, but that distinction yields no subscore. And faced with a missing item, the manual asks that it be replaced by the mean of its subscale rather than counted as zero.
One last thing about the name, because it sheds light on the validation population. The instrument was only called the Patient-Rated Tennis Elbow Evaluation from 2005: it was born in 1999 as the Patient-Rated Forearm Evaluation Questionnaire, and its author changed the title because the former one was misleading.
Nine points of importance against seventeen of noise
This is the point that decides what can be concluded from a follow-up, and it was measured within a single cohort, which is rare.
In 100 patients followed test-retest, the minimal important change of the total score is 9 points. The smallest detectable change at 95% is 17. The difference the patient judges important is therefore nearly twice as small as what the instrument can tell apart from noise.
The same study gives both subscales, and the inversion appears again for function: importance at 4 points, detection at 12. For pain, the order is reversed, 11 for importance against 8 for detection. In other words, the pain subscale is the only one of the three where a change judged important exceeds the measurement noise.
Those same authors add a result the systematic synthesis did not suggest: 8 of the 15 items fall below content validity criteria. Their conclusion is cautious and worth quoting as it stands: content validity is questionable, and interpretation of the minimal important change is made difficult by the size of the measurement errors.
The consequence for reading a trial is harsh. The network meta-analysis of treatments for this condition is built entirely on the pain subscale, and the only intervention beating placebo in the medium term is physiotherapy and exercise, with a mean difference of -4.32 points. That is less than half the minimal important change of 11 points established for that same subscale. The best demonstrated effect in the field is smaller than its own importance threshold.
The published thresholds, and the cohorts they come from
The thresholds published on this instrument do not agree, and one must know where the one being used comes from. The table below puts each in front of its cohort.
A second, independent study, in 143 patients followed at three months, places both bounds elsewhere: minimal clinically important difference at 14.8 points and smallest detectable change at 90% at 9.7. It also measures a markedly lower reliability than the pooled value in the literature, 0.62 against 0.96, with an interval running from 0.21 to 0.86. Its authors in fact open their abstract by speaking of contradictory clinimetric data on this instrument.
The systematic synthesis of 21 studies gives distinctly more reassuring pooled values: reliability of 0.96, standard error of measurement of 3.1 and smallest detectable change at 95% of 8.9. Two reading caveats apply nonetheless. The ranges attached to those last two values, 1.8 to 4.4 and 5.3 to 12.5, are not called confidence intervals by the abstract, which reserves that term for the three other quantities. And structural validity and internal consistency are both rated indeterminate there.
One last fact, present from the original publication and never picked up: reliability drops in patients whose epicondylalgia is work-related, 0.80 against 0.94, with a p of 0.018. That is precisely the community caseload. The reference validation, meanwhile, was conducted in 78 tennis players.
| What is measured | Value | Cohort |
|---|---|---|
| Minimal important change, total score | 9 | 100 patients, prospective test-retest |
| Smallest detectable change at 95%, total score | 17 | the same 100 patients |
| The same, pain subscale out of 50 | 11 and 8 | the same 100 patients |
| The same, function subscale out of 50 | 4 and 12 | the same 100 patients |
| Minimal clinically important difference at 3 months | 14.8 | 143 patients, second cohort |
| Smallest detectable change at 90% | 9.7 | the same 143 patients |
| Pooled standard error of measurement | 3.1 | synthesis of 21 studies, 1999 to 2021 |
| Pooled smallest detectable change at 95% | 8.9 | the same synthesis |
Two French-language versions, and no French threshold
Two validated French-language versions exist, which is rather comfortable, and neither carries a French name: both keep the English acronym, where the Swedish version, for instance, received a full title in its own language.
The better documented is the European French-language version, produced by a Liège team in 115 participants, with a test-retest reliability of 0.86 for the global score, an internal consistency of 0.98, and no floor or ceiling effect. The Canadian French one, earlier, covers 32 patients, with an internal consistency of 0.93; it is the one referenced as the French version on the author's official page.
But here is what matters for a report. Neither publishes a numerical threshold. The Liège one writes that it computed the smallest detectable change and does not give its value; the Quebec one reports effect sizes without saying which follow-up time each corresponds to. Every threshold used in French is therefore imported: the 9 and 17 from a Norwegian cohort, the 14.8 and 9.7 from a second cohort, the 8.9 from a synthesis of twenty-one studies.
Why this page does not display the questionnaire
The PRTEE is under its author's copyright, and the official wording was read in a browser: the English questionnaire is 'free to use with permission from the developer', with a credit line to be added to any reprint. That is free use conditioned on individual permission, not an open licence, and that permission has not been requested for this site: the fifteen items are therefore not reproduced here. A useful detail for anyone wishing to request it: the university page that carried this wording now returns a 404 error, and the author's contact address appears in her manual.
Three administration pitfalls
Forgetting to halve the function sum
The raw sum of the ten function items runs from 0 to 100. It is divided by two before entering the total, failing which function weighs twice as much as pain and the score can reach 150.
Concluding from a ten-point gain in a patient
It exceeds the minimal important change, 9 points, but stays below the smallest detectable change at 95%, 17 points, measured in the same cohort. In an individual, that difference supports no conclusion.
Transferring the reference validation's measurement properties to a community patient
It was conducted in 78 tennis players, and the original publication already reported that reliability drops in patients whose epicondylalgia is work-related, 0.80 against 0.94.
Frequently asked questions
What does the PRTEE measure?
It rates pain and disability in lateral epicondylalgia across fifteen items, for a total of 0 to 100 where 100 is the worst state. Pain and function weigh equally there, 50 points each.
How is the PRTEE score computed?
The five pain items are added, giving a score out of 50. The ten function items give a sum from 0 to 100 which must be divided by two. The total is the sum of both, out of 100.
How much change counts?
The minimal important change is 9 points, but the smallest detectable change at 95% is 17 in the same cohort. Below 17, an individual gain cannot be told apart from measurement error.
Is there a French version of the PRTEE?
There are two, a European French-language one and a Canadian French one, both validated. But neither publishes a numerical threshold: those used in France come from elsewhere.
Is the PRTEE free of rights?
No. It is under copyright, and its author's official wording says it is free to use with permission from the developer, with a credit line to be added to reprints.
References
9 sources, PMIDs included
- Overend TJ, Wuori-Fearn JL, Kramer JF, MacDermid JC. Reliability of a patient-rated forearm evaluation questionnaire for patients with lateral epicondylitis. J Hand Ther 1999;12(1):31-7. PMID 10192633. The original publication, under the instrument's first name, Patient-Rated Forearm Evaluation Questionnaire. In 47 patients, test-retest reliability is 0.89 for the global score, 0.89 for pain and 0.83 for function. Above all it carries a result the rest of the literature confirmed and nobody quotes: reliability drops in patients whose epicondylalgia is work-related, 0.80 against 0.94, p = 0.018. Its authors conclude that the size of the measurement error limits prediction of an individual score.
- Macdermid J. Update: The Patient-rated Forearm Evaluation Questionnaire is now the Patient-rated Tennis Elbow Evaluation. J Hand Ther 2005;18(4):407-10. PMID 16271687. The note recording the change of name, in 2005. Its abstract is not indexed in PubMed: its title establishes the renaming, it does not give the reason. That is written in the author's manual, which explains that the former title was misleading and that it was changed to indicate that the measure specifically targets lateral epicondylalgia.
- Rompe JD, Overend TJ, MacDermid JC. Validation of the Patient-rated Tennis Elbow Evaluation Questionnaire. J Hand Ther 2007;20(1):3-10; quiz 11. PMID 17254903. The reference validation under the new name, and the one the author's official page gave as the citation to carry. It establishes an internal consistency of 0.94 for pain, 0.93 and 0.85 for the two function blocks, and a responsiveness above the four instruments compared, 2.1 against 1.5 to 1.7. A crucial point for interpreting everything that follows: its cohort is made of 78 tennis players, only a fraction of the patients seen in practice.
- Sveinall H, Brox JI, Engebretsen KB, Hoksrud AF, Røe C, Johnsen MB. Measurement properties of core outcomes in patients with tennis elbow. Shoulder Elbow 2026;18(4):835-48. PMID 40453542. The most awkward reference in this file, and the only one to put both quantities side by side in a single cohort of 100 patients. The minimal important change of the total score is 9 points; the smallest detectable change at 95% is 17. It also gives both subscales: 11 and 4 for importance, 8 and 12 for detection. And it adds that 8 of the 15 items fall below content validity criteria, which the systematic synthesis did not suggest.
- Young I, Dunning J, Mourad F, Escaloni J, Bliton P, Fernández-de-Las-Peñas C. Clinimetric analysis of the numeric pain rating scale, patient-rated tennis elbow evaluation, and tennis elbow function scale in patients with lateral elbow tendinopathy. Physiother Theory Pract 2025;41(8):1712-20. PMID 39793982. A second, independent set of thresholds, in 143 patients followed at three months: minimal clinically important difference of 14.8 points and smallest detectable change at 90% of 9.7. Reliability there is markedly lower than the pooled value of the meta-analysis, 0.62 against 0.96, with an interval running from 0.21 to 0.86. Its authors in fact open their abstract by speaking of contradictory clinimetric data on this instrument.
- Shafiee E, MacDermid JC, Walton D, Vincent JI, Grewal R. Psychometric properties and cross-cultural adaptation of the Patient-Rated Tennis Elbow Evaluation (PRTEE); a systematic review and meta-analysis. Disabil Rehabil 2022;44(19):5402-17. PMID 34196231. The consensus synthesis, 21 studies from 1999 to 2021 and 13 languages of adaptation. It gives a pooled test-retest reliability of 0.96, a construct validity of 0.81 against the DASH, a standard error of measurement of 3.1 and a smallest detectable change at 95% of 8.9. Two reading points: the ranges attached to those last two values, 1.8 to 4.4 and 5.3 to 12.5, are not called confidence intervals by the abstract, which reserves that term for the correlation quantities it reports, reliability and the construct validities; and structural validity and internal consistency are both rated indeterminate there.
- Lowdon H, Chong HH, Dhingra M, Gomaa AR, Teece L, Booth S, Watts AC, Singh HP. Comparison of Interventions for Lateral Elbow Tendinopathy: A Systematic Review and Network Meta-Analysis for Patient-Rated Tennis Elbow Evaluation Pain Outcome. J Hand Surg Am 2024;49(7):639-48. PMID 38678448. The network meta-analysis of treatments, built entirely on the pain subscale of this instrument. In the short term, no treatment beats placebo. In the medium term, the only intervention that does is physiotherapy and exercise, with a mean difference of -4.32 points, interval -7.58 to -1.07. To be set against the minimal important change of 11 points established for that same subscale: the best demonstrated effect in the field is smaller than its own importance threshold.
- Kaux JF, Delvaux F, Schaus J, Demoulin C, Locquet M, Buckinx F, Beaudart C, Dardenne N, Van Beveren J, Croisier JL, Forthomme B, Bruyère O. Cross-cultural adaptation and validation of the Patient-Rated Tennis Elbow Evaluation Questionnaire on lateral elbow tendinopathy for French-speaking patients. J Hand Ther 2016;29(4):496-504. PMID 27769841. The European French-language version, the PRTEE-F, produced by a Liège team in 115 participants: test-retest reliability of 0.86 for the global score and 0.8 to 0.96 per item, internal consistency of 0.98, high correlations with the DASH and low to moderate ones with the divergent SF-36 subscales, with no floor or ceiling effect. Worth knowing before quoting a threshold in French: its abstract states that the smallest detectable change was computed but does not publish its value.
- Blanchette MA, Normand MC. Cross-cultural adaptation of the patient-rated tennis elbow evaluation to Canadian French. J Hand Ther 2010;23(3):290-9; quiz 300. PMID 20400264. The Canadian French version, the earlier one, in 32 patients: internal consistency of 0.93, item-total correlations of 0.58 to 0.85, correlation with the visual analogue scale of 0.64 to 0.77. This is the one referenced as the French version on the author's official page. A reading point: its abstract gives effect sizes of 0.8 and 1.0 and standardised response means of 0.9 and 1.0 without saying which corresponds to six weeks and which to three months.
Page written by Anthony Baillon, physiotherapist, co-founder of Physio Learning. The questionnaire itself is under copyright of Joy C. MacDermid: this page documents and interprets it, it reproduces none of its items.
