The Kujala score rates patellofemoral function out of 100. In the same cohort, the change the patient judges important is worth 11 points and the scale's measurement error is worth 13. Its authors write it out. The interpreter below keeps the two quantities apart, because they do not measure the same thing.
A validated French version exists, in 101 patients from Belgium and France, but its study publishes no threshold: every one used in France is imported. The detail is below.
Figures taken from the studies listed at the end of this page, every value carrying its source where it is written. The questionnaire itself is not reproduced.
What the Kujala measures, and who it was built on
The Kujala score, published as the Anterior Knee Pain Scale, is a self-report questionnaire of thirteen items with a variable ordinal response format: each item does not carry the same number of points, and the number of options differs from one to the next. The total runs from 0 to 100, with no subscale, and the direction is the reverse of a pain scale: 100 is the best function.
Its original publication in fact gives it no name. It is simply titled Scoring of patellofemoral disorders and speaks of a new questionnaire. It tests it in 68 exclusively female subjects across four groups: controls score 100 on average, anterior knee pain 83, patellar subluxation 68, dislocation 62.
That breakdown deserves a pause, because it explains what follows. Of those 68 subjects, 35 were unstable, 16 subluxations and 19 dislocations, against only 16 genuinely in pain. The founding population of what everyone now uses as a patellofemoral pain scale was therefore, in the majority, a population of patellar instability.
And the trace of it stayed in the instrument. The original publication names six items that separated its groups most clearly, all at p < 0.0001 and with no ranking between them: abnormal painful kneecap movements, limping, pain, running, stair-climbing and prolonged sitting with the knees flexed. Twenty-seven years later, a Rasch analysis in 646 adolescents with patellofemoral pain flags a single item in significant misfit: that of abnormal kneecap movements. The item that founded the scale is the one that fits it worst in the population where it is used.
Eleven points of importance against thirteen of error
This is the point that decides what can be concluded from a follow-up, and two quantities that practice conflates must first be separated.
An importance threshold answers the question 'does this change matter to the patient?'. A measurement error answers 'is this change real, or did the instrument move?'. These are not the same objects, and they do not belong in a single interval.
Only one cohort measured both together, in 112 Norwegian patients, 50 of them stable for test-retest. The minimal important change there is 11 points. The smallest detectable change there is 13. Its authors write the consequence themselves: a change in AKPS total score of 11 points would be considered important by the patient, although changes up to 13 points may be due to measurement error.
In other words, there is a zone, between eleven and thirteen points, where the patient reports that something has changed for them and where the scale is unable to confirm it. That is the zone the tool at the top of this page names, rather than smoothing it over.
Two other results from the same study are worth knowing. The total score has neither floor nor ceiling effect, which is reassuring, but eight of its thirteen items hit the ceiling taken alone, which is no longer reassuring. And its authors conclude that the weak correlation between pain and the scale questions the validity of the score, recommending that a revision be considered.
The published thresholds, and the populations they come from
The threshold French practice quotes most is 10 points out of 100. It comes from a single study, comparing five instruments in patellofemoral pain, and two clarifications matter before reusing it: it comes from a group analysis in 71 people, not from an individual threshold study; and test-retest reliability there ran from poor to good, 0.49 to 0.83.
Nor is that figure a constant of the instrument. A Thai cohort of 47 patients finds a minimal clinically important improvement at 3.5 points, close to a third. The same study provides the only published target score for this scale in conservatively managed patellofemoral pain: an acceptable symptom state from 72. That is not a change, it is a level, and the two are not interchangeable.
A systematic review of 24 studies adds a value of yet another kind: the smallest detectable change of this scale expressed as a percentage, 9.0%. A percentage does not convert mechanically into points, and that review does not rank the Kujala first either: it recommends as first choice another questionnaire, the Activities of Daily Living Scale of the Knee Outcome Survey.
Finally, a fact that governs all the others: the scale is not an interval measure. The Rasch analysis concludes it must be treated as ordinal data. A ten-point gain is therefore not worth the same depending on where one starts, and a difference in scores cannot be handled like a difference in centimetres.
| What is measured | Value | Population |
|---|---|---|
| Minimal clinically important improvement | 3.5 | 47 patients, Thai cohort |
| Difference judged important, GROUP analysis | 10 | 71 people, randomised trial |
| Minimal important change | 11 | 112 patients, Norwegian randomised trial |
| Smallest detectable change, that is the MEASUREMENT ERROR | 13 | 50 stable patients from the same cohort |
| The same, expressed as a percentage | 9.0% | review of 24 studies, five questionnaires |
| Patient acceptable symptom state, a LEVEL not a change | 72 | the same 47 patients |
The French version, and its imported thresholds
A validated French version exists, and that is good news: it was produced by a Liège team in 101 patients recruited in French-speaking Belgium and in France, mean age 34.5, with a six-step translation following international recommendations.
Its properties are solid: test-retest reliability of 0.97, internal consistency of 0.87, coherent construct validity against the SF-36, with strong correlations on physical function and role physical and weak ones on mental health, and no floor or ceiling effect.
But what that validation does not establish matters just as much. Its abstract reports neither a smallest detectable change, nor a minimal important change, nor a target score for the French version. Every threshold used in France is therefore imported: the 10 points from a study comparing five instruments, the 11 and 13 from a Norwegian one, the 3.5 and the 72 from a Thai one.
A detail that matters for a report: the scale has no French name. The validation itself is published under the English title, and there is no French-language publication on this instrument. Unlike the Roland-Morris, which became the EIFEL, no French title has ever been introduced into the indexed literature.
Why this page does not name the items in question
Two studies challenge the structure of this scale, and one of them shows that a form reduced to eleven items fits better than the full version, while correlating at 0.966 with it. Its abstract designates those two items by number and does not name them. This page therefore does not name them either. The adversarial review of this file had flagged, as a blocking fault, an earlier version that attributed to them labels absent from the source, and inverted relative to the original numbering as well. As no licence is established for this instrument, its thirteen items are not reproduced here.
Three administration pitfalls
Treating the 10 points as an individual threshold
They come from a group analysis in 71 people. The published importance thresholds run from 3.5 to 11 points depending on population and method, and the scale's measurement error is 13 points, more than the largest of them.
Filing the measurement error among the importance thresholds
These are two different constructs. Three and a half, ten and eleven answer 'does this matter to the patient?'; thirteen answers 'can the instrument see it?'. An interval running from 3.5 to 13 mixes the two questions.
Treating a difference in points as a distance
The Rasch analysis concludes that the scale does not meet interval measurement criteria: a ten-point gain is not worth the same at 55 and at 85. The score is read, it is not subtracted like a length.
Frequently asked questions
What does the Kujala score measure?
It rates patellofemoral function across thirteen items with unequal weights, for a total of 0 to 100 where 100 is the best function. The direction is the reverse of a pain scale.
How much change counts?
Published importance thresholds run from 3.5 to 11 points depending on the population. But the scale's measurement error is 13 points: below that, an individual change supports no conclusion.
What score should be aimed for at discharge?
The only published target score is an acceptable symptom state from 72, established in 47 patients followed under conservative treatment. It is not a change but a level reached.
Is there a French version of the Kujala score?
Yes, validated in 2019 in 101 patients from French-speaking Belgium and France, with a reliability of 0.97. That study, however, publishes no threshold: every one used in France comes from elsewhere.
Does the Kujala score measure pain?
Poorly, and its own evaluators say so: the weak correlation between pain and the scale questions the validity of the score. The instrument was in fact built on a population made mostly of patellar instabilities.
References
9 sources, PMIDs included
- Kujala UM, Jaakkola LH, Koskinen SK, Taimela S, Hurme M, Nelimarkka O. Scoring of patellofemoral disorders. Arthroscopy 1993;9(2):159-63. PMID 8461073. The original publication, which gives the scale no name and speaks only of a new questionnaire. It tests it in 68 exclusively female subjects across four groups: 17 controls at 100 points on average, 16 anterior knee pain at 83, 16 patellar subluxations at 68 and 19 dislocations at 62. It names six items that separated the groups most clearly, all at p < 0.0001 and with no ranking between them, including abnormal painful kneecap movements, limping, pain, running, stair-climbing and prolonged sitting with the knees flexed.
- Ittenbach RF, Huang G, Barber Foss KD, Hewett TE, Myer GD. Reliability and Validity of the Anterior Knee Pain Scale: Applications for Use as an Epidemiologic Screener. PLoS One 2016;11(7):e0159204. PMID 27441381. The open-access text carrying the description of the scoring, and the only one read in full. The instrument has 13 items with a variable ordinal response format: each item does not carry the same number of points and the number of options differs from one to the next. The total runs from 0 to 100. That same text contains no occurrence of 'permission', 'freely available' or 'copyright holder'.
- Crossley KM, Bennell KL, Cowan SM, Green S. Analysis of outcome measures for persons with patellofemoral pain: which are reliable and valid? Arch Phys Med Rehabil 2004;85(5):815-22. PMID 15129407. The single source of the 10 points out of 100 that became the French reflex. A decisive clarification before reusing that figure: it comes from a group analysis in 71 people, not from an individual threshold study, and test-retest reliability there runs from poor to good, 0.49 to 0.83. The scale is however the most responsive of the five compared, effect size 0.98 against 0.37 for the functional index questionnaire.
- Hott A, Liavaag S, Juel NG, Brox JI, Ekeberg OM. The reliability, validity, interpretability, and responsiveness of the Norwegian version of the Anterior Knee Pain Scale in patellofemoral pain. Disabil Rehabil 2021;43(11):1605-14. PMID 31583918. The most awkward reference in this file, because it measures both quantities together and finds the wrong one larger. In 112 patients, 50 of them stable for test-retest, the minimal important change is 11 points and the smallest detectable change is 13. Its authors write it out: a change in AKPS total score of 11 points would be considered important by the patient, although changes up to 13 points may be due to measurement error. It also shows that the absence of a ceiling effect on the total hides a ceiling on eight items out of thirteen, and concludes that the weak correlation between pain and the scale questions its validity.
- Selhorst M, Fernandez-Fernandez A, Cheng MS. Rasch analysis of the anterior knee pain scale in adolescents with patellofemoral pain. Clin Rehabil 2020;34(12):1512-9. PMID 32674606. The Rasch analysis in 646 adolescents with patellofemoral pain, median score 73, observed range 7 to 100. It establishes that the scale does not meet interval measurement criteria: a ten-point gain is not worth the same at 55 and at 85. It flags a single item in significant misfit, that of abnormal painful kneecap movements, precisely the one the original publication listed first among its six most discriminating items. And only five items out of thirteen order their response options correctly.
- Apivatgaroon A, Tungbanjerdsook P, Chernchujit B, Sanguanjit P. Clinically meaningful thresholds for the Thai Kujala score in patellofemoral pain syndrome: Minimal clinically important improvement and patient acceptable symptom state. J ISAKOS 2026;19:101169. PMID 42398829. The numerical proof that the 'plus ten' is not a property of the instrument. In 47 patients followed at one and three months, the minimal clinically important improvement falls to 3.5 points, close to a third of the most quoted figure. This study also provides the only patient acceptable symptom state published for this scale in conservatively managed patellofemoral pain: a target score of 72, area under the curve 0.708.
- Buckinx F, Bornheim S, Remy G, Van Beveren J, Reginster J, Bruyère O, Dardenne N, Kaux JF. French translation and validation of the "Anterior Knee Pain Scale" (AKPS). Disabil Rehabil 2019;41(9):1089-94. PMID 29264931. The only validated French version, produced by a Liège team in 101 patients recruited in French-speaking Belgium and in France, mean age 34.5. It establishes excellent test-retest reliability, 0.97, an internal consistency of 0.87, coherent construct validity against the SF-36, and no floor or ceiling effect. What it does not establish matters as much: its abstract reports neither a smallest detectable change nor a minimal important change. Every numerical threshold applied in France is therefore imported.
- Esculier JF, Roy JS, Bouyer LJ. Psychometric evidence of self-reported questionnaires for patellofemoral pain syndrome: a systematic review. Disabil Rehabil 2013;35(26):2181-90. PMID 23627531. The systematic review that puts this scale in competition with four others in patellofemoral pain syndrome, and which does not rank it first: it recommends as first choice the Activities of Daily Living Scale of the Knee Outcome Survey, the Kujala remaining adequate. It also gives its smallest detectable change expressed as a percentage, 9.0%, to be compared with the 8.3% of the first and the 30% of another knee score. A percentage does not convert mechanically into points.
- da Silva-Júnior FB, Dibai-Filho AV, Barros DCC, Dos Reis-Júnior JR, Gonçalves MBS, Soares AR, Cabido CET, Pontes-Silva A, Fidelis-de-Paula-Gomes CA, Pires FO. Anterior Knee Pain Scale (AKPS): structural and criterion validity in Brazilian population with patellofemoral pain. BMC Musculoskelet Disord 2024;25(1):39. PMID 38191375. A second challenge to the structure, in 101 participants: two items carry a factor loading below 0.23, and an 11-item form excluding them fits better while correlating at 0.966 with the long version. Its abstract does not name those two items, and this page therefore does not name them either: this is the blocking correction raised by the adversarial review of this file.
Page written by Anthony Baillon, physiotherapist, co-founder of Physio Learning. The questionnaire itself is by Kujala et al., 1993: this page documents and interprets it, it reproduces none of its items.
