The STarT Back sorts the low back pain patient into three risk groups to steer treatment. Its algorithm is sequential, and that changes everything: a patient who is severely limited physically but untroubled can never reach the high-risk group. The classifier is right below.
A patient ticking all four physical items reaches 4 out of 9 and stays at medium risk, definitively: the high-risk group is a group of distress, not of severity. The detail is below.
Figures taken from the studies listed at the end of this page, every value carrying its source where it is written. The questionnaire itself is not reproduced.
What the STarT Back decides
The STarT Back does not measure the severity of low back pain, it forecasts its course in order to decide the care pathway. Published in 2008 by Hill et al. in British primary care, it sorts the patient into low, medium or high risk of persistent disabling low back pain.
It has nine items covering the last two weeks. The first four are physical: pain referred to the leg, pain associated with the shoulder or neck, walking limited to short distances, slowed dressing. The next five form the psychosocial subscale: belief that activity is dangerous, worry, catastrophising, loss of enjoyment, overall bothersomeness.
One point of sourcing is worth stating, because it changes how the original study is cited. The thresholds were established in the development sample of 131 patients; the external sample of 500 patients served to evaluate predictive validity, not to derive the thresholds. Writing that the thresholds come from 631 patients would be wrong.
The algorithm, in two steps
This is the point almost everyone skips, and it decides half the classifications.
The algorithm reads in two steps. First step: a total of 3 or less classifies the patient as low risk, and the subscale is not even consulted. Second step, only if the total reaches 4: it is the psychosocial subscore alone that decides, 4 or 5 for high risk, 3 or less for medium risk.
The consequence is sharp and counter-intuitive. A patient ticking all four physical items, with referred pain, associated neck pain, limited walking and slowed dressing, reaches 4 out of 9 with a psychosocial subscore of zero: they are classified medium risk, and nothing in their physical picture will ever classify them high risk. The high-risk group is a group of distress, not a group of severity. Note in passing that the condition 'total of 4 or more' is logically redundant at the second step, since a subscore of 4 already implies a total of 4.
The questionnaire is not the intervention
This is the question every practitioner who starts handing it out asks, and one trial answers it directly.
The founding British trial, in 851 patients, did not test the questionnaire but stratification with matched care pathways. Its benefit was real and modest, and it decayed: effect size 0.32 at four months, 0.19 at twelve. Its primary outcome was the Roland-Morris, never the STarT Back itself.
The American MATCH trial isolates the mechanism by showing what happens when it is missing. Clinicians there were trained seriously, six sessions for physicians and five days for physiotherapists, and the tool was built into the electronic record. They used it for about half their patients, without changing the treatments they recommended. The intervention had no significant effect, neither on patients nor on healthcare use. The Danish trial points the same way differently: no difference in Roland-Morris at three or twelve months, but 3.5 sessions against 4.5 and a lower cost. What the tool best allows is to subtract, not to add.
| What is measured | Value | Sample |
|---|---|---|
| Items, and psychosocial subscale | 9 and 5 | original publication |
| Sample where the thresholds were established | 131 | British primary care |
| Predictive validity sample | 500 | external and independent |
| Effect size of stratification, 4 months | 0.32 | 851 randomised patients |
| The same, at 12 months | 0.19 | same patients |
| Area under the curve, pain prognosis | 0.59 | 1,153 patients, non-informative |
| Area under the curve, disability prognosis | 0.74 | 821 patients, acceptable |
| Test-retest reliability, French version | ICC 0.90 | 108 patients |
| Sessions, Danish de-escalation trial | 3.5 against 4.5 | 169 against 164 patients |
The French version was modified
This is a detail French presentations almost always leave out, and it bears precisely on the part that decides.
During the translation, tested in 44 patients, six subjects out of forty-four, that is 14 %, wondered whether two questions were about the back or about general health. After discussion with the tool's developer, those two questions were modified to add an explicit reference to back pain.
Both items belong to the psychosocial subscale, the one that alone decides high risk. The French questionnaire therefore asks about worry and loss of enjoyment referred to the back, where the English asks about general worry and anhedonia. The French validation that followed, in 108 patients, establishes good reliability, ICC 0.90, and adequate construct validity, but it does not establish predictive validity, and its patients came from a rehabilitation centre, a back school, a practice or a gym, not from the general practice where the tool was built.
Why this page does not display the questionnaire
Keele University distributes the STarT Back without registration or payment, in thirty-eight languages plus a version for children and adolescents. That de facto openness is not a licence, however: the only legal text the questionnaire carries is a copyright notice, that is a claim of rights and not a permission. No item is therefore reproduced here. The official form, French included, downloads from Keele, and that is the one to use, the French version carrying two items deliberately different from the English.
Three administration pitfalls
Classifying high risk on a high total score
The total never opens the door to high risk: only the psychosocial subscore does. A patient at 7 out of 9 with a subscore of 3 is medium risk, exactly like a patient at 4 out of 9 with a subscore of 0.
Announcing a pain prognosis
To discriminate the course of pain, the pooled area under the curve is 0.59, judged non-informative. It reaches 0.74 for disability. The score speaks of future disability, not of future pain.
Using it as a follow-up measure
None of the references on this page reports a minimal clinically important difference for the STarT Back, and none establishes its responsiveness in French. Re-scoring a patient to show progress is supported by none of them.
Frequently asked questions
What does STarT Back stand for?
STarT is short for Subgroups for Targeted Treatment. The full tool is called the Keele STarT Back Screening Tool, and it has no distinct French name: Keele University itself distributes the French version.
How is the STarT Back scored?
Nine items give a total score out of 9, the last five of which form a psychosocial subscore out of 5. A total of 3 or less classifies as low risk; beyond that, a subscore of 4 or 5 classifies as high risk, otherwise medium risk.
Can a patient be high risk without psychosocial distress?
No, and that is the central point. High risk requires a psychosocial subscore of 4 or 5. All four physical items together cap the patient at medium risk.
Is there a validated French version?
Yes, published in 2012 for the translation and in 2014 for the validation, by a Liège team. It establishes a test-retest reliability of 0.90 but not predictive validity, and two of its items were deliberately re-anchored on back pain.
Is the questionnaire enough to improve outcomes?
No. In the American MATCH trial, trained clinicians used it without changing their treatments, and the intervention had no effect. It is the care pathway that acts, not the questionnaire.
References
8 sources, PMIDs included
- Hill JC, Dunn KM, Lewis M, Mullis R, Main CJ, Foster NE, Hay EM. A primary care back pain screening tool: identifying patient subgroups for initial treatment. Arthritis Rheum 2008;59(5):632-41. PMID 18438893. The original publication. It sets the nine items, the five-item psychosocial subscale and the classification algorithm. The thresholds were established in the development sample of 131 patients; the independent external sample of 500 patients served to evaluate predictive validity, not to derive the thresholds.
- Hill JC, Whitehurst DG, Lewis M, Bryan S, Dunn KM, Foster NE, Konstantinou K, Main CJ, Mason E, Somerville S, Sowden G, Vohora K, Hay EM. Comparison of stratified primary care management for low back pain with current best practice (STarT Back): a randomised controlled trial. Lancet 2011;378(9802):1560-71. PMID 21963002. The randomised trial that made the tool famous, in 851 patients. It does not test the questionnaire but stratification with matched care pathways, and its primary outcome is the Roland-Morris at twelve months. The benefit is real and modest, and it decays: effect size 0.32 at four months, 0.19 at twelve.
- Cherkin D, Balderson B, Wellman R, Hsu C, Sherman KJ, Evers SC, Hawkes R, Cook A, Levine MD, Piekara D, Rock P, Estlin KT, Brewer G, Jensen M, LaPorte AM, Yeoman J, Sowden G, Hill JC, Foster NE. Effect of low back pain risk-stratification strategy on patient outcomes and care processes: the MATCH randomized trial in primary care. J Gen Intern Med 2018;33(8):1324-36. PMID 29790073. The American replication, with Hill and Foster among the authors. Despite six training sessions for physicians and five days for physiotherapists, clinicians used the tool for about half their patients without changing the treatments they recommended, and the intervention had no effect, neither on patients nor on healthcare use.
- Morsø L, Olsen Rose K, Schiøttz-Christensen B, Sowden G, Søndergaard J, Christiansen DH. Effectiveness of stratified treatment for back pain in Danish primary care: A randomized controlled trial. Eur J Pain 2021;25(9):2020-38. PMID 34101953. The first large randomised replication in Europe outside the United Kingdom, in 169 against 164 patients. No difference in Roland-Morris at three months nor at twelve, nor in sick leave. The only gain is subtractive: 3.5 sessions against 4.5 and a lower healthcare cost.
- Karran EL, McAuley JH, Traeger AC, Hillier SL, Grabherr L, Russek LN, Moseley GL. Can screening instruments accurately determine poor outcome risk in adults with recent onset low back pain? A systematic review and meta-analysis. BMC Med 2017;15(1):13. PMID 28100231. The meta-analysis that separates what the score predicts from what it does not. To discriminate the course of pain, the pooled area under the curve is 0.59, judged non-informative; for disability, it is 0.74, judged acceptable.
- Bruyère O, Demoulin M, Brereton C, Humblet F, Flynn D, Hill JC, Maquet D, Van Beveren J, Reginster JY, Crielaard JM, Demoulin C. Translation validation of a new back pain screening questionnaire (the STarT Back Screening Tool) in French. Arch Public Health 2012;70(1):12. PMID 22958224. The French translation, tested in 44 patients. A decisive and rarely quoted point: 6 of the 44 subjects, that is 14 %, wondered whether two questions were about back pain or general health. After discussion with the developer, those two questions were modified to add an explicit reference to back pain.
- Bruyère O, Demoulin M, Beaudart C, Hill JC, Maquet D, Genevay S, Mahieu G, Reginster JY, Crielaard JM, Demoulin C. Validity and reliability of the French version of the STarT Back screening tool for patients with low back pain. Spine (Phila Pa 1976) 2014;39(2):E123-8. PMID 24108286. The French validation, in 108 patients. Test-retest reliability of the total score ICC 0.90, internal consistency of the psychological subscale 0.73, correlations of 0.74 with the Roland-Morris and the Örebro. What it does not establish: predictive validity, and the patients came from a rehabilitation centre, a back school, a practice or a gym, not from general practice.
- Al Zoubi FM, Eilayyan O, Mayo NE, Bussières AE. Evaluation of Cross-Cultural Adaptation and Measurement Properties of STarT Back Screening Tool: A Systematic Review. J Manipulative Physiol Ther 2017;40(8):558-72. PMID 29187307. The review of adaptations, across 17 studies and 11 versions in 10 languages. Only two versions, the Belgian-French and the Mandarin, meet all translation requirements. None tested the full set of measurement properties, and the overall quality score is judged poor everywhere except for two versions.
Page written by Anthony Baillon, physiotherapist, co-founder of Physio Learning. The questionnaire itself is by Hill et al., 2008, and Keele University claims its rights: this page documents it and applies its algorithm, it reproduces none of its items.
