hit counter

SDQ Had the Strongest Evidence Among Broad ADHD Rating Scales for Children

Among broad child behavior questionnaires used to screen for ADHD, the short hyperactivity-inattention subscale of the Strengths and Difficulties Questionnaire (SDQ) has the strongest measurement evidence, according to a 2026 review of 36 studies. None of the 12 scales reviewed met every quality standard.1

Research Highlights

  • 36 studies, 12 scales: a systematic review in BMC Pediatrics graded the ADHD subscales of 12 broad child behavior rating scales against the COSMIN standards for judging health questionnaires.1
  • SDQ came out ahead: its 5-item hyperactivity-inattention subscale earned a passing overall rating on every property that was tested, with high-quality evidence for 4 of them, including reliability over time and telling children with and without ADHD apart.1
  • CBCL red flag: high-quality evidence showed the Child Behavior Checklist (CBCL) Attention Problems subscale did not measure the same thing across countries, so COSMIN rules would recommend against it; the authors advise caution outside North America.1
  • No scale was fully tested: no study of any scale checked measurement error, sensitivity to change over time, or agreement with a gold standard.1
  • Screening, not diagnosis: an earlier study found the SDQ hyperactivity rating missed more than half of children who were diagnosed with ADHD.2

Why Broad Behavior Questionnaires Matter in ADHD Assessment

ADHD has no biological test, so diagnosis leans heavily on how parents, teachers and children describe behavior. Rating scales turn those descriptions into scores.1

These scales come in 2 broad types:

  • Narrow-band scales, such as the Conners Rating Scales and the ADHD Rating Scale-IV, ask only about ADHD symptoms and rate their severity.
  • Broad-band scales, such as the SDQ and the CBCL, cover a child’s overall behavior and emotions, which helps catch problems that often travel with ADHD, such as anxiety, depression and oppositional behavior.

A broad-band scale is useful for ADHD only if its attention and hyperactivity items work well on their own. The SDQ, for example, has 25 items split into 5 subscales, and just 5 of them cover hyperactivity and inattention.1

The latest COSMIN guidance (Consensus-Based Standards for the Selection of Health Measurement Instruments) stresses that each subscale of a questionnaire has to be evaluated separately.4 Chen and colleagues did exactly that, rating only the ADHD-related subscales rather than each questionnaire as a whole.1

How the Review Graded 12 Rating Scales

The team searched PubMed, Web of Science, Scopus and ProQuest up to December 1, 2025. Of 17,361 records screened, 36 English-language studies of children 18 or younger with ADHD made the cut, covering 12 broad-band instruments.1

Each subscale was judged on a set of measurement properties. In plain terms, these ask:

  • Content validity: do the questions actually cover ADHD, in words parents and children understand?
  • Structural validity and internal consistency: do the items hang together as one attention or hyperactivity score?
  • Cross-cultural validity (measurement invariance): do the items mean the same thing in different countries, languages or groups?
  • Test-retest reliability: does the same child get a similar score a little later?
  • Construct validity: do scores line up with other ADHD measures and separate children with ADHD from those without it?

Each result was rated sufficient or insufficient, and the quality of the evidence behind it was graded high, moderate, low or very low. Under COSMIN rules, a scale earns a full recommendation only with high-quality evidence of good results on every relevant property. High-quality evidence of a poor result on any single property means a recommendation against it.1

SDQ Hyperactivity-Inattention Subscale Had the Strongest Evidence

Seven studies with a combined 19,477 children covered the SDQ. Its hyperactivity-inattention subscale received a sufficient overall rating on every property that was tested, although not every individual study agreed:1

  • High-quality evidence for cross-cultural validity, test-retest reliability, scores that tracked the CBCL’s ADHD-related scores, and separating children with ADHD from those without it
  • Moderate-quality evidence for structural validity and internal consistency, downgraded because results were inconsistent: 2 studies that used a weaker method (principal component analysis) got poorer results
  • Low-quality evidence for content validity, because no study tested it directly
Grid comparing the ADHD subscales of 5 broad child behavior scales (SDQ, CADBI, FTF, CBCL and CABI) on 6 measurement properties. The SDQ has green cells for all 6, 4 of them backed by high-quality evidence. The CBCL has a red cell backed by high-quality evidence for measuring the same thing across groups. Many cells for CADBI, FTF, CBCL and CABI are gray because the property was never studied.
Evidence quality for each measurement property of 5 well-studied ADHD subscales. Green means the property passed, red means it failed, and gray means no included study tested it.1

The authors also point to practical strengths: the SDQ is short, has matching parent, teacher and self-report versions, and is freely available in many languages.1

They are careful not to oversell it. The SDQ is a general screening tool, and its hyperactivity-inattention subscale is not a dedicated ADHD diagnostic measure or a substitute for one. The review also did not compare instruments head-to-head, so it does not show the SDQ is the best tool for ADHD overall.1

One included study shows why. Hall and colleagues gave the SDQ to 250 children aged 6 to 17 who had been referred for an ADHD assessment. A “probable” hyperactivity rating correctly ruled out about 75% to 85% of children who did not have ADHD, but it flagged only about 43% to 45% of those who did.2 A low score, in other words, does not rule ADHD out.

CBCL Attention Problems Subscale Failed the Cross-Cultural Test

The CBCL is one of the most widely used child behavior checklists. Its 11-item Attention Problems subscale held up on internal consistency (moderate-quality evidence) and separated children with ADHD from controls in Brazilian and German samples (high-quality evidence).1

The problem was cross-cultural validity. Two well-conducted studies found the subscale did not behave the same way in US and Australian samples, or in Dutch and Israeli samples. Because that evidence was rated high quality, COSMIN rules would recommend against using this subscale.1

One of those studies tested the CBCL’s structure in about 6,700 clinic-referred children from the US, the Netherlands and Australia. About 90% of items held up across models and countries, but the attention factor and especially the social problems factor found the least support.3

The authors link this to the CBCL’s origins in a mostly North American sample, and they advise users outside North America to be cautious. The verdict covers only the Attention Problems syndrome scale. The review did not evaluate the CBCL’s separate DSM-oriented ADHD scale or the questionnaire as a whole.1

CADBI, FTF and CABI Look Promising but Incomplete

  • CADBI (Child and Adolescent Disruptive Behavior Inventory): separate 9-item inattention and hyperactivity-impulsivity subscales had high-quality evidence for structural validity and for working the same way across time, sex and countries. No included study checked its reliability over time or how it compares with other measures.1
  • FTF (Five to Fifteen): a Nordic parent questionnaire with moderate-quality evidence for structure and internal consistency. One study found its scores did not differ between boys and girls in the expected way.1
  • CABI (Child and Adolescent Behavior Inventory): a newer, expanded version of the CADBI with high-quality evidence for test-retest reliability, but only low-quality evidence for structure and internal consistency, mostly from small samples.1

The other 7 instruments had limited or unsatisfactory evidence for their ADHD subscales. Four of them (ANSER-PQ, CBAI, YI-4 and DSMD) rested on a single study each. Two more, the ABC and CSI-4, had 2 studies apiece, both from the same research group and published more than a decade ago.1

Big Gaps in How ADHD Rating Scales Have Been Tested

Research has favored the easiest properties to measure. Across the 36 studies:1

  • Internal consistency: 28 studies (77.8%)
  • Structural validity and convergent validity: 16 studies each (44.4%)
  • Discriminant validity: 13 studies (36.1%)
  • Cross-cultural validity: 9 studies (25.0%)
  • Test-retest reliability: 8 studies (22.2%)
  • Measurement error, responsiveness to change, and criterion validity: none

Responsiveness matters if a scale is used to track whether treatment is working, and none of these subscales has been tested for it. Criterion validity needs a gold standard to compare against, and ADHD has no universally accepted one.1

Content validity was also weak across the board. Most scales described how their items were written only vaguely, and for several older instruments the original development manuals could not be found.1

What This Means for Parents and Clinicians

A broad behavior questionnaire is a starting point. It can flag attention problems and other difficulties worth a closer look, but an ADHD diagnosis rests on a full clinical assessment, often with ADHD-specific scales and reports from more than one setting.

Among broad-band options, the SDQ hyperactivity-inattention subscale currently has the best-documented measurement properties. Clinicians using the CBCL Attention Problems subscale with children from different cultural backgrounds have reason to interpret scores carefully. The authors note that a separate review of narrow-band ADHD scales is underway.1

Limitations

  • Ratings, not accuracy: the review graded the evidence behind each property; it did not pool sensitivity or specificity or compare scales directly.
  • English only: validation work published in other languages was excluded, which may matter most for cross-cultural ratings.
  • Reviewer judgment: content validity ratings relied heavily on the reviewers when development records were missing, and the authors acknowledge some subjectivity in applying COSMIN.
  • Rigid criteria: COSMIN treats different factor-analysis methods and correlation types the same way, and some newer statistical approaches do not fit its framework.
  • Protocol change: the registered plan was to rate whole scales; the team switched to rating ADHD subscales, in line with COSMIN guidance, and documented the change.

References

  1. Chen Y, Cai Q, Hu S, Zhao Z, Wang J, Niu L. How do the measurement properties of the broad-band scales for attention-deficit/hyperactivity disorder in children perform based on the COSMIN guidelines? BMC Pediatrics. 2026;26:907. doi:10.1186/s12887-026-07661-1
  2. Hall CL, Guo B, Valentine AZ, et al. The validity of the Strengths and Difficulties Questionnaire (SDQ) for children with ADHD symptoms. PLoS ONE. 2019;14(6):e0218518. doi:10.1371/journal.pone.0218518
  3. Heubeck BG. Cross-cultural generalizability of CBCL syndromes across three continents: from the USA and Holland to Australia. Journal of Abnormal Child Psychology. 2000;28(5):439–450. doi:10.1023/A:1005131605891
  4. Mokkink LB, Elsman EBM, Terwee CB. COSMIN guideline for systematic reviews of patient-reported outcome measures version 2.0. Quality of Life Research. 2024;33(11):2929–2939. doi:10.1007/s11136-024-03761-6

Related Posts:

Mental Health Research Updates

Weekly insights on medications, supplements, and brain health.

We respect your privacy. Unsubscribe anytime.

Leave a Comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.