The assumption behind every "positive" or "negative" screen
Somewhere on most caseloads sits a folder with a completed Social Communication Questionnaire (SCQ) or Autism Spectrum Rating Scale (ASRS) form, a number circled at the bottom, and an implicit decision already half-made: screen positive, refer for full evaluation; screen negative, move on. That shorthand treats a parent-report score as if it were a clean yes/no signal. Five sources — a Norwegian population cohort, a Prader-Willi syndrome sample, a tertiary ASD-specialty clinic, a national comorbidity registry, and two CDC surveillance reports — show that parent-report screening and community diagnosis break down in several different ways, and only one of them directly links screener accuracy to a child's language or cognitive level.
The defensible rule: don't let a single below-cutoff SCQ score, in a child with typical or near-typical language around age three, rule out ASD on its own — and don't let any single parent-report score, at any age, rule ASD in or out without a closer look. The studies below show why, and where that rule's evidence runs out.
Same construct, very different tests
Two well-known parent-report tools designed to flag the same construct — autism spectrum symptoms — are often discussed as though their numbers sit on the same scale. They don't, and the difference matters before comparing anything else about them.
A 2019 study drawn from the Norwegian Mother and Child Cohort tested the Sensitivity and specificity of early screening for autism using the SCQ at 36 months. Of 58,520 mothers who responded to the questionnaire (a 58% response rate), 385 children (0.7%) were later identified with ASD. At the manual's recommended cutoff of 15, sensitivity was just 20%, while specificity was 99%. In plain terms: at its standard cutoff, the SCQ almost never produces a false alarm, but it fails to flag roughly four out of five children who will later be identified with ASD.
The ASRS comes from a different kind of sample entirely. A 2022 psychometric study of 490 children at a tertiary ASD-specialty clinic, Psychometric Evaluation of the Autism Spectrum Rating Scales (6–18 Years Parent Report) in a Clinical Sample, found the ASRS parent form had favorable sensitivity but poor specificity, and concluded it did not perform effectively to differentiate ASD from other clinical disorders. True-positive results were more likely for children with a multiracial background and less likely for children with high social capital — the study's own explanation for some of its variability, and one that has nothing to do with cognitive level.
It is tempting to line these two findings up as opposite failure modes of the same measurement problem — one instrument that under-flags, one that over-flags. Resist that temptation, or at least qualify it heavily. The SCQ study screened a general population at 36 months, where ASD prevalence was 0.7%; the ASRS study evaluated children already referred to a specialty ASD clinic, where the comparison group was other clinical diagnoses, not typically developing peers. One study reports a specificity number; the other reports only that specificity was "poor," without a comparable figure. These are not head-to-head numbers, and a sentence like "the ASRS is more likely than the SCQ to produce a false positive" is not something this data can actually support. What each study does support, on its own terms, is narrower: the SCQ at its standard cutoff is a poor way to rule ASD in, and the ASRS on its own is a poor way to rule other clinical presentations out.
The subgroup one population study found hardest to catch
The Norwegian SCQ study did not just report an overall sensitivity number — it broke results down by developmental profile at 36 months. Sensitivity for children with ASD who had not yet developed phrase speech was 46% (cutoff 15), compared with only 13% for children with ASD who had phrase speech. The authors were explicit about the implication: the SCQ "identified mainly individuals with ASD with significant developmental delay and captured very few children with ASD with cognitive skills in the normal range."
It's worth being precise about what that phrase actually measures, because different proxies for "cognitive level" show up across this set of studies and they are not interchangeable. Study 1 stratifies by whether a child had developed phrase speech by 36 months — a marker of global developmental delay, not a standardized IQ score. The Prader-Willi study below measures verbal and composite IQ directly, on a standardized test. And the phrase clinicians use in practice — "average or near-average cognitive and language skills" — blends both ideas together. Keep them separate: what follows about phrase speech is a language-delay marker, not an IQ score, even though the study's own authors connect it to "cognitive skills in the normal range" in their conclusion.
Lowering the cutoff to 11 raised sensitivity overall to 42%, with specificity falling to 89%. Broken out by subgroup, sensitivity for children without phrase speech rose from 46% to 69% — a 1.5x increase — while sensitivity for children with phrase speech rose from 13% to 34%, a proportionally larger, 2.6x increase. So the lower cutoff did help the phrase-speech group more, proportionally, than it helped the more delayed group. What it didn't do is close the gap: even at 34%, two-thirds of children with phrase speech who had ASD were still missed, compared with roughly a third of children without phrase speech, and the specificity cost of the lower cutoff — from 99% down to 89% — applied to the entire sample, not just to this subgroup. The authors' own conclusion was that sensitivity could not be raised without severely compromising specificity; the data don't support the stronger claim that lowering the cutoff left the higher-language subgroup completely unrescued, but they do support a real, persistent gap between the two subgroups at both cutoffs tested.
A genetically distinct sample adds a different piece of evidence about cognitive level, though not the piece it's sometimes assumed to add. A 2017 study assessed 146 children and youth with Prader-Willi syndrome (PWS), a condition with a known elevated rate of autism features, in Diagnoses and characteristics of autism spectrum disorders in children with Prader-Willi syndrome. Best-estimate clinical diagnoses, built from ADOS-2 videotape review, calibrated severity scores, developmental history, and current functioning, identified ASD in only 18 children (12.3%) — far lower than prior estimates of 25–41% that had relied solely on parent screeners. The study attributes that gap mainly to the unreliability of prior estimates that relied solely on parent screeners. Separately, it reports that compulsivity and insistence on sameness in routines, seen in 76–100% of children regardless of ASD status, was robustly correlated with lower adaptive functioning. Whether that same overlap also inflates screener scores independent of a child's cognitive profile is a plausible extrapolation, not a mechanism the study tested. That's a different explanation than "average-cognition children get over-flagged" — it's a PWS-specific behavioral overlap, and the study doesn't isolate cognitive level as the driver.
Where this study does speak to cognitive level directly is in the opposite direction from a "blind spot" narrative: children with PWS plus ASD had lower verbal and composite IQ, and lower adaptive daily living and socialization scores, than children with PWS only. In this sample, lower cognitive scores tracked with an ASD diagnosis, not the reverse. Most children with PWS alone showed only sub-threshold social difficulties, which the study's authors say "could signal risks for other psychopathologies." It's a reasonable extension, though not something the study tested, that some of those same sub-threshold behaviors are the kind of items that show up on a parent checklist and could nudge a score upward without an ASD diagnosis being warranted. But that's a plausible mechanism, not something the study tested.
When parent-report drives more than screening
Parent-report doesn't only influence whether a child gets referred for evaluation — it can also shape which comorbid diagnoses accumulate in a child's community record. A 2011 study used a national online registry to examine Parent Report of Community Psychiatric Comorbid Diagnoses in Autism Spectrum Disorders in 4,343 children with ASD. Adjusted odds of a parent-reported lifetime psychiatric comorbidity (anxiety, depression, bipolar disorder, or ADHD/ADD) were significantly higher with each additional year of life, with increasing autism severity, and with Asperger syndrome or PDD-NOS diagnoses compared to autistic disorder. That detail matters less as a cognitive-level clue than as a labeling one: this study did not measure screener accuracy or IQ, and its authors were careful to note their findings could reflect both real comorbidity trends and variation in how community providers apply diagnostic labels. The practical point is narrower — community-assigned comorbidity labels track diagnostic category and severity, which is itself a reason to verify an inherited label rather than assume it reflects the child in front of you.
Community evaluation practices also determine who gets counted as a case at all, independent of any instrument or any child's presentation. The CDC's ADDM Network's two reports on the 2012 surveillance year — a 2016 release and a 2018 revision — both found significantly higher prevalence at sites reviewing both education and health records than at sites reviewing health records alone: 17.1 versus 10.7 per 1,000 in the 2016 report, and 17.1 versus 10.4 per 1,000 in the 2018 revision. Individual sites ranged more widely, from 8.2 to 24.6 per 1,000, but that range reflects different states with different populations and service systems, not a clean record-type effect. Case counts, in short, are downstream of system design as well as of any single instrument's accuracy.
Why direct observation outperforms checklists, even at genetically elevated risk
The PWS study gives the cleanest head-to-head comparison available in this set. Within the same 146-child sample, the SCQ yielded only a 29–49% chance that a screen-positive case actually had ASD by best-estimate diagnosis. The ADOS-2, a structured direct-observation instrument, had higher sensitivity, specificity, and predictive values than the SCQ in that same population. The study also flagged the opposite error: some children were ADOS-2 positive but were judged by the clinical team not to have ASD, and these children tended to have communication problems that could mimic autism-specific behaviors on a structured observation without reflecting the full diagnostic picture. Even the stronger instrument wasn't infallible — it simply outperformed the parent checklist by a wide margin, in a population where PWS-related compulsivity and insistence on sameness make screening-only judgments especially unreliable.
A decision rule for your caseload
None of these five studies were designed to tell you how to run your clinic, and none of them were conducted in a school setting — the Norwegian sample was a general population screened by mail at 36 months, the ASRS sample and the PWS sample were both clinic-based, and the comorbidity and surveillance data come from national registries and CDC records. That's a real limit on how far any of this generalizes to a school-based caseload. But the one study that did test for a cognitive-level effect found something specific enough to act on: a parent checklist, at both cutoffs tested, was worse at catching ASD in children who already had phrase speech at 36 months than in children who didn't.
In practice, that supports a few concrete habits. Treat an SCQ score above cutoff as a reason to look more closely, not as confirmation — its 99% specificity at the standard cutoff means false alarms are rare, but its 20% sensitivity means real cases are missed far more often than not. Treat a below-cutoff SCQ score as insufficient grounds to close a case for a child with typical or near-typical language — but note that this specific recommendation rests on one instrument (the SCQ), one age (36 months), and one cohort (a Norwegian mail survey with a 58% response rate), and hasn't been established for the ASRS, for older children, or for a school-based population. Treat an ASRS positive as a reason to differentiate carefully between ASD and other clinical presentations before accepting it, given its documented poor specificity in a 490-child clinical sample — and keep in mind the study's own moderators (multiracial background, social capital) rather than assuming cognitive level explains its false-positive rate, because that study never measured cognitive level at all. And treat comorbidity labels already in a child's file as worth verifying rather than inheriting, since the registry study of 4,343 children found they track diagnostic category and severity in ways that may reflect community labeling habits as much as clinical reality.
Here is what this set of studies actually converges on, and where it doesn't. One study — the Norwegian population cohort — directly measured how screening accuracy relates to a proxy for language and developmental level, and found a real, persistent gap: a parent checklist missed the majority of children with typical language development even after the cutoff was loosened. That's a documented finding from a single study using a single proxy variable, not a consensus reached by four different samples. The other studies in this set show something different, and just as worth knowing: that parent-report and community-based diagnostic processes are unreliable for reasons that have nothing to do with a child's cognitive level at all — a genetic syndrome's own behavioral profile can inflate screener scores, a psychiatric label in a chart can reflect diagnostic culture as much as a child's presentation, and which records a surveillance site reviews is associated with a meaningfully different prevalence estimate (17.1 versus 10.4–10.7 per 1,000 in the aggregate comparison). The defensible claim is narrower than "checklists always fail typical-cognition kids": parent-report tools fail for several different, non-overlapping reasons, only one of which has been shown to track cognitive or language level directly — and none of which a five-minute checklist score can substitute for your own observation.