Assessment and Monitoring
01What is this guide for?
Two basic questions have to be answered in a child's auditory rehabilitation: where to begin and how to know whether there is progress. The answer to both rests on measurement.
This guide addresses how that measurement is set up. Which tool measures what, which one is informative in which age range, how the resulting number is to be read, and which decisions follow from it — these are the subject of this page.
The frame. A scale does not show what the child can do; it shows as much as that scale is able to measure.
The reason for that careful wording is this: most tools in this field are informative over a narrow age range and stop discriminating beyond it; some measure a construct other than the one they are assumed to measure; and for many there is no evidence to decide whether a change in score is genuine growth or measurement error. Using these tools with their limits in mind gives markedly more reliable results than using them without.
Suggested reading first: Foundations of Auditory Rehabilitation and Auditory Perception Development. The second in particular sets up the Erber levels and the typical developmental line used throughout this page.
02What does assessment measure?
The most common conceptual confusion in this field comes from three different questions being used interchangeably. They measure three independent constructs.
Can the child hear it? Do speech sounds fall within the area the child can hear? This is a question of audibility; it belongs to the device and the audiogram. It concerns the ear, not the child's skill.
Can the child tell it apart? What can be done with the sound that arrives? Can sounds be told apart, words recognised, sentences understood? This is a question of performance and it is measured under controlled conditions, in the test room.
Does the child use it? Are these skills actually used in daily life — in a noisy classroom, in a crowd, when called from behind? This is a question of function, and only the people who observe the child through the day can answer it.
No construct substitutes for another
This is the point that most needs emphasis on this page, because it is so often overlooked in clinical practice.
The audiogram does not adequately predict language development. In a study following 182 children who use hearing aids, unaided pure-tone average failed to predict the risk of language delay; aided audibility, speech recognition in noise and daily hours of device use, by contrast, predicted that risk significantly (Wiseman, McCreery & Walker, 2023). The inference "the loss is moderate, so a delay of this size is not expected" is therefore riskier than it appears.
A functional scale may be measuring something other than hearing. In a longitudinal study with monthly measurement over two years, LittlEARS scores correlated moderately to strongly with a language development inventory but showed no significant relationship with PEACH. The authors' reading is that LittlEARS may be measuring general communicative development rather than hearing specifically (Persson et al., 2019). When a family is asked "does the child hear?", the question actually answered may be "does the child communicate?".
Conventional measures underestimate the difficulty of daily life. Pure-tone average, speech detection threshold in quiet and word recognition in quiet have been reported to make the real communication difficulty of children with hearing loss look milder than it is (Hillock-Dunn et al., 2015).
Clinical note. A child who identifies 90% of words correctly in a quiet test room but cannot follow the teacher in class is not a contradiction; two different constructs have been measured. The task is not to argue which one is "right" but to read the two together.
03What should a measurement tool offer?
Before turning to the tools themselves, a few concepts need to be in place so that the tools can be judged.
Normative data. Does the tool have a reference group to compare against? If there is no answer to "what score is expected of a 24-month-old?", the value obtained can only be compared with the child's own earlier measurements, not with other children.
Validity. Does the tool actually measure the construct it sets out to measure? As the LittlEARS example shows, this is a harder question than it looks.
Reliability. Does the same child, measured twice under the same conditions and a short interval apart, produce a similar result? Do two different examiners arrive at the same result?
Responsiveness. When the child genuinely progresses, can the tool capture that change? A tool can be reliable and still be blind to change. This concept should not be confused with the diagnostic sensitivity of a test.
Two further concepts are vital for monitoring.
Ceiling and floor effects. A child who has reached the highest score a tool can give can no longer be monitored with it; the skill keeps growing but the score does not move. That is the ceiling effect, and it is present in almost every tool in this field. The floor effect is its opposite: if an infant's pre-implant score sits at zero, the tool is carrying no information about that child.
Critical difference. Is a rise from 68 to 74 progress, or is it no more than the natural variability of the measurement? Answering that requires knowing the tool's measurement error. PEACH is the only tool for which this has been calculated; the topic returns in section 9.
Student note. Most of these tools were developed in the 1990s as clinical monitoring instruments, not as psychometric tests. That does not mean they should not be used; it means they should not be asked for something they cannot deliver. That a scale is widely used does not make it psychometrically strong.
04Functional scales
In the youngest age group, measurement in the test room is limited; an 18-month-old cannot be given a closed-set word recognition test. That gap is filled by putting structured questions to the person who observes the child through the day. In the literature these tools are called parent-report measures.
The LittlEARS Auditory Questionnaire
What it measures. The emergence of auditory behaviour in the first two years. Thirty-five questions of the form "does the child turn towards the source of a sound?", "does the child respond to their name?" are answered yes or no; the total score is the number of "yes" answers.
Who it is for. Children aged 0–24 months chronologically, or whose hearing age has not passed 24 months. The second criterion is often overlooked: a child implanted at three years is eligible for LittlEARS through the first two years after implantation.
The evidence. A normative study of 3,309 children across sixteen countries and fifteen languages. Internal consistency is high (α = 0.96) and the relationship between score and age is strong (r ≈ 0.90). The norm curves of different languages overlap almost exactly, which is why the tool is treated as language-independent (Coninx et al., 2009).
How to read it. The score is judged not on its own but against the expected value and the lower-limit curve for age. The real information lies not at a single measurement point but in the direction of the child's curve relative to the normative band. A study following seventy infants shows this plainly: 88% of infants with mild to moderate loss sat within the expected range at an early measurement, yet 29% of them fell below the band at the next one (Visram et al., 2022). A first measurement within normal limits therefore does not rule out later deviation.
Its limits. The items are weighted towards the detection level; few items measure discrimination. In addition, normally hearing children reach 35 points at about 24 months, and the expected score at 18–19 months is already 31. The discriminating range of the tool is therefore narrower than it appears.
PEACH and TEACH
What they measure. The child's auditory behaviour in quiet and in noise, judged from the family's observations over the past week. PEACH scores eleven items: six on the quiet condition, five on the noisy one.
Who they are for. After LittlEARS reaches ceiling, from about two years to the preschool period. TEACH, the teacher version, is used in the school setting.
The evidence. A normative study with the families of ninety normally hearing and ninety hearing-impaired children. Internal consistency α = 0.88; test–retest reliability r = 0.93; inter-rater agreement 0.95 (Ching & Hill, 2007). An independent validation showed the norms hold for children who use hearing aids, but added that scores come out significantly lower in children under 20 months and that norm interpretation calls for care below that age (Bagatto & Scollie, 2013).
What sets it apart. It is the only tool in this field with a calculated measurement error. The critical difference at the 95% level is 11 points; when a child's change in PEACH percentage falls below that value it cannot be told apart from measurement error. This single figure is what separates PEACH from the other scales for monitoring purposes.
Its limits. Normally hearing children reach nearly the maximum score at about 40 months; for school-age monitoring PEACH too falls short.
IT-MAIS, MAIS and MUSS
What they measure. IT-MAIS and MAIS assess meaningful auditory integration — whether the child brings sound into daily life — across ten items. MUSS addresses the use of speech in communication. All three are administered as a structured parent interview; they are not questionnaires handed over to be filled in.
Who they are for. IT-MAIS for 0–3 years, MAIS for the preschool period and beyond. They are the scales most often encountered in the cochlear implant literature.
What to know. An item-level analysis of IT-MAIS explains why the tool loses its discriminating power so quickly in monitoring. Across fifty-six measurements from twenty-three children with implants, two of the ten items were dropped from the analysis for not fitting the scale structure, and the difficulty range of the items sat clearly below the ability range of the children. In other words, the scale contains no items difficult enough to discriminate a high-performing child. Families were also found not to use the full 0–4 response range, with the middle categories rarely marked. The authors' conclusion is unambiguous: IT-MAIS should not be used on its own in implant candidacy decisions; LittlEARS and PEACH are to be preferred (Barker et al., 2016).
Clinical note. The limitation of IT-MAIS is usually reported as a "ceiling effect"; yet the same analysis found no ceiling effect. The problem is item targeting: the scale does not become hard enough to measure a child who is developing. The distinction is fine but it changes what follows — with a ceiling effect you change the tool, with a targeting problem you keep using it while knowing the range in which it is informative.
CAP, CAP-II and SIR
What they measure. CAP places auditory performance on an eight-category ordering, from being unaware of environmental sounds to holding a telephone conversation with a familiar person. CAP-II adds two steps to that ordering. SIR rates speech intelligibility in five categories.
Their strengths. They are short, easy to administer and well suited to comparison at group level. That is why implant centres use them so often in aggregate outcome reports.
Their limits. They are inadequate for individual monitoring. In early-implanted children the median CAP at 24 months is seven — the ceiling of the scale — and SIR is five, again the ceiling value. Within two years both tools stop discriminating. These scales are also non-linear and low in resolution; moving up one category does not correspond to the same size of gain everywhere on the scale.
There is a further practical problem: families do not read the ordering hierarchically. In an adaptation study with 107 families, 28% of participants answered inconsistently, endorsing both the most basic and the most advanced category at once (Bustos-Rubilar et al., 2022).
SSQ and its child versions
This family of scales, which treats speech, spatial hearing and sound quality as separate domains, is a strong instrument in adults. In the child versions two points need attention.
First, the child form (SSQ-C) cannot be used before about eleven years of age; it was designed for adolescents, not young children.
Second, and more importantly, the parent form (SSQ-P) does not deliver the information expected of it when handed over as a questionnaire. In an analysis with the families of 145 children using bilateral implants, only 45% of questionnaires contained a numerical answer for every item; the rest included at least one "don't know" or "have not observed this". Item-level analysis found 13 of the 23 items inadequate. The authors' recommendation is explicit: SSQ-P should not be used without an interview and a week-long observation period (Killan et al., 2020).
Practice note. This finding bears directly on practice in Türkiye, where SSQ-P is used in most centres in exactly the way the authors criticise: a form filled in while waiting. If the tool is to be used, it should be worked through item by item, sitting with the family.
Tools without norms that are nonetheless useful
Some tools make no claim to be measurement instruments; their purpose is to give shape to the conversation with the family.
ELF (Early Listening Function) tests twelve listening activities at five different distances — from fifteen centimetres to the next room — in infants aged 4 months to 3 years. The resulting chart lets the family see their child's listening distance for themselves. Its own documentation states plainly that it is neither diagnostic nor a screening tool.
CHILD covers home communication situations across fifteen items for ages 3–12 and allows the parent and child versions to be compared; that comparison alone can be instructive.
FLE (Functional Listening Evaluation) is not a scale but a behavioural protocol. The same material is presented under eight conditions (near/far × quiet/noisy × auditory-only/auditory-visual), so the effect of noise, distance and visual support can each be seen separately. It is one of the most telling demonstrations available for classroom layout and remote-microphone decisions.
Clinical note. These three tools should be used for counselling and goal setting, not as outcome measures. A concrete demonstration — "your child does not hear their name from three metres away" — carries a persuasive force no score can match.
How reliable is parent report?
The question comes up unavoidably. The available evidence points to bias in several directions:
- A scale may be measuring a construct other than the one assumed.
- Families do not use the full set of response options and gravitate to the extremes.
- Hierarchical scales are not read hierarchically.
- Situations the family has had no chance to observe are answered by guesswork.
- The one bias that can be quantified is sizeable: families report device use as 2–3 hours a day more than the device's own record.
None of this means the family should not be asked. Parent report belongs alongside objective measurement, not in place of it. When the two sources conflict, the reading should not be that one is wrong but that two different constructs have been measured.
05Where each tool stops discriminating
The figure and table below answer the question that actually arises in practice: for the child in front of you, which tool is still informative?
| Tool | Where it stops discriminating |
|---|---|
| LittlEARS | ~24 months in normal hearing; ~22 months of hearing age in a child with an implant. The expected score at 18–19 months is already 31/35 |
| IT-MAIS | After the first 12–18 months post-implant; cannot be used in candidacy decisions because pre-implant scores cluster at the floor |
| CAP / CAP-II | Ceiling at 24 months in early-implanted children; the top step of CAP-II is not developmentally appropriate for young children |
| SIR | Ceiling at 24 months in early-implanted children |
| PEACH | ~40 months in normal hearing; insufficient for school age |
| SSQ-C | A floor problem rather than a ceiling one: cannot be used before about 11 years |
The sequence that follows from this table is:
LittlEARS 0–24 months → PEACH 2–6 years → PEACH / TEACH · CHILD preschool – primary school → SSQ-P · FLE school age and beyond
That sequence is not arbitrary; it derives from a critical review that judged twelve tools against eleven criteria. Only LittlEARS and PEACH were found to have sufficient normative data and psychometric evidence to be included in the protocol (Bagatto et al., 2011).
06Performance tests
Once the child is mature enough to follow instructions, auditory skill can be measured directly. The governing principle of this section is: the child should be measured on the hardest task they can manage.
The logic of a hierarchical battery
Speech perception tests can be ordered by difficulty. In a large-scale study following 188 children with cochlear implants over two years, the tests were placed into seven steps by empirical difficulty: from functional scales through closed-set word tests to open-set word and sentence tests (Wang et al., 2008). The principle is this: move one step above the step at which the child reaches ceiling, and one step below the step at which the child hits the floor.
The proposed minimum battery for the preschool period contains an important detail: tests are chosen by the child's language age, not chronological age (Uhler et al., 2017). Giving a five-year-old whose language sits at a three-year level the five-year test measures language level, not auditory skill.
When to move up a step
Two different values appear in the literature for moving up a step, and both are correct; they were calculated for different list lengths.
| Accuracy | Decision |
|---|---|
| Above 75–80% | Move up one step |
| 25–79% | Stay at the same step |
| Below 25% | Return to a simpler task (chance is 25% on a four-alternative test) |
| ~85% (on 25-word lists) | Move up one step |
The reason for the difference is statistical. On a 25-word list, the difference between 84% and 96% is not statistically significant (Thornton & Raffin, 1978). The shorter the list, the less information the same score difference carries.
Clinical note. Categorical labels of the form "scored 72%, so performance is moderate" should therefore be avoided. On a 25-word list, 72% and 84% are statistically indistinguishable, yet one is labelled "moderate" and the other "good". The score should always be recorded together with the list length.
Ling 6: what it measures and what it does not
Ling 6 is among the most used and most often misread tools in clinical practice. It has a guide of its own in preparation, so only its position as an assessment tool is covered here.
What it measures. Whether six sounds (/m/, /a/, /u/, /i/, /ʃ/, /s/) are detected at a given distance and level. With the calibrated version, Ling-6(HL), that measurement becomes a threshold value and has been found usable across the 3–18 year range (Glista et al., 2014).
What it does not measure. Comprehension. This is a limit stated plainly in the tool's own documentation. The difference is also numerical: in a study with fifty preschool children, identification thresholds were on average 5–6 dB higher than detection thresholds (Gaikwad et al., 2019). Hearing a sound does not entail recognising it.
| Sound | Detection (dB HL) | Identification (dB HL) |
|---|---|---|
| /a/ | 10.3 | 15.4 |
| /i/ | 11.4 | 19.1 |
| /u/ | 10.4 | 16.8 |
| /m/ | 14.5 | 21.0 |
| /ʃ/ | 15.7 | 21.6 |
| /s/ | 18.2 | 24.0 |
A caveat about a frequently repeated claim. The statement "Ling's six sounds represent the whole speech spectrum" is widespread but weakly supported. In a recent comparison with children using amplification, thresholds for the Ling sounds correlated significantly with warble-tone thresholds at the corresponding frequencies — except for /s/. The assumption of spectral representation weakens precisely where it matters most: in the high frequencies.
Practice note. "The child passed the Ling test" is an incomplete record. Unless the level at which it was passed (detection or identification), the distance and the presentation level are stated, the record cannot be compared with the next session.
Speech in noise and the age norm
A speech-in-noise test belongs in every paediatric battery; but interpreting the result without taking the child's age into account leads almost inevitably to an error. These skills mature late:
| Condition | Reaches adult level |
|---|---|
| Speech in steady noise | 9–10 years |
| With competing talkers in the background | After 13 years |
| Making use of the direction of the sound | 14 years |
(Leibold & Buss, 2019; Corbin et al., 2016; Brown et al., 2010)
Clinical note. An eight-year-old performing below the adult norm in noise is an expected finding, not evidence of an auditory processing disorder. Without an age-specific norm, no performance in noise can be called abnormal.
Holding the test conditions steady
Monitoring rests on comparison, and comparison is only meaningful when the conditions are held constant. Practice in the field is far from that: in a survey of 101 paediatric audiologists, presentation level in quiet ranged from 30 to 65 dB HL, and even the most common value covered only 42% of respondents (Muñoz et al., 2012).
A set of conditions should be chosen and applied consistently. In the minimum battery proposal that set is: 60 dBA in quiet, 65 dBA with a +15 dB signal-to-noise ratio in noise, 50 dBA for soft speech, and recorded stimuli wherever possible. With live-voice presentation, changing the talker effectively changes the test.
07The device side: audibility
Everything covered so far concerns the child's performance. Yet the first question to ask is often a different one: is the sound actually reaching the child?
SII — the audible portion of speech
The Speech Intelligibility Index (SII) summarises how much of the speech signal falls within the child's audible area. Its value comes from being measurable even in children who cannot be tested behaviourally.
On a norm curve built from 161 paediatric ears, aided SII falls from roughly 100% to 40% as hearing level rises from 20 dB HL to 90 dB HL (Bagatto et al., 2011). Practical reference values are: in mild loss an SII of 80% or above is expected at average speech level, and each additional 10 dB HL of loss lowers the value by about 8% (Scollie, 2018).
Why it matters. In a large cohort following more than three hundred children, one child in three was found to have a hearing aid that did not deliver the audibility expected for the degree of loss (McCreery et al., 2015). The assumption "a device is being worn, so the sound is being heard" fails in one child out of every three.
A high SII is not, however, favourable in every case. The index weights the mid frequencies and gives almost no weight to high-frequency audibility; the formula also tends to understate amplified values in severe losses. SII is an indicator to be read alongside loudness targets, not a value to be maximised on its own.
Auditory dosage
A concept proposed in recent years and highly useful for monitoring is auditory dosage: the hours the child spends aided multiplied by aided audibility, plus the hours spent unaided multiplied by unaided audibility (Wiseman & McCreery, 2023).
The value of the concept comes from the two variables combining multiplicatively: if a well-fitted device is worn only two hours a day, the quality of the fitting does not determine the outcome. The converse holds too; a device worn twelve hours that does not deliver adequate audibility will not provide the expected benefit either.
What the data show about hours of use.
- Observed averages are lower than expected. Around 4.5 hours a day in infancy and toddlerhood; below six hours in more than half of children under five.
- In a study following forty children, every child who reached full-time use — 80% of waking hours — by the twenty-fourth month had language scores within normal limits at three years; but only twenty-one of the forty reached that level. The same study found that the age at which full-time use is reached is a stronger predictor than the age at implantation (Park et al., 2019).
- The "at least ten hours a day" criterion often cited in clinic derives from a breakpoint observed in a large cohort; it is not an age-independent biological threshold. It is too strict for a twelve-month-old and too lenient for a school-age child. The age-referenced criterion of "80% of waking hours" is more meaningful.
Practice note. Hours of use should be measured, not asked about. Families report 2–3 hours a day more than the device's own record. With an implant, the time the coil is attached and the time the processor is switched on are also not the same thing; unless the metric is stated, two measurements cannot be compared.
Cortical responses
Where behavioural testing is not possible or the findings conflict, cortical auditory evoked potentials (CAEP) can join the assessment. The latency of the P1 wave is a marker of maturation in the auditory pathways: about 300 milliseconds in the newborn, falling to about 125 ms at three years and about 60 ms in adults.
Its clinical value lies in showing whether a child is within normal limits for their own age group. It has limits too: electrophysiological interpretation requires expertise, implant artefact can corrupt the measurement, and the natural change in waveform between 7 and 11 years can mimic a deprivation pattern. For that reason the inference "P1 is normal, so there is no problem" cannot be drawn.
08Applying the tools by age group
The three snapshots below show how a typical monitoring package is assembled. The tool at the end of the section filters this list automatically according to the age and device status entered.
0–2 years
Behavioural testing options are limited at this stage; the weight falls on the device side and on the family's observations.
- Functional assessment: LittlEARS, at monthly or three-monthly intervals; the direction of the curve is judged, not the score alone.
- Objective assessment: Aided SII at average and soft speech inputs; real-ear measurement (REM) repeated at short intervals as the earmould outgrows its fit.
- Use: Datalogging data at every visit.
- Behavioural assessment: Visual reinforcement audiometry (VRA), from around five months of developmental age.
- Counselling: ELF, so the family can see the listening distance for themselves.
Clinical note. The most commonly missed problem in this age group is the earmould lowering audibility unnoticed as the ear grows. Changing the mould is not only a maintenance step; it is also a reason to measure.
2–6 years
Assessment in the test room becomes possible and the battery diversifies.
- Functional assessment: PEACH; TEACH once institutional education has started.
- Performance: Start with closed-set word tests appropriate to language age, moving to open-set tests as ceilings are reached.
- Ling 6: As a quick screen at the start of every session, recording the detection–identification distinction.
- Language: Regular measurement with age-appropriate Turkish language tests.
- Objective assessment: Aided SII and verification measures; datalogging.
School age
The questions change at this stage; the issue is no longer recognising words in quiet but following the teacher in class.
- Functional assessment: CHILD (parent and child versions together), TEACH; SSQ-P at older ages, through interview.
- Performance: Speech-in-noise tests, with age-specific norms.
- Protocol: FLE, for classroom layout and remote-microphone decisions.
- Academic monitoring: Reading skills and academic attainment; at this age this is where auditory difficulty shows itself most visibly.
Practice note. The most frequent interpretive error at school age is the inference "the audiogram is stable, so there is no problem". At this stage difficulty usually appears not in thresholds but under conditions of noise and distance, neither of which is visible in a standard audiological assessment.
This tool is a decision-support list; it does not replace clinical judgement. For a child with an implant the reference is time with auditory access (hearing age), not chronological age.
09Monitoring: measurement spread over time
A single measurement gives a snapshot. Monitoring shows change, and clinical decisions are drawn from that change.
How often to measure
International guidelines diverge here; some recommend an individualised approach, others define a schedule. The most widely used schedule is: every 4–6 weeks until a full audiogram is obtained; then every three months until three years; and every six months from four years on. With a cochlear implant, re-mapping every three months in the first year is recommended, followed by performance assessment every 6–12 months if progress is adequate.
For developmental monitoring the clearest recommendation is measurement every six months until thirty-six months, and yearly thereafter. The expected level is that the child stays within one standard deviation of their chronological age or cognitive development; when that boundary is crossed, services are adjusted.
A workable rhythm for device monitoring is: initial fitting → month 3 → month 6 → yearly, plus event-triggered visits (mould change, a concern raised by the family, teacher feedback, the period after an illness).
Clinical note. "If the family has a concern, assessment should not be delayed" appears in almost every guideline and is very likely the item most often ignored. A family's concern is by itself a reason for an appointment.
Why monitoring must be routine: progressive loss
The strongest answer to this comes from two population-based studies.
In a study following 207 children diagnosed with mild bilateral loss, the loss progressed in at least one ear in 47% of them; in half of that group the drop exceeded 20 dB. Among cases followed for more than five years the figure rose to 56%. One third of children initially diagnosed with "mild loss in the better ear" showed moderate or greater loss in both ears by the end of follow-up.
The most striking finding is this: none of the risk factors examined predicted progression (Fitzpatrick et al., 2020). A strategy of monitoring only the cases judged to be at risk is therefore not workable; monitoring must be routine, not selective.
The picture in unilateral loss is similar: 47% of 177 children deteriorated, in 12% the loss became bilateral, and most of the deterioration appeared within the first four years (Fitzpatrick et al., 2023). This supports weighting monitoring intensity towards the early years.
Reading progress: the growth slope
Knowing how many points a child gained in six months is not enough on its own. The real question is: is the gap closing, holding steady, or widening?
The measure used for this is the growth slope: the child's gain in language age divided by the calendar time elapsed.
| Slope | What it means |
|---|---|
| = 1.0 | Twelve months of language gain in twelve months. Growth at the same rate as peers; the gap is held steady (gap maintainer) |
| > 1.0 | Faster growth than peers; the gap is closing (gap closer) |
| < 1.0 | Slower growth than peers; the gap is widening (gap opener) |
In a study following eighty-seven children, the slopes of the cochlear implant and hearing aid groups fell between 1.04 and 1.33 — that is, on average a rate at or somewhat above age-equivalent. The group average is misleading, however: the same study found that roughly 80% of the children did not close the gap but held it steady (Yoshinaga-Itano et al., 2010).
Clinical note. This finding provides the most honest ground for the conversation about expectations with a family. The realistic answer to "will my child catch up?" is, for most children: the rate of progress will be similar to that of peers, but the initial gap may not close by itself. A slope falling below 1.0 is the most usable numerical definition of a plateau and is grounds for review.
What counts as a real change
Can a rise in a child's PEACH percentage from 68 to 74 be counted as progress?
Answering that requires knowing the tool's measurement error, and PEACH is the only tool for which this has been calculated: the critical difference at the 95% level is 11 points. A six-point rise therefore cannot be told apart from the natural variability of the measurement.
For the other scales this threshold is unknown. That does not mean they should not be used; it means small score changes should not be reported as progress. Judgement should rest on the trend, not on a single step.
10Red flags
The purpose of monitoring is not to accumulate data but to change the intervention when needed. This section defines what "when needed" means.
Thresholds that call for a review of the intervention. In 182 children using hearing aids, the indicators that best predicted the risk of language delay, and their threshold values, were (Wiseman et al., 2023):
| Indicator | Risk threshold |
|---|---|
| Aided audibility (SII) | Below 0.61 |
| Speech recognition in noise | Below 59% phoneme accuracy |
| Auditory dosage | Below 6.0 |
| Unaided pure-tone average | not predictive |
The last row matters especially: the audiogram alone cannot answer this question.
Referral for cochlear implant evaluation. A recent guideline treats meeting any one of three criteria as grounds for referral (Holder et al., 2026):
- Unaided four-frequency pure-tone average ≥ 60 dB HL
- Aided audibility ≤ 0.60
- Word recognition score ≤ 60% (where obtainable)
In the data behind the guideline, 94% of ears recommended for implantation met at least one criterion. One practical detail matters: only 23% of the children could complete an aided open-set speech test, which is why two of the criteria were kept independent of speech testing.
Clinical note. The referral door should be kept open in both directions. A functional plateau is grounds for referral even when no audiometric criterion is met; and if a criterion is met, referral should follow even when the functional picture looks favourable.
The order to follow when progress is inadequate.
- Evaluate the device first. Is audibility adequate? Are the hours of use adequate? No other explanation should be entertained until these two conditions are satisfied.
- Then review the intervention. Are the goals matched to the child's level? How is the family carrying it through at home?
- Finally, consider additional diagnoses. If audibility and use are adequate and the language delay persists, an additional diagnosis comes into question. If the delay is not isolated and joint attention, play and social reciprocity are also affected, an autism assessment; if all developmental domains lag equally, a cognitive assessment; if speech discrimination is disproportionately poor relative to what the thresholds predict, an investigation for auditory neuropathy spectrum disorder (ANSD).
11Tools used in Türkiye
This section summarises where the tools covered so far stand in Türkiye.
Functional scales and ÇİAT
The functional scales most often encountered in cochlear implant and auditory rehabilitation studies published in Türkiye are LittlEARS, IT-MAIS/MAIS and CAP. Turkish translations of these tools are in clinical use.
Turkish validity studies of second-generation tools have also been carried out in recent years:
- ÇİPDÖ (Children's Auditory Performance Scale; the Turkish version of CHAPS) — ages 7–15, 150 children, internal consistency α = 0.97 (Baydan et al., 2020). Open access.
- FAPCI (functional communication in the preschool period) — ages 2–6; validated with 155 children using implants and 34 controls (Özkan et al., 2020).
- GYİD (Auditory Behaviour in Everyday Life; the Turkish version of ABEL) — Özses et al. (2022).
- C.H.I.L.D Turkish version — Yıldırım Gökay et al. (2024).
ÇİAT (the Auditory Perception Test for Children) is a comprehensive speech perception test developed directly for Turkish. It comprises six categories and seventeen subtests and has been validated with 100 children using cochlear implants and 80 with normal hearing across the 2–15 year range. Internal consistency is high (α = 0.913) (İçöz et al., 2025). Its significance lies in the long-standing call for a perception test built around the phonological structure of Turkish.
Language tests
Auditory monitoring cannot be separated from language monitoring. The main tools used in Türkiye are:
| Tool | Age | Note |
|---|---|---|
| TİFALDİ | 2–12 years | Normed on 3,755 children across 61 provinces; internal consistency 0.99, test–retest 0.97 (Berument & Güven, 2013) |
| TEDİL | 2;0–7;11 years | Has two parallel forms |
| TODİL | 4;0–8;11 years | Comprises nine subtests |
| TPLS-5 | 0–7;11 years | Validated with 1,320 children (Şahlı & Belgin, 2017); useful in early monitoring |
Practice note. Tools such as TEDİL, TODİL and TİGE are paid instruments and most carry an examiner-training requirement. When a monitoring protocol is being built, which tests can actually be obtained should be planned from the outset; a protocol built on a test that cannot be obtained cannot be run.
For the bibliographic details of these tools, see the ODAK scale catalogue.
Assessment in the MEB support education programme
Auditory rehabilitation carried out in special education and rehabilitation centres in Türkiye takes place within the framework of the Support Education Programme for Individuals With Hearing Impairment. The programme's Auditory Training module runs to 200 lesson hours and is divided directly along the Erber levels: fifty hours each for detection, discrimination, identification and auditory comprehension.
The programme defines assessment in four stages, which map onto the cycle covered on this page:
- Coarse assessment — determining which module is needed
- Baseline assessment — establishing the starting level with criterion-referenced tests and checklists
- Assessment during instruction — monitoring progress and charting it
- End-of-year assessment — portfolio analysis
The emphasis on charting at the third stage is not incidental; as the next section shows, putting progress on a chart markedly increases the effect of measurement.
Session structure is set out in regulation as well: individual and group sessions run 60 minutes — 40 minutes of teaching, 10 minutes of rest and preparation, and 10 minutes of feedback to the family. At least eight lesson hours of individual and/or four lesson hours of group education are provided per month.
Practice note. Those ten minutes are an allowance the regulation grants. Family guidance is a defined part of the session, not an extra to be fitted in if time remains.
Educational assessment: RAM and the IEP
Educational assessment and identification is carried out by the special education assessment board within the Guidance and Research Centres (RAM) and is repeated at every transition between types and levels of education; it can also be renewed at the written request of a parent or the school. The Individualised Education Programme prepared in line with the board's decision sets out annual and short-term goals, the type and duration of support services, the methods to be used and the environmental arrangements.
How the measurements covered on this page are carried into IEP goals is the subject of the Auditory Goals in the IEP guide.
12Measuring alone is not enough
Everything covered on this page converges on one final point.
Measurement by itself produces a smaller effect than expected. Two meta-analyses from education research show this plainly.
The first is the classic study of the effect of systematic formative evaluation; the effect size pooled from twenty-one controlled studies was reported as 0.70 (Fuchs & Fuchs, 1986). A more recent and more tightly designed second meta-analysis, across twenty-five studies and 7,379 students, found g = 0.30 (Fuchs et al., 2024); the realistic expectation is therefore more modest.
What is genuinely instructive is the conditions under which the effect grows:
- Frequency of measurement is decisive. Measuring once or twice a week yields meaningful benefit (g ≈ 0.55); less frequent measurement shows no meaningful benefit.
- When practitioners are taught how to act on the data, the effect almost doubles (g = 0.73).
- In the first meta-analysis too, how the data were evaluated and whether they were charted stood out among the moderating variables.
In other words, the power of measurement comes not from the number obtained but from the decision that number is tied to.
Closing. A measurement that does not change the path taken may as well not have been made. For every value measured, a threshold and an action should be defined in advance: what will be done when a given value falls below a given limit. Results should be charted for both practitioner and family; and when a plateau is identified, something should change rather than be waited out.
13Summary
- Assessment answers three separate questions: can the child hear it (audibility), can the child tell it apart (performance), does the child use it (function). No answer substitutes for another; the audiogram alone does not adequately predict language development.
- Most tools are informative over a narrow range: LittlEARS stops discriminating at around 24 months, CAP and SIR at 24 months in early-implanted children, PEACH at around 40 months. The tool is chosen not for the child but for the range the child is in.
- Tests are chosen by language age, not chronological age; the child is measured on the hardest task they can manage.
- Ling 6 measures detection, not comprehension; identification thresholds sit 5–6 dB above detection thresholds.
- Speech-in-noise skills mature late: 9–10 years in steady noise, 13 years with competing talkers, 14 years for using directional information. Without an age-specific norm, no performance in noise can be called abnormal.
- In one child in three, the device does not deliver the audibility expected for the degree of loss.
- Hours of use should be measured, not asked about; families report 2–3 hours a day more.
- Monitoring must be routine, not selective: the loss progresses in roughly half of children diagnosed with mild loss, and no risk factor predicts it.
- Progress should be judged by growth slope rather than a single measurement. Roughly 80% of children hold the gap steady rather than closing it.
- PEACH is the only tool with a calculated measurement error; its critical difference is 11 points.
- The effect of measurement comes from tying it to a decision rule. The threshold should be defined in advance and the data should be charted.
14References
- Bagatto, M. P., Moodie, S. T., Seewald, R. C., Bartlett, D. J. & Scollie, S. D. (2011). A critical review of audiological outcome measures for infants and children. Trends in Amplification, 15(1–2), 23–33.
- Bagatto, M. P., Moodie, S. T., Malandrino, A. C., Richert, F. M., Clench, D. A. & Scollie, S. D. (2011). The University of Western Ontario Pediatric Audiological Monitoring Protocol (UWO PedAMP). Trends in Amplification, 15(1), 57–76.
- Bagatto, M. P. & Scollie, S. D. (2013). Validation of the Parents' Evaluation of Aural/Oral Performance of Children (PEACH) Rating Scale. Journal of the American Academy of Audiology, 24(2), 121–125.
- Barker, B. A., Donovan, N. J., Schubert, A. D. & Walker, E. A. (2016). Using Rasch analysis to examine the item-level psychometrics of the Infant-Toddler Meaningful Auditory Integration Scales. Speech, Language and Hearing, 20(3), 130–143.
- Baydan, M., Aslan, F., Yılmaz, S. & Yalçınkaya, F. (2020). Çocuklar İçin İşitsel Performans Değerlendirme Ölçeği'nin geçerlik ve güvenirlik çalışması. Hacettepe Üniversitesi Sağlık Bilimleri Fakültesi Dergisi, 7(1), 32–40.
- Berument, S. K. & Güven, A. G. (2013). Türkçe İfade Edici ve Alıcı Dil (TİFALDİ) Testi: I. Alıcı dil kelime alt testi standardizasyon ve güvenilirlik-geçerlik çalışması. Türk Psikiyatri Dergisi, 24(3).
- Brown, D. K. et al. (2010). Effects of long-term musical training on cortical auditory evoked potentials. Journal of the American Academy of Audiology, 21, 629–641.
- Bustos-Rubilar, M. et al. (2022). Adaptación y validación del cuestionario CAP-II en población chilena. Revista Chilena de Fonoaudiología, 21, 1–14.
- Ching, T. Y. C. & Hill, M. (2007). The Parents' Evaluation of Aural/Oral Performance of Children (PEACH) scale: Normative data. Journal of the American Academy of Audiology, 18(3), 220–235.
- Coninx, F., Weichbold, V., Tsiakpini, L. et al. (2009). Validation of the LittlEARS Auditory Questionnaire in children with normal hearing. International Journal of Pediatric Otorhinolaryngology, 73(12), 1761–1768.
- Corbin, N. E. et al. (2016). Development of open-set word recognition in children: Speech-shaped noise and two-talker speech maskers. Ear and Hearing, 37, 55–63.
- Fitzpatrick, E. M., Nassrallah, F., Vos, B., Whittingham, J. & Fitzpatrick, J. (2020). Progressive hearing loss in children with mild bilateral hearing loss. Language, Speech, and Hearing Services in Schools.
- Fitzpatrick, E. M. et al. (2023). Trajectory of hearing loss in children with unilateral hearing loss. Frontiers in Pediatrics, 11, 1149477.
- Fuchs, L. S. & Fuchs, D. (1986). Effects of systematic formative evaluation: A meta-analysis. Exceptional Children, 53(3), 199–208.
- Fuchs, A., Radkowitsch, A. & Sommerhoff, D. (2024). Using learning progress monitoring to promote academic performance: A meta-analysis of the effectiveness. Educational Research Review, 46, 100648.
- Gaikwad, S. M., Patil, V. A. & Nandurkar, A. (2019). Performance of normal hearing preschool children on audiometric Ling's six sound test. Journal of Otolaryngology-ENT Research, 11(6), 62–65.
- Glista, D., Scollie, S., Moodie, S. & Easwar, V. (2014). The Ling 6(HL) test: Typical pediatric performance data and clinical use evaluation. Journal of the American Academy of Audiology, 25(10), 1008–1021.
- Hillock-Dunn, A., Taylor, C., Buss, E. & Leibold, L. J. (2015). Assessing speech perception in children with hearing loss: What conventional clinical tools may miss. Ear and Hearing, 36(2), e57–e60.
- Holder, J. T., Warner-Czyz, A. D. & Park, L. R. (2026). Pediatric CI 3-60 Guideline: When to refer children for a cochlear implant candidacy evaluation. Ear and Hearing, 47(5), 1187–1196.
- İçöz, Ö., Işıkhan, S. Y. & Yücel, E. (2025). Development and validation of the Auditory Perception Test for Children. Brain and Behavior, 15(3).
- Killan, C. F., Baxter, P. D. & Killan, E. C. (2020). Face and content validity analysis of the Speech, Spatial and Qualities of Hearing Scale for Parents (SSQ-P). International Journal of Pediatric Otorhinolaryngology, 133, 109964.
- Leibold, L. J. & Buss, E. (2019). Masked speech recognition in school-age children. Frontiers in Psychology, 10, 1981.
- McCreery, R. W., Walker, E. A., Spratford, M. et al. (2015). Longitudinal predictors of aided speech audibility in infants and children. Ear and Hearing, 36, 24S–37S.
- Ministry of National Education (MEB). Support Education Programme for Individuals With Hearing Impairment. Directorate General for Special Education and Guidance Services.
- Muñoz, K., Blaiser, K. & Schofield, H. (2012). Aided speech perception testing practices for three-to-six-year old children with permanent hearing loss. Journal of Educational Audiology, 18, 53–60.
- Özkan, H. B., Aslan, F., Karakaya, J. & Yücel, E. (2020). Okul öncesi dönem çocuklarında işlevsel iletişim becerilerinin değerlendirilmesi. Türkiye Klinikleri Journal of Health Sciences, 5(2), 264–271.
- Özses, M., Özbal Batuk, M., Yılmaz Işıkhan, S. & Çiçek Çınar, B. (2022). Turkish adaptation of the Auditory Behavior in Everyday Life questionnaire. American Journal of Audiology, 31(1), 155–165.
- Park, L. R. et al. (2019). Age at full-time use predicts language outcomes better than age of surgery in children who use cochlear implants. American Journal of Audiology.
- Persson, A., Miniscalco, C., Lohmander, A. & Flynn, T. (2019). Validation of the Swedish version of the LittlEARS Auditory Questionnaire in children with normal hearing — a longitudinal study. International Journal of Audiology, 58, 635–642.
- Scollie, S. (2018). 20Q: Using the aided Speech Intelligibility Index in hearing aid fittings. AudiologyOnline, Article 23707.
- Şahlı, A. S. & Belgin, E. (2017). Adaptation, validity, and reliability of the Preschool Language Scale–Fifth Edition (PLS-5) in the Turkish population. International Journal of Pediatric Otorhinolaryngology.
- Thornton, A. R. & Raffin, M. J. (1978). Speech-discrimination scores modeled as a binomial variable. Journal of Speech and Hearing Research, 21(3), 507–518.
- Uhler, K., Warner-Czyz, A. & Gifford, R. (2017). Pediatric Minimum Speech Test Battery. Journal of the American Academy of Audiology, 28(3), 232–247.
- Visram, A. S., Purdy, S. C., Kelly, J. & Munro, K. J. (2022). Longitudinal changes in LittlEARS scores in infants with hearing aids. International Journal of Audiology, 62(4), 334–342.
- Wang, N.-Y., Eisenberg, L. S., Johnson, K. C. et al. (2008). Tracking development of speech recognition: Longitudinal data from hierarchical assessments in the Childhood Development after Cochlear Implantation study. Otology & Neurotology, 29(2), 240–245.
- Wiseman, K. B. & McCreery, R. W. (2023). Quantifying access to speech in children with hearing loss. Seminars in Hearing, 44(Suppl 1), S17–S28.
- Wiseman, K. B., McCreery, R. W. & Walker, E. A. (2023). Hearing thresholds, speech recognition, and audibility as indicators for modifying intervention in children with hearing aids. Ear and Hearing, 44(4), 787–802.
- Yıldırım Gökay, N., Baysan, C. & Yücel, E. (2024). Turkish adaptation of the Children's Home Inventory for Listening Difficulties. Cochlear Implants International, 25(3).
- Yoshinaga-Itano, C., Baca, R. L. & Sedey, A. L. (2010). Describing the trajectory of language development in the presence of severe-to-profound hearing loss. Otology & Neurotology, 31(8), 1268–1274.
What are the three questions assessment answers?
Tap the card to see the answerCan the child hear it (audibility), can the child tell it apart (performance), does the child use it (function). The three measure different constructs and none substitutes for another.
What does critical difference mean, and for which tool is it known?
Tap the card to see the answerIt is the smallest score change that can be told apart from measurement error. PEACH is the only tool in this field with a published figure: 11 points at the 95% level.
Which age governs the choice of a speech perception test?
Tap the card to see the answerLanguage age, not chronological age. Giving a five-year-old whose language sits at a three-year level the five-year test measures language level, not auditory skill.
Which level does Ling 6 measure?
Tap the card to see the answerDetection. It does not measure comprehension. Identification thresholds sit on average 5–6 dB above detection thresholds, so the level at which the child passed must be recorded.
What does a growth slope of 1.0 mean?
Tap the card to see the answerTwelve months of language gain in twelve months; growth at the same rate as peers. The existing delay is held steady, however, not closed.
How is auditory dosage calculated, and why is it multiplicative?
Tap the card to see the answerIt is aided hours × aided audibility plus unaided hours × unaided audibility. Because it is multiplicative, when one term approaches zero the other cannot rescue the result.
Why should device use be measured rather than asked about?
Tap the card to see the answerFamilies report 2–3 hours a day more than the device's own record.
1How well does unaided pure-tone average predict the risk of language delay?
2At roughly what age does LittlEARS stop discriminating in normally hearing children?
3What can be said about the difference between 84% and 96% on a 25-word list?
4Which risk factor predicts progression in children diagnosed with mild bilateral loss?
5Which condition adds most to the effect of progress monitoring?
İşitmeAtölyesi