Auditory Perception Development
01What is this guide for?
The question asked most often when planning a session is this: "Which level is this child at?" At first glance the question seems reasonable; the Erber hierarchy, the shared language of the field, consists of four levels, and placing a child on one of them makes planning easier.
The question is, however, wrongly framed. Erber's four levels define not where the child is, but how difficult the task is. On the same day, the same child may understand connected speech yet be unable to discriminate a single minimal pair. This is not a contradiction but two tasks of differing difficulty; the level is a property of the task, not of the child.
This guide has been prepared in order to separate three things from one another: how typical auditory perception develops, what the Erber framework actually says and how this development unfolds in a child with hearing loss. The intended gain is not to place a child on a level, but to be able to define the task in front of you and to write the next goal accordingly.
02What is auditory perception, and what development are we tracking?
Auditory perception is the conversion of the acoustic energy that reaches the ear into a meaningful event. This is a different process from measuring thresholds. The audiogram shows whether a sound was heard; perception covers whether that sound can be separated from the others, named, and turned into an instruction.
What we track the development of is not a single skill but a cluster of skills that mature together:
Picking up the presence of sound. Detecting that there is a sound in the environment, turning towards it, locating its source.
Resolving the difference between sounds. First differences of duration, intensity, pitch and timbre; then the fine spectral differences between speech sounds.
Linking what is heard to a representation. Mapping a sound onto an object, a word or a person.
Extracting and using meaning. Carrying out an instruction, answering a question, sustaining a conversation.
Two further components accompany this cluster and are most often overlooked: auditory attention (being able to choose which sound to attend to and which to suppress) and auditory memory (being able to hold what is heard for long enough). Without these two the higher levels cannot be built; as will be seen in the sections that follow, it is in these two areas that children with hearing loss have been reported to fall behind most in the long term.
03Typical development: from before birth to school age
Before birth
The auditory system starts working long before birth. It has been reported that the organ of Corti appears at around the 9th week of gestation and that myelination of the cochlear nerve begins at around week 22 (Moore & Linthicum, 2007). The reliable window for functional responses, however, is weeks 23–30: in a systematic review of eight controlled studies it was stated that the first fetal responses emerge at around week 23 and that consistent responding corresponds to weeks 28–30; that responses to low frequencies (250–500 Hz) were obtained at weeks 25–27, whereas responses to 1000–3000 Hz were obtained only at weeks 29–31 (Movalled et al., 2023).
A common misconception. The statement "hearing begins at week 20" is often repeated, but it is at the optimistic end. The available evidence supports saying that functional hearing begins from around weeks 23–25 and becomes reliable after week 28.
Textbook knowledge about the intrauterine sound environment has also changed recently. The classic description, according to which "sounds arrive attenuated by about 30 dB and cut off above 500 Hz", rested on sheep models. In a recent computational model built on MRI from four pregnancies, it has been shown that attenuation below 1 kHz can be as little as 6 dB, and that above 3 kHz the transmitted pressure at times exceeds the incident pressure because of intrauterine resonance (Gélat et al., 2025). Accordingly, the acoustic environment to which the foetus is exposed is richer than has long been assumed.
Learning takes place in this environment too. It has been shown that newborns alter their sucking behaviour so as to hear their mother's voice (DeCasper & Fifer, 1980). A similar discrimination has been reported in utero as well: foetuses at weeks 33–41 were reported to respond to the mother's voice with heart-rate acceleration and to an unfamiliar voice with deceleration (Kisilevsky et al., 2003). It is stated that the foetus also responds to the father's voice, but that the postnatal preference is specific to the mother (Lee & Kisilevsky, 2014).
Speech perception in infancy
An infant is born equipped to discriminate the sound contrasts of every language in the world and, over the first year, narrows that equipment to the native language. According to Werker and Tees's classic finding (1984), infants aged 6–8 months can discriminate contrasts that do not occur in their native language, whereas by 10–12 months most can no longer do so.
Presenting this perceptual narrowing as a loss is misleading. The two-way picture set out by Kuhl and colleagues (2006) is more useful in the clinic: between 6 and 12 months discrimination of native-language contrasts improves, while discrimination of non-native contrasts declines. The infant does not lose a skill; it specialises.
A note of caution. This timetable is not as sharp as is supposed. In a meta-analysis covering 97 effect sizes and 1,613 infants, it was confirmed that native-language vowel discrimination improved significantly with age, whereas the non-native decline was not statistically significant, and signs of publication bias were detected (Tsuji & Cristia, 2014). A categorical statement such as "the window closes at 10–12 months" should therefore be avoided. Similarly, Kuhl's perceptual magnet effect is a theory; the effect itself has attracted serious methodological criticism (Lotto, Kluender & Holt, 1998).
The most robust mechanism-level finding of this period is statistical learning. It has been shown that eight-month-old infants, after listening to a two-minute artificial syllable stream, can distinguish "words" from part-words by using the transitional probabilities within the stream (Saffran, Aslin & Newport, 1996). Here the amount of language input takes on a technical meaning: the infant can extract the pattern only once it has seen enough samples.
The best-replicated finding in the field, meanwhile, is the preference for infant-directed speech. In a multi-site study conducted with 69 laboratories, 16 countries and 2,329 infants aged 3–15 months, this preference was confirmed; it has been shown that the preference strengthens with age and when speech is presented in the native language (ManyBabies Consortium, 2020). This is the basis for the advice given in family guidance to "speak slowly, at a high pitch and with exaggerated intonation".
The maturation timetable for psychoacoustic skills
Clinically, the most functional information is which skill matures early and which matures late. The table below sets out this timetable in one place.
Foremost among those that mature early is frequency resolution: at 500 and 1000 Hz it is close to adult level at three months and completes its maturation at around six months. By contrast, understanding speech in noise matures very late. It has been reported that children need a signal-to-noise ratio (SNR) about 2.3 dB better than adults in steady noise and that this difference closes only at ages 9–10; and that in two-talker babble the required difference is ~8 dB and persists until after age 13. It is stated that absolute thresholds at low frequencies do not reach adult level until about age 10.
The source of these differences is often not the ear. The interpretation Werner drew from a series of studies applies directly to the clinic: the deficits in the infant are largely a matter of processing efficiency and listening strategy. It has been shown that a six-month-old infant produces brainstem and cortical responses to gaps similar to those of an adult, but listens to a broad band of frequencies all at once — that is, does not narrow attention to the expected frequency region. This factor alone has been reported to account for 2–3 dB of the difference from adults (Werner, 2002; Bargones & Werner, 1994).
Localisation, too, follows a measurable maturation curve: the latency of turning towards a sound source has been reported to be about one second at six months and ~500 ms at three years, with the steepest change occurring before age 1.5 (Eklöf, Asp & Berninger, 2022).
| Skill | Adult level | Clinical note |
|---|---|---|
| Frequency resolution (500–1000 Hz) | ~6 months | Close to adult level at three months; foremost among the early-maturing skills. |
| Auditory scene analysis | from birth | The mechanism is functional, the accuracy low. A six-month-old discriminates a 4% mistuned harmonic; in adults the threshold is 1–2%. |
| Brainstem myelination | 6–12 months | Coding precision continues to improve between ages 3 and 8 (Thompson et al., 2021). |
| Localisation latency | ~age 3 | ~1 second at six months, ~500 ms at three years; the steepest change is before age 1.5. |
| Low-frequency absolute thresholds | ~age 10 | Matures markedly later than the high frequencies. |
| Speech in noise — steady noise | ages 9–10 | The child needs a signal-to-noise ratio about 2.3 dB better than an adult's. |
| Speech in noise — two-talker babble | after age 13 | The required difference is ~8 dB; it is the main source of difficulty in the classroom. |
| Selective (spectral) attention | after age 12 | Children aged 9–12 need a larger frequency separation than adults for stream segregation. |
| Gap detection | up to age 19 | Improvement has been shown to continue between ages 8 and 19 (Gay, Rosen & Huyck, 2020). |
| Cortical axonal maturation | ~age 12 | P1 latency falls to 50–70 ms in the second decade. |
A common over-generalisation. It is often written that gap detection reaches adult level at six years. Yet improvement has been shown to continue between ages 8 and 19 (Gay, Rosen & Huyck, 2020).
Auditory scene analysis and selective attention
Real listening takes place not in a quiet booth but in environments where several sounds overlap. It has been reported that the mechanisms that separate sounds into their sources — that is, auditory scene analysis — are functional from birth, and that the newborn uses the same cues as an adult, only with lower accuracy (Calcus, 2024). It is stated that a six-month-old infant can discriminate a 4% mistuned harmonic, whereas in adults this threshold is 1–2%.
What develops late is not segregation itself but the selectivity of attention. It has been shown that from six months onwards infants use temporal expectancy but not spectral expectancy — that is, they cannot set up a filter of the form "I shall listen to this frequency region". It has been reported that even children aged nine to twelve need a larger frequency separation than adults for auditory stream segregation (Sussman & Steinschneider, 2009).
This is the best answer to the question of why a child who hears well still struggles in the classroom: the ear is adequate, the filter is not yet.
The developmental trajectories covered in this section are shown together on two different timescales in Figure 1.
04Central maturation and the critical period
The different layers of the auditory system reach maturity at different rates. It has been reported that myelination in the cochlear nerve and brainstem pathways approaches adult levels at 6–12 months and that axonal maturation of the auditory cortex continues until about 12 years of age (Moore & Linthicum, 2007). This is the anatomical basis of the principle "brainstem early, cortex late".
The most established way of tracking this maturation clinically is the P1 cortical auditory evoked potential (CAEP). P1 latency has been reported to fall from ~300 ms in infancy to 50–70 ms in the second decade (Sharma et al., 2015), and to be ~105 ms at three years and ~54 ms in adults, shortening by roughly 3 ms a year (Jeon et al., 2025). P1 is therefore interpreted not against a single cut-off value but against age-normed confidence intervals.
What makes P1 genuinely valuable is that it shows maturation to be experience-dependent in nature. Sharma, Dorman and Spahr (2002) reported that children implanted before about 3.5 years of age reached P1 latencies within normal limits within 3–6 months of device use, whereas most of those implanted at 7 years or later did not reach that level; this is the single most cited finding in the field. The 3.5–7 year range, meanwhile, is described as a transition zone in which variability is high (Figure 2).
These findings are the concrete counterpart of the critical (sensitive) period: there is a window during which the central auditory system remains open to stimulation, and once that window has closed the same input no longer produces the same outcome.
Maturation does not, on the other hand, come to an end on a particular date. In a study following 175 children between three and eight years of age, the latency of brainstem responses was shown to go on shortening by 0.018 ms a year while response stability increased (Thompson et al., 2021). The statement "the brainstem completes its maturation at one year of age" is therefore correct in terms of myelination but wrong in terms of coding precision.
05The Erber hierarchy: four levels
The shared language of the field was systematised by Erber (1982) in his book Auditory Training as four levels:
Detection. Being able to perceive the presence — or absence — of sound. The answer to the question "Is there a sound?"
Discrimination. Being able to determine whether two stimuli are the same or different. The child does not need to know what they heard; noticing that the two stimuli are not the same is enough.
Identification. Being able to specify what was heard by naming it, pointing to it or repeating it. At this level the sound is now tied to a representation.
Comprehension. Grasping the meaning of the auditory input; following an instruction, answering a question, sustaining a conversation. It is the highest level, because language knowledge, memory and context all come into play together.
The auditory-only testing principle
Erber's experimental studies in the 1970s form the evidence base of the framework, and its least disputed principle comes from there: auditory-visual perception is superior to auditory-only or visual-only perception (Erber, 1972, 1975). The direct consequence is this: every claim about a child's auditory skill is uninterpretable unless the visual channel has been controlled.
In practice this requires the mouth to be covered with a hand or a small screen and the modality to be stated explicitly in every goal. "Ahmet follows the instruction" is not a goal; "follows it by listening alone" is.
Three common mistakes
Mistaking detection for identification. A child who turns to their name may not have identified it; they may merely have detected that there was a sound. The distinction is made by checking whether they also turn when another word is said in the same setting.
Mistaking discrimination for identification. A child who can say that two stimuli are different may not know which one is which. "Same or different?" and "Which one?" are different tasks.
Mistaking closed-set success for general skill. Being able to choose among four pictures does not predict open-set identification. The step in between is taken up in the next section.
06The matrix: level × stimulus unit
Erber's real contribution was not ordering the four levels but crossing them with a second axis. The stimulus–response matrix is a two-dimensional grid:
One axis holds the four perceptual tasks: detection → discrimination → identification → comprehension. The other holds speech units of increasing linguistic size: speech elements (phonemes, Ling sounds) → syllables → words → phrases → sentences → connected speech.
This yields 24 cells, and every cell corresponds to a definable clinical task. Identification × word: pointing to the spoken word from a closed set. Discrimination × syllable: a same–different judgement for /ba-ba/ and /ba-bi/. Comprehension × connected speech: answering questions about a story that has been listened to. In the matrix in Figure 3, clicking a cell shows that cell's clinical counterpart and the next target.
interactive24 cells, 24 defined tasks. Click a cell; that cell's clinical counterpart and the next target will open below. You can also navigate with the arrow keys.
The clinical function of the matrix is both diagnostic and directive: you mark where performance breaks down, and the next target is written from a neighbouring cell — either to the right (a longer stimulus) or downwards (a harder task). Moving in two directions at once obscures the source of failure.
Difficulty variables
Once the cell has been chosen, the variables that fine-tune difficulty are presented together in Figure 4:
Set structure. Closed set → bridging set (or limited set) → open set. The bridging set is a set the child knows but does not see in front of them (for example, "one of the animals we worked on this week"). This is the step most often skipped in the clinic; it is exactly where success in a closed set fails to transfer to an open set.
Set size. It is increased as success is achieved and reduced in the event of failure.
Contrast salience. From coarse to fine: a difference in syllable count (top / kelebek) → the same syllable count with different phonemes → pairs that share a phoneme → minimal pairs (bal / dal). For consonants the order generally runs manner → place → voicing.
Acoustic conditions. Intensity, signal-to-noise ratio, distance, reverberation.
Context and familiarity. Context makes identification easier; it is reduced step by step.
Visual cue. Moving from auditory-visual presentation to auditory-only presentation.
As a recent addition, De Raeve and colleagues' (2012) Listening Cube model argues that two dimensions are not enough and adds a third axis: material (non-verbal/verbal, including prosody) and listening conditions (quiet/noisy environment, distance, device configuration). In a child with a cochlear implant this third axis is often more decisive than the first two.
07Is the hierarchy really a developmental sequence?
Almost all Turkish sources present the four levels as a validated sequence of developmental steps. The literature does not support this reading.
The assessment in the textbooks. Although the levels are often presented in the form of a developmental hierarchy, in reality they overlap; children with hearing loss do not develop these skills in a strict order. It has been noted that a child can work simultaneously at the phoneme, word, sentence and discourse levels within the same period (Perigoe & Paterson, 2015).
The psychometric finding. Barker and colleagues (2016) applied a Rasch analysis to the IT-MAIS items. The empirically obtained order of item difficulty was reported not to "reflect the order in which children would be expected to develop functional listening skills". A strong rank correlation was obtained only when the items were re-ordered by experts, and two items were removed from the scale because they did not fit. This is the strongest evidence we have that an Erber-aligned instrument does not behave like a unidimensional ladder.
The theoretical framework. Lalonde and Holt (2016) describe the levels as a "hierarchical continuum" — identifying something requires detecting and discriminating it first — but emphasise that the levels are "not entirely discrete from one another".
The structure of the curricula themselves. Among structured curricula, DASL II runs its phonetic listening and auditory comprehension strands in parallel after the sound-awareness stage. In other words, it does not structurally adopt a linear progression.
The position supported by the current data is this: Erber's four levels form a logically nested task-difficulty hierarchy — not a validated sequence of developmental stages.
The nesting is not in dispute — identification already presupposes detection. The developmental order, however, has not been established empirically; the only psychometric study to test it at item level did not confirm the order.
This has three consequences in the clinic. First, rather than placing a child on a single step, mapping where they stand on which tasks is what is required. Secondly, there is no need to wait for a lower level to be "completed" before working at a higher one; the goal of understanding connected speech can also be worked on with a newly implanted child. Thirdly, progress should be reported not as "has skipped a step" but as "can perform this task under these conditions".
08The course of development in a child with hearing loss
Hearing age and catching up
The time that has elapsed from the moment consistent auditory access through a device or implant begins is called the hearing age. Auditory skill inventories are interpreted in relation to hearing age, whereas language norms are interpreted in relation to chronological age. The clinical question is whether the two meet.
The clearest data on this come from the change in the rate of growth after cochlear implantation. In a multicentre study following 188 implanted and 97 normally hearing children, comprehension after implantation was reported to gain 10.4 points per year (95% CI 9.6–11.2), while the modelled pre-implant baseline was 5.4 points (4.1–6.7); for expressive language these values were reported as 8.4 and 5.8 (Niparko et al., 2010). This finding indicates that implantation roughly doubled the rate of growth; however, it was also reported that the same cohort had not reached age-appropriate norms within three years.
Catching up requires not growth at a normal rate but growth faster than normal. Because chronological age advances by one year every year, closing the gap between the two is possible only when the slope is steeper.
Catching up has also been shown to be possible, but conditionally. In a study in which 21 children who received simultaneous bilateral implants between five and eighteen months of age were assessed ten times over six years, it was reported that the overall receptive and expressive language gap closed within the first four years, whereas expressive grammar was still low in the sixth year and the receptive vocabulary gap re-opened between the fourth and the sixth years (Wie et al., 2020). Catching up on a composite score does not mean catching up on every sub-skill.
The dose–response relationship in early intervention
The benefit of early intervention is not a binary "present/absent" matter but a relationship of a dose–response nature, and it grows with the degree of hearing loss. In a study following 350 children with permanent hearing loss, fitting the hearing aid at 24 months rather than at three months was reported to correspond to a loss of 6.8 points at 50 dB HL and 11.8 points at 70 dB HL in the global language score at year five. Activating the cochlear implant at 24 months rather than at six months, in turn, was reported to point to a loss of 21.4 points, that is, of about 1.4 standard deviations (Ching et al., 2017).
In the same direction, in a study in which 403 children were grouped by age at implantation, the language standard scores of those implanted before 12 months were reported to be superior to those of all other groups, and the proportion of children reaching the normative range by school entry was reported to be significantly higher in this group (Dettman et al., 2016).
In children who use hearing aids the corresponding variable is aided audibility. In a large cohort following children with mild to severe hearing loss, it has been shown that the language gap grows with the degree of loss and that audibility determines the rate of growth; it has also been reported that children fitted after 18 months continued to progress in proportion to the amount of device use — the system remains open to experience (Tomblin et al., 2015). At school age, mild and moderate losses were reported to perform similarly to their peers, whereas moderate-to-severe losses were reported to lag significantly behind in spoken language and reading (Tomblin et al., 2020).
A frequently overlooked variable: the actual hours of device use
Early fitting means nothing on its own if the device is not worn regularly. The gap between family report and the device log is striking: families reported an average of 10.63 hours of use per day, whereas the device's own log was reported to show 8.44 hours (Walker et al., 2015).
The picture becomes even clearer in studies that scale the measure to waking hours: the mean percentage of daily listening hours has been reported as 63% (range 18–117%), and this single measure was reported to explain roughly 40% of the variance in receptive language in the first year (Gagnon, Eskridge & Brown, 2020).
The clinical implication. Interpreting progress without reading the datalogging is like judging the effect of a treatment without knowing whether the medicine has been taken regularly.
On the input side, too, quantity alone is not enough. It has been reported that it was not the quantity but the conversation-initiating quality of the language input provided by the caregiver at 18 months that predicted language at three years (Ambrose et al., 2015).
09Late-emerging skills and the ceiling effect
After the first two or three years, functional scales become saturated. A ceiling effect appears in instruments such as IT-MAIS, CAP and LittlEARS, and these instruments stop showing the differences between children. This does not mean that development has stopped; reaching the limit of what we are measuring is all it means.
The skills that need to be monitored beyond this point are the following:
Understanding speech in noise. This is not only descriptive but also predictive: speech recognition in noise in the preschool period has been reported to predict language at nine years significantly, even after early language level has been statistically controlled for (Ching et al., 2022).
Spatial hearing. In bilaterally implanted adolescents, spatial hearing has been reported to display a pattern that is not normal but "still developing": good thresholds in quiet yet marked deterioration in noise, more errors in localising a stationary sound and in detecting the direction of movement, and low sensitivity to the interaural time difference (ITD) (Alemu et al., 2025).
Spectral resolution. This skill, which improves until adolescence in normally hearing children, has been reported to appear unrelated to chronological age, age at implantation or implant experience in implant users (Landsberger et al., 2018; Jahn et al., 2022). Because the samples are small, this finding should be read as a trend rather than as a definitive conclusion.
Auditory working memory. In school-age implant users, the deficit has been reported to appear only in the auditory modality; no deficit was found in the visual modality, and the ability to benefit from pattern regularity was reported to be preserved (Pesnot Lerousseau et al., 2025). On this account, the problem is not a general memory weakness but a bottleneck in auditory and lexical encoding.
Its relationship with literacy. In 47 implant users, phonological processing has been reported to be markedly low (z = −0.95 for words, z = −1.90 for non-words), whereas reading comprehension was reported to remain almost within normal limits (z = −0.20). In the regression analysis it was reported that the variable predicting reading comprehension was spoken language and that phonological processing was not a significant predictor (Camarata et al., 2026). The practical conclusion that follows is this: phonological practice is a means, not an end; the target should be spoken language.
A caution about normative curves. MAIS/IT-MAIS/CAP growth curves arranged by month are frequently shared; however, there is no large-sample, methodologically sound normative curve in the field. The tables in circulation are mostly uncontrolled single-centre series. For this reason it is more appropriate to use these tables not as norms but as your own centre's range of expectation.
10Assessment: how do we measure the level?
Behavioural audiometry measures detection only
This distinction is constantly confused in the clinic. Behavioural observation audiometry (BOA) is used below a developmental age of about six months, and the relevant guideline states explicitly that this method is not suitable for screening, threshold estimation or verification of amplification. Visual reinforcement audiometry (VRA) is for approximately 5–24 months, and conditioned play audiometry (CPA) for approximately 2–5 years (AAA, 2020). The VRA guideline emphasises an important point: every clear head turn towards the signal should be reinforced, because lateralisation ability is not being tested in this test (BSA, 2025).
In the end, all three of these methods answer the question "can the child hear?". The question "what does the child hear, and how much of it does the child understand?" can be answered only with level-based speech perception tasks and with functional scales.
Ling 6
The six sounds sample the speech spectrum: /m/ ~50–350 Hz; /u/ first formant 250–500, second formant 700–1200 Hz; /a/ 500–700 and ~1000–1400 Hz; /i/ 200–400 and 2300–3500 Hz; /ʃ/ above 2–4 kHz; /s/ above 4–8 kHz.
The critical nuance is this: being able to repeat a sound is significantly harder than being able to detect it — copying a sound requires hearing it at least ~10 dB above threshold. For this reason the two versions of the test do not measure the same thing; which one is used must be written explicitly in the goal.
The live-voice version depends on the speaker, the distance and the ambient noise, and it is not calibrated. It is therefore not a diagnostic threshold tool but a daily device check and functional verification tool. The calibrated version, Ling-6(HL), presents pre-recorded stimuli in dB HL and can be administered in the sound field with VRA or CPA (Scollie & Glista, 2012).
An instrument map by level
Closed-set pattern perception tasks measure the discrimination level; picture-choice batteries such as ESP and MTP measure pattern and word identification; GASP covers three levels, from detection to sentence comprehension, in a single battery; PBK, LNT and MLNT measure open-set identification; and HINT-C and BKB-SIN measure identification and comprehension in noise.
Functional scales complete this picture: LittlEARS (0–24 months, chronological or hearing age), IT-MAIS/MAIS, MUSS, CAP and its reduced-ceiling version CAP-II, SIR, PEACH/TEACH, ABEL, CHAPS and ELF. Which level each instrument measures is summarised in Table 2.
| Instrument | Level measured | Description |
|---|---|---|
| Pattern perception tasks (closed set) | Discrimination | Based on differences in syllable number and stress pattern; does not require word recognition. |
| ESP · MTP | Identification | Picture-choice closed-set batteries; measure pattern and word identification together. |
| GASP | Detection → Comprehension | Brings detection, word identification and sentence comprehension together in a single battery. |
| PBK · LNT · MLNT | Identification (open set) | Open-set word identification; word frequency and neighbourhood density are controlled. |
| HINT-C · BKB-SIN | Comprehension (in noise) | Sentence-level identification and comprehension in noise; yields an SNR threshold. |
| LittlEARS | Functional (0–24 months) | Interpreted by chronological or hearing age; sensitive in the early period. |
| IT-MAIS · MAIS | Functional | A listening behaviour index based on family report. It is not evidence of an Erber step. |
| MUSS | Functional (spoken communication) | Assesses the use of spoken communication; read together with language goals. |
| CAP · CAP-II | Functional | An ordinal scale with eight categories (extended in CAP-II); the ceiling effect is reduced. |
| SIR | Functional (speech intelligibility) | Rates how intelligible the child's speech is to an unfamiliar listener. |
| PEACH · TEACH | Functional (everyday settings) | Functional performance in home and classroom settings under quiet and noisy conditions. |
| ABEL · CHAPS · ELF | Functional (behaviour / listening) | Auditory behaviour in everyday life, listening difficulty in the classroom and early listening function. |
A point to watch. In the item analysis applied to IT-MAIS, the two items removed from the scale for misfit were reported to be the non-perceptual item and the single comprehension item (Barker et al., 2016). IT-MAIS is a functional listening index; it cannot be used as evidence of where a child stands on the Erber steps.
Turkish instruments
The strongest Turkish-developed instrument for assessing auditory perception by developmental level in Türkiye is ÇİAT (Children's Auditory Perception Test). The revised version is stated to contain 17 subtests in six hierarchical categories: phoneme detection; pattern identification; closed-set speech recognition; auditory-visual integration; modified open-set identification; open-set identification. Cronbach's α has been reported as 0.913 and predictive validity as R² = 0.78 (İçöz, 2022; İçöz et al., 2025).
Another instrument, aimed at the recognition of speech sounds, is KSTT, which offers a 58-item screening based on 29 Turkish speech sounds; in the 6;5–6;11 age range the area under the ROC curve has been reported as 0.923 and the cut-off score as 48.5 (Küçükünal & Yücel).
Among the scales whose Turkish adaptations have been published are ABEL (Özses et al., 2022), CHAPS (Baydan et al., 2020), PEACH/EÇİPED (Eroğlu et al., 2021), C.H.I.L.D. (Yıldırım Gökay et al., 2024) and FAPCI (Özkan et al., 2020).
Alongside these, instruments such as LittlEARS, IT-MAIS/MAIS and MUSS, CAP and CAP-II, and Ling 6 are also widely used in Türkiye. Table 3 shows these instruments together.
| Instrument | Type | Scope and psychometric data |
|---|---|---|
| ÇİAT — Children's Auditory Perception Test | Turkish-developed | 17 subtests in six hierarchical categories: phoneme detection, pattern identification, closed-set identification, auditory-visual integration, modified open-set and open-set identification. Cronbach's α = 0.913; predictive validity R² = 0.78 (İçöz, 2022; İçöz et al., 2025). |
| KSTT — Speech Sound Recognition Test | Turkish-developed | 58 items over 29 Turkish speech sounds. In the 6;5–6;11 age range the area under the ROC curve is 0.923; cut-off score 48.5 (Küçükünal & Yücel). |
| ABEL | Adaptation | Auditory behaviour in everyday life; its Turkish validity and reliability have been published (Özses et al., 2022). |
| CHAPS | Adaptation | Screening for classroom listening difficulty by teacher report (Baydan et al., 2020). |
| PEACH / EÇİPED | Adaptation | Assessment of functional performance in everyday settings by family report (Eroğlu et al., 2021). |
| C.H.I.L.D. | Adaptation | Assessment of home and school listening conditions from the child's and the family's perspective (Yıldırım Gökay et al., 2024). |
| FAPCI | Adaptation | Functional spoken communication performance in young children (Özkan et al., 2020). |
| LittlEARS · IT-MAIS / MAIS · MUSS · CAP / CAP-II · Ling 6 | In common use | These instruments are widely used in Türkiye. |
Mapping onto the MEB framework
In the Ministry of National Education's support education programme, the Auditory Training module consists of four units, each of 50 lesson hours (MEB, 2021): Sound Detection, Sound Discrimination, Sound Identification, Auditory Comprehension. The structure maps one-to-one onto the Erber hierarchy (Table 4). For an audiologist working in Türkiye this is a considerable convenience: the official framework of the institution and the language of the international literature are the same.
| MEB Auditory Training unit | Duration | Erber equivalent and content |
|---|---|---|
| Sound Detection | 50 lesson hours | Detection. Perceiving the presence or absence of sound; establishing the conditioned response. |
| Sound Discrimination | 50 lesson hours | Discrimination. Deciding whether two stimuli are the same or different; from gross to fine contrasts. |
| Sound Identification | 50 lesson hours | Identification. Pointing to, naming or repeating what is heard; from closed set to open set. |
| Auditory Comprehension | 50 lesson hours | Comprehension. Carrying out an instruction, answering a question, sustaining a conversation. |
11From level to goal
A goal is the intersection of three independent parameters (Garber & Nevins, 2012):
1. Auditory function — detection, discrimination, identification, comprehension.
2. Stimulus unit — sounds, words, phrases, sentences, conversation.
3. Situational context — structured task → closed set → bridging set → open set → routine activity → natural interaction.
Only one parameter is advanced at a time. When two parameters are made harder together, the source of failure becomes unclear.
As a fourth axis, the number of critical elements can be added: "take the car" carries one, "take the car, put it in the box" two, and "take the red car, put it in the blue box" three critical elements. This axis indexes auditory memory rather than acoustic difficulty.
A properly written goal: From a closed set of four items, in a quiet room, at a distance of one metre, by listening alone, Ahmet points to the target word on 8 of 10 trials; over three consecutive sessions.
Examples of poorly written goals: "Listening skills will improve." (No level, unit, context or criterion is specified.) "Will understand what is said with 80% accuracy." (The level is specified; the unit, the set and the acoustic condition are missing.)
On the "80% accuracy over three consecutive sessions" criterion
This criterion is widely used in Türkiye as well, and it is often written down without being thought through. The criticism levelled at it is this: the criterion has turned into a habit and may not be logically compatible with the target behaviour (Diehm, 2017).
Percentage accuracy is appropriate for tasks with countable trials — "responds to their name in noise at three metres" is a measurable goal. For goals such as "takes part in conversation", "uses a repair strategy" or "follows an instruction in class", by contrast, a percentage offers no meaningful criterion; for goals of this kind, frequency, duration, consistency or a rubric are the more appropriate measures.
In children with additional disabilities this is even more pronounced: it has been reported that at 12 months after implantation the median CAP score rose only from 0 to 2, and that in two children it did not change at all (Martínez-Pantanalli & Bravo-Torres, 2025). In such a child, the criterion needs to be shifted from percentage accuracy towards the frequency and consistency of the response instead.
How well does auditory training work?
Under this heading the limits of the evidence must be stated plainly. A systematic review of auditory training in paediatric implant users established that all of the studies reported improvement on the trained tasks but that transfer occurred only in part, and that the quality of the evidence remained low; quality of life was not measured in any of the studies (Rayes, Al-Malky & Vickers, 2019). In a review of computer-based training, only six of 920 records were reported to have met the inclusion criteria (Silva et al., 2023).
The strongest paediatric randomised data in the field come from a study conducted with 99 school-age children: it has been reported that 16 hours of gamified training produced gains in both modalities and that transfer to an untrained talker did occur, but remained smaller than for the trained one (Tye-Murray et al., 2022). In another study, children who received no additional training were reported to have retained their gains in full at 4–6 weeks, with booster training making only a limited contribution (Spehar et al., 2024).
Adult data complete the picture: it has been reported that a gain of ~14.8 points in vowel and consonant identification was retained after roughly 10 hours of computer-assisted home training, whereas no change at all was found on the self-report scale (Kerneis et al., 2023).
These findings can be summarised in a single sentence: auditory training reliably improves the task that is trained; transfer to an untrained talker is smaller; transfer to real-life self-report has most often not been demonstrated. This picture is not a reason to give up on training, but a reason to set the goal close to real life from the outset.
One final reminder: in the classroom, remote microphone systems are an auditory access intervention with far stronger evidence behind them than any training curriculum. They are not an alternative to working through the hierarchy but a precondition for it.
12Summary
Auditory perception is meaning, not a threshold. The audiogram shows that a sound has been heard; perception reveals that the sound can be separated out, named and turned into an instruction.
In typical development some skills reach maturation very early and others very late. Frequency resolution is at adult level by six months; speech understanding in noise takes until 9–10 years of age, and in two-talker babble until after 13 years. For the child who struggles in class, the problem is often not the ear but the filter of attention.
Central maturation is experience-dependent. P1 latency makes this measurable; it has been reported that auditory access provided before 3.5 years of age allows maturation within normal limits, whereas access after 7 years of age largely does not.
Erber's four levels define task difficulty, not the child's place. The logical nesting is not in dispute; that it forms a developmental sequence of stages, however, has not been confirmed. Rather than "which step is this child on?", the question to ask is "which cell is this task in?".
The matrix dictates the goal. Level × stimulus unit × context; only one parameter is advanced at a time. The most frequently skipped step is the bridging set.
Catching up requires faster-than-normal growth. It has been reported that implantation roughly doubles the rate of growth, but that this alone does not bring children up to the norms; the benefit of early intervention increases with the degree of loss.
If the device is not worn consistently, no goal is valid. Family report has been shown to overestimate wear time consistently; progress should not be interpreted without reading the datalogging.
After the first two to three years, a ceiling effect appears on functional scales. From this point on, monitoring is carried forward through speech understanding in noise, spatial hearing, auditory memory and literacy.
Locally developed tools are available in Türkiye. ÇİAT and KSTT have published psychometric data; tools such as LittlEARS, IT-MAIS, CAP and the Ling 6 are also in widespread use.
Auditory training works, but its effect has limits. The trained task improves; transfer shrinks; transfer to self-report has most often not been demonstrated. This is why the goal must be set close to real life from the outset.
13References
Alemu, R. Z., Blakeman, A., Fung, A. L., Hazen, M., Negandhi, J., Papsin, B. C., Cushing, S. L., & Gordon, K. A. (2025). Children with bilateral cochlear implants show emerging spatial hearing of stationary and moving sound. Trends in Hearing, 29.
Ambrose, S. E., Walker, E. A., Unflat-Berry, L. M., Oleson, J. J., & Moeller, M. P. (2015). Quantity and quality of caregivers’ linguistic input to 18-month and 3-year-old children who are hard of hearing. Ear and Hearing, 36(Suppl. 1), 48S–59S.
American Academy of Audiology. (2020). Clinical guidance document: Assessment of hearing in infants and young children.
Bargones, J. Y., & Werner, L. A. (1994). Adults listen selectively; infants do not. Psychological Science, 5(3), 170–174.
Barker, B. A., Donovan, N. J., Schubert, A. D., & Walker, E. A. (2016). Using Rasch analysis to examine the item-level psychometrics of the Infant-Toddler Meaningful Auditory Integration Scales. Speech, Language and Hearing, 20(3), 130–143.
Baydan, M., Aslan, F., Yılmaz, S., & Yalçınkaya, F. (2020). Çocukların İşitsel Performans Ölçeği’nin Türkçe geçerlik ve güvenirliği. Hacettepe Üniversitesi Sağlık Bilimleri Fakültesi Dergisi, 7(1), 32–40.
British Society of Audiology. (2025). Recommended procedure: Visual reinforcement audiometry (v1.2).
Calcus, A. (2024). Development of auditory scene analysis: A mini-review. Frontiers in Human Neuroscience, 18, 1352247.
Camarata, S., Lighterink, M., Sunderhaus, L., Labadie, R., & Gifford, R. (2026). Phonological processing, oral language abilities, and reading comprehension in children with cochlear implants. Scientific Reports, 15, 45800.
Ching, T. Y. C., Cupples, L., & Zhang, V. W. (2022). Predicting 9-year language ability from preschool speech recognition in noise in children using cochlear implants. Trends in Hearing, 26.
Ching, T. Y. C., Dillon, H., Button, L., Seeto, M., Van Buynder, P., Marnane, V., Cupples, L., & Leigh, G. (2017). Age at intervention for permanent hearing loss and 5-year language outcomes. Pediatrics, 140(3), e20164274.
De Raeve, L., Baerts, J., Colleye, E., & Croux, E. (2012). The Listening Cube: A three-dimensional auditory training program. Clinical and Experimental Otorhinolaryngology, 5(Suppl. 1), S1–S4.
DeCasper, A. J., & Fifer, W. P. (1980). Of human bonding: Newborns prefer their mothers’ voices. Science, 208(4448), 1174–1176.
Dettman, S. J., Dowell, R. C., Choo, D., ve ark. (2016). Long-term communication outcomes for children receiving cochlear implants younger than 12 months: A multicenter study. Otology & Neurotology, 37(2), e82–e95.
Diehm, E. (2017). Writing measurable and academically relevant IEP goals with 80% accuracy over three consecutive trials. Perspectives of the ASHA Special Interest Groups, 2(16), 34–44.
Eklöf, M., Asp, F., & Berninger, E. (2022). The development of sound localization latency in infants and young children with normal hearing. Trends in Hearing, 26.
Erber, N. P. (1972). Auditory, visual, and auditory-visual recognition of consonants by children with normal and impaired hearing. Journal of Speech and Hearing Research, 15(2), 413–422.
Erber, N. P. (1975). Auditory-visual perception of speech. Journal of Speech and Hearing Disorders, 40(4), 481–492.
Erber, N. P. (1982). Auditory training. Alexander Graham Bell Association for the Deaf.
Eroğlu, K., Şahin Kamışlı, G. İ., Altınyay, Ş., Gökdoğan, Ç., Bayramoğlu, İ., & Kemaloğlu, Y. K. (2021). Validation of the Turkish version of the PEACH rating scale. ENT Updates, 11(3).
Gagnon, E. B., Eskridge, H., & Brown, K. D. (2020). Pediatric cochlear implant wear time and early language development. Cochlear Implants International, 21(2), 92–97.
Garber, A., & Nevins, M. E. (2012). Getting started with auditory skills. AudiologyOnline, Article 7034.
Gay, J. D., Rosen, M. J., & Huyck, J. J. (2020). Effects of gap position on perceptual gap detection across late childhood and adolescence. JARO, 21(3), 243–258.
Gélat, P., van ’t Wout, E., Haqshenas, R., ve ark. (2025). Evaluation of fetal exposure to environmental noise using a computer-generated model. Nature Communications, 16, 3916.
İçöz, Ö. (2022). Çocuklar için İşitsel Algı Testi geçerlik ve güvenirlik çalışması [Doktora tezi]. Hacettepe Üniversitesi Sağlık Bilimleri Enstitüsü.
İçöz, Ö., & Yücel, E. (2024). İşitme kayıplı çocuklarda işitsel algı gelişimi ve değerlendirilmesi. Türk Odyoloji ve İşitme Araştırmaları Dergisi, 7(2), 21–26.
İçöz, Ö., ve ark. (2025). Psychometric properties of the Children’s Auditory Perception Test: Reliability and validity analysis. Brain and Behavior.
Jahn, K. N., Arenberg, J. G., & Horn, D. L. (2022). Spectral resolution development in children with normal hearing and with cochlear implants: A review of behavioral studies. Journal of Speech, Language, and Hearing Research, 65(4).
Jeon, E. K., Brown, C., Abbas, P., & Gantz, B. (2025). The effect of development on cortical auditory evoked potentials in normal hearing listeners and cochlear implant users. Frontiers in Human Neuroscience, 19, 1473365.
Kerneis, S., Galvin, J. J., Borel, S., Baqué, J., Fu, Q.-J., & Bakhos, D. (2023). Preliminary evaluation of computer-assisted home training for French cochlear implant recipients. PLOS ONE, 18(4), e0285154.
Kisilevsky, B. S., Hains, S. M. J., Lee, K., ve ark. (2003). Effects of experience on fetal voice recognition. Psychological Science, 14(3), 220–224.
Kuhl, P. K., Stevens, E., Hayashi, A., Deguchi, T., Kiritani, S., & Iverson, P. (2006). Infants show a facilitation effect for native language phonetic perception between 6 and 12 months. Developmental Science, 9(2), F13–F21.
Küçükünal, I. S., & Yücel, E. Konuşma seslerini tanıma testi Türkçe geçerlik güvenirlik çalışması. Türk Odyoloji ve İşitme Araştırmaları Dergisi, 6(2), 51–56.
Lalonde, K., & Holt, R. F. (2016). Audiovisual speech perception development at varying levels of perceptual processing. The Journal of the Acoustical Society of America, 139(4), 1713–1723.
Landsberger, D. M., Padilla, M., Martinez, A. S., & Eisenberg, L. S. (2018). Spectral-temporal modulated ripple discrimination by children with cochlear implants. Ear and Hearing, 39(1), 60–68.
Lee, G. Y., & Kisilevsky, B. S. (2014). Fetuses respond to father’s voice but prefer mother’s voice after birth. Developmental Psychobiology, 56(1), 1–11.
Lotto, A. J., Kluender, K. R., & Holt, L. L. (1998). Depolarizing the perceptual magnet effect. The Journal of the Acoustical Society of America, 103(6), 3648–3655.
ManyBabies Consortium. (2020). Quantifying sources of variability in infancy research using the infant-directed-speech preference. Advances in Methods and Practices in Psychological Science, 3(1), 24–52.
Martínez-Pantanalli, C., & Bravo-Torres, S. (2025). Auditory perception outcomes in children with deafness and additional disabilities 12 months after cochlear implant activation. Audiology Research, 15(3), 47.
Millî Eğitim Bakanlığı. (2021). İşitme yetersizliği olan bireyler için destek eğitim programı. Özel Öğretim Kurumları Genel Müdürlüğü.
Moore, J. K., & Linthicum, F. H. (2007). The human auditory system: A timeline of development. International Journal of Audiology, 46(9), 460–478.
Movalled, K., Sani, A., Nikniaz, L., ve ark. (2023). The impact of sound stimulations during pregnancy on fetal learning: A systematic review. BMC Pediatrics, 23, 183.
Niparko, J. K., Tobey, E. A., Thal, D. J., ve ark. (2010). Spoken language development in children following cochlear implantation. JAMA, 303(15), 1498–1506.
Özkan, S., Aslan, F., Karakaya, J., & Yücel, E. (2020). FAPCI ölçeğinin Türkçe geçerlik ve güvenirliği. Türkiye Klinikleri Sağlık Bilimleri Dergisi, 5(2), 264–271.
Özses, M., Özbal Batuk, M., Yılmaz Işıkhan, S., & Çiçek Çınar, B. (2022). Validity and reliability of Turkish version of the Auditory Behavior in Everyday Life questionnaire. American Journal of Audiology, 31(1), 155–165.
Perigoe, C. B., & Paterson, M. M. (2015). Understanding auditory development and the child with hearing loss. In D. R. Welling & C. A. Ukstins (Eds.), Fundamentals of audiology for the speech-language pathologist. Jones & Bartlett.
Pesnot Lerousseau, J., Denis, M., Roman, S., & Schön, D. (2025). Working memory deficits in school-age children with cochlear implants. Journal of Speech, Language, and Hearing Research.
Rayes, H., Al-Malky, G., & Vickers, D. (2019). Systematic review of auditory training in pediatric cochlear implant recipients. Journal of Speech, Language, and Hearing Research, 62(5), 1574–1593.
Saffran, J. R., Aslin, R. N., & Newport, E. L. (1996). Statistical learning by 8-month-old infants. Science, 274(5294), 1926–1928.
Scollie, S., & Glista, D. (2012). The Ling-6(HL): Instructions. Western University.
Sharma, A., Dorman, M. F., & Spahr, A. J. (2002). A sensitive period for the development of the central auditory system in children with cochlear implants. Ear and Hearing, 23(6), 532–539.
Sharma, A., Glick, H., Deeves, E., & Duncan, E. (2015). The P1 biomarker for assessing cortical maturation in pediatric hearing loss: A review. Otorinolaringologia, 65(4), 103–114.
Silva, J. de M., Sordi Silva, B. C., Frederigue Lopes, N. B., Jacob, R. T. de S., & Moret, A. L. M. (2023). Effectiveness of computerized auditory training on speech perception in children with hearing loss: A systematic review. American Journal of Audiology, 32(4), 990–1004.
Spehar, B., Murray, N., Mauzé, E., Sommers, M., & Barcroft, J. (2024). Speech perception training in children: The retention of benefits and booster training. Ear and Hearing, 45(1), 164–173.
Sussman, E. S., & Steinschneider, M. (2009). Attention effects on auditory scene analysis in children. Neuropsychologia, 47(3), 771–785.
Thompson, E. C., ve ark. (2021). Auditory neurophysiological development in early childhood: A growth curve modeling approach. Clinical Neurophysiology, 132(9), 2110–2122.
Tomblin, J. B., Harrison, M., Ambrose, S. E., Walker, E. A., Oleson, J. J., & Moeller, M. P. (2015). Language outcomes in young children with mild to severe hearing loss. Ear and Hearing, 36(Suppl. 1), 76S–91S.
Tomblin, J. B., Oleson, J., Ambrose, S. E., Walker, E. A., McCreery, R. W., & Moeller, M. P. (2020). Aided hearing moderates the academic outcomes of children with mild to severe hearing loss. Ear and Hearing, 41(4), 775–789.
Tsuji, S., & Cristia, A. (2014). Perceptual attunement in vowels: A meta-analysis. Developmental Psychobiology, 56(2), 179–191.
Tye-Murray, N., Spehar, B., Sommers, M., Mauzé, E., Barcroft, J., & Grantham, H. (2022). Teaching children with hearing loss to recognize speech. Ear and Hearing, 43(1), 181–191.
Walker, E. A., McCreery, R. W., Spratford, M., ve ark. (2015). Trends and predictors of longitudinal hearing aid use for children who are hard of hearing. Ear and Hearing, 36(Suppl. 1), 38S–47S.
Werker, J. F., & Tees, R. C. (1984). Cross-language speech perception: Evidence for perceptual reorganization during the first year of life. Infant Behavior and Development, 7(1), 49–63.
Werner, L. A. (2002). Infant auditory capabilities. University of Washington.
Wie, O. B., Torkildsen, J. V. K., Schauber, S., Busch, T., & Litovsky, R. (2020). Long-term language development in children with early simultaneous bilateral cochlear implants. Ear and Hearing, 41(5), 1294–1305.
Yıldırım Gökay, N., Baysan, C., & Yücel, E. (2024). The validity and reliability of Turkish version of the C.H.I.L.D. questionnaire. Cochlear Implants International.
What do Erber's four levels define?
Tap the card to see the answerThe difficulty of the task — not the step the child is on. On the same day, the same child may understand connected speech and yet be unable to discriminate a minimal pair.
What are the two axes of the stimulus–response matrix?
Tap the card to see the answerThe four levels of perception (detection → discrimination → identification → comprehension) and the stimulus unit in increasing linguistic size (element → syllable → word → phrase → sentence → connected speech). 24 cells, 24 defined tasks.
Which is the skipped step between the closed set and the open set?
Tap the card to see the answerThe bridging set: a set the child knows but does not see in front of them — for example, “one of the animals we worked on this week”. This is precisely where success in the closed set fails to transfer.
Why is P1 latency clinically valuable?
Tap the card to see the answerIt shows that cortical maturation is experience-dependent: it has been reported that children implanted before ~3.5 years of age reach normal limits within 3–6 months, whereas most of those implanted at 7 years of age or later do not.
Which psychoacoustic skill reaches maturation early, and which one late?
Tap the card to see the answerFrequency resolution is at adult level by ~6 months. Speech understanding in noise, by contrast, takes until 9–10 years of age, and in two-talker babble until after 13 years.
What does “catching up” mean and why is it difficult?
Tap the card to see the answerBecause chronological age advances by one year every year, closing the gap requires faster-than-normal growth. Implantation has been reported to roughly double the rate of growth; on its own, however, this does not bring a child up to the norms.
Which data should be looked at before progress is interpreted?
Tap the card to see the answerThe device's datalogging record. It has been reported that families report ~10.6 hours a day while the record shows ~8.4 hours, and that the percentage of listening hours explains ~40% of the variance in receptive language in the first year.
1What do Erber's four levels tell you about?
2How is the next goal chosen in the stimulus–response matrix?
3When does speech understanding in noise reach adult level in typical development?
4Which is the intermediate step between the closed set and the open set?
5What kind of tool is the live-voice version of the Ling 6?
İşitmeAtölyesi