celpipmocks
Speaking · Lesson 2 of 28

The Four Speaking Criteria

Four criteria, each scored from 1 to 12, each carrying exactly the same weight. Knowing which one a mistake lands on is what turns vague practice into targeted practice.

Four criteria, equal weight, one average

CELPIP Speaking is assessed on four criteria: Content and Coherence, Vocabulary, Listenability, and Task Fulfilment. Each one is scored on the same 1 to 12 scale, and your overall speaking level is the rounded average of the four. No criterion counts for more than another.

That arithmetic has a consequence worth sitting with. A candidate at 10, 10, 10 and 6 averages to 9. A candidate at 8 across the board also averages to 8, and is two criteria away from the same place with far less work to do. Your weakest criterion is dragging three strong ones down, and the fastest route to a higher band is almost always to find the one that is lagging rather than to polish the ones that are not.

The four are not independent in practice. Hesitating while you hunt for a word is a Vocabulary problem that shows up as a Listenability score. Running out of things to say at forty seconds is a Content problem that shows up as a Task Fulfilment score. Diagnosing which criterion an error actually lands on is most of the skill this lesson is for.

Content and Coherence: how many ideas, and how well built

This criterion asks how well your ideas are organised and developed. It looks at the number of ideas, the quality of them, the order you put them in, and whether you supported them with examples and detail. It is the criterion that punishes an answer made of three unconnected assertions and rewards one made of two assertions with something concrete hanging off each.

The commonest way to lose here is refusing to commit. On Task 1 and Task 7 in particular, listing the arguments on both sides and never landing on one caps this criterion at 7 no matter how fluent the listing was. The task asked what you think or what they should do. A survey of the possibilities has not answered it.

The second commonest way is quantity over development. Eight quick pieces of advice in ninety seconds gives each of them about eleven seconds, which is one clause and no support. Three pieces of advice with a reason and a concrete detail each will use the same ninety seconds and score higher, because the criterion is about quality and organisation of ideas rather than a count of them.

Vocabulary: range, precision and the ladder between bands

Vocabulary looks at word choice, whether words and phrases are used suitably, the range available to you, and precision. The published ladder is unusually clear about what separates the levels, and it is not about knowing rare words.

Around band 5 the entitlement is common words and phrases only, with no expectation of abstract language at all. Around band 9 you are adding vocabulary specific to the context you are talking about, and some idiom. At the top, around 11 and 12, there is a broad range of both concrete and abstract language, and figures of speech appear naturally rather than as decorations.

The practical reading of that ladder is this. Moving from a middle band to a high one is not about swapping ordinary words for impressive ones. It is about being able to leave the concrete level at all. The bakery was busy is concrete and fine. It had the kind of busyness where nobody is quite in a queue is the same observation with an abstraction sitting on top of it, and that is the move the higher bands are built from.

Listenability: how hard the listener is working

Listenability covers rhythm, pronunciation and intonation, pauses and interjections and self-correction, grammar and sentence structure, and variety of sentence structure. The single question behind it is how easy the answer is to listen to and understand.

Accent is not in it. There is no version of this criterion where sounding Canadian earns a mark or sounding like anywhere else loses one. What the criterion measures is effort on the listener's side, which comes from pace, from where you pause, and from whether the words are articulated clearly enough to be caught first time. Clear speech in a strong accent scores well. Rushed, mumbled speech in any accent does not.

The grammar half of Listenability is a ladder of control and then range, not a count of errors. Around band 5 there is some control of simple structures. Around 7 there is good control of simple structures. Around 9 there is some control of complex ones. Around 11 there is good control of a broad and varied range of complex structures. Notice what that means: reliably correct simple sentences will not on their own reach the top, and attempting complex sentences you cannot land will not either. The top asks for both.

Quick tip

After every practice answer, write down which of the four criteria your mistake belongs to before you write down the mistake. Hunting for a word is Vocabulary appearing as Listenability, and treating it as a fluency problem means you practise the wrong thing for a month.

Task Fulfilment: relevance, completeness, tone and length

Task Fulfilment has four named factors: relevance, completeness, tone, and length. Tone being listed there is the part most candidates miss. The register you use with a cousin is not the register you use with a supervisor, and using the wrong one is a scored error rather than a stylistic preference.

This criterion also has its own ladder, and it runs convey, then adjust, then adapt. Around band 5 you can convey information on a familiar topic. Around band 9 you adjust your style and tone for a range of different audiences. Around band 11 you adapt to the situation, to the purpose, and to your relationship with the listener, all three at once. That is why the tasks keep naming who you are talking to: they are giving you the material this criterion is scored on.

Length sits here too, and it is measured against the time you were given rather than against a word count. An answer using well under a third of the allowed recording time cannot score above 6 on Task Fulfilment, because a task you stopped answering after eighteen seconds has not been fulfilled. Completeness means every part of the prompt, including the part at the end that asks you to make a request or say why it mattered.

Penalties are caps, not subtractions

This is the part of the system that changes how you should practise. Problems do not stack up as deductions from a starting score. Each identified problem sets a ceiling on the criterion it affects, and the lowest applicable ceiling is where that criterion lands.

So failing to compare the two options on Task 5 caps Task Fulfilment at 5 for that task. Re-describing the picture instead of predicting on Task 4 caps it at 6. Addressing the wrong person on Task 6, or wobbling between both of them, caps it at 6. Never making the closing request on Task 8 caps it at 8. One underlying mistake produces one cap, not three separate deductions.

The useful inference is about where your effort belongs. If you have a habit that triggers a cap, no amount of vocabulary or fluency work will get you past it, because the ceiling is applied after everything else. Find the caps you personally trigger, remove them, and only then start polishing. A candidate who speaks beautifully and never makes the Task 8 request is holding their own score down by a mechanism that takes one sentence to fix.

Quick tip

Check length against the clock, not against how it felt. Anything under roughly a third of the allowed recording time caps Task Fulfilment at 6, and twenty seconds on a sixty second task feels much longer from the inside than it looks in the file.

Practise it

Five questions on which criterion is doing the damage, which is the part most candidates get wrong.

1Your four criterion scores are 10, 10, 10 and 6. What is your overall speaking level?
2A candidate pauses for five seconds mid answer while searching for a noun. Which criterion takes the hit?
3Two answers to a Task 7 opinion prompt. Which scores better on Content and Coherence?
4On Task 6 you are told you may speak to your neighbour or to the building manager. You spend thirty seconds on each. What does that trigger?
5Which statement about the grammar half of Listenability is correct?