How we evaluate Writing and Speaking
The criteria, the method, the turnaround, and the things an automated assessment genuinely cannot do.
Last updated 12 September 2026
Two different kinds of marking
Reading and Listening have a correct answer. Each has 38 scored questions, there is no penalty for a wrong answer, and your raw score converts to a CLB band through a published conversion. That marking is arithmetic: it happens the instant you submit, it costs nothing to run, and it is identical for every candidate.
Writing and Speaking have no raw score. They are rated against four criteria, and that is the part this page is about.
The four Writing criteria
Writing Task 1 and Task 2 are each scored 1 to 12 on the four official CELPIP Writing criteria, weighted as the official criteria are weighted.
- Content and Coherence, 30 per cent. Whether the ideas are relevant, clear and logically ordered, whether reasons and examples support them, and whether the paragraphs connect.
- Readability, 25 per cent. Grammar, sentence variety, spelling, punctuation and paragraphing, judged by whether the errors get in the way of understanding.
- Task Fulfillment, 25 per cent. Whether every instruction and bullet is addressed, whether the purpose and opinion are clear, whether the tone fits, whether Task 1 is a properly formatted email, and whether Task 2 picks one option and argues for that one.
- Vocabulary, 20 per cent. Range, accuracy of word choice, repetition, collocation, and whether the register suits the topic.
You get a band per criterion, not a single number, because the four rarely move together and the lowest one is the one worth working on.
The four Speaking criteria
All eight speaking tasks are scored on the four official CELPIP Speaking criteria, each 1 to 12 and each weighted equally, so the overall band is the mean of the four, rounded.
- Content and Coherence. How many ideas, how good, how well organised, and whether they are supported with examples.
- Vocabulary. Range and precision, and whether the words are used naturally.
- Listenability. How much work the listener has to do: flow, hesitation, self-correction, grammatical control and sentence variety.
- Task Fulfillment. Relevance, completeness, tone and length. Tone is scored, so addressing the wrong person or misjudging the relationship costs marks.
What happens to a speaking recording
Your recording is uploaded to a private storage bucket, transcribed by an automatic speech recognition service, and measured. The measurements are taken from the audio itself, not from the transcript: words per minute, filler words per hundred words, the longest silence inside an answer, how much of the allowed recording time you actually used, and the recogniser's own confidence.
Those numbers matter because they are the part of Listenability a transcript cannot show. A perfectly worded answer delivered at 70 words per minute with five-second gaps is not a band 11 answer, and the transcript alone would not know that.
The scorer then works from the transcript plus those measurements. It is told explicitly that it cannot hear the audio and must never claim to have judged accent, tone of voice or emotion. Low recogniser confidence is reported as clarity, never as accent.
How the assessment works
Writing and Speaking are assessed against a detailed rubric that encodes the official criteria, the band descriptors, the specific failure modes of each CELPIP task, and a set of caps. Penalties are applied as caps rather than stacked deductions, so one underlying mistake is not punished four times. Before a result is returned it is checked for internal consistency: every weakness it lists has to be reflected in the criterion it affects, and the overall band has to equal the arithmetic of the criterion bands.
The rubrics are calibrated against the published band descriptors and deliberately anchored low, because the common failure of automated scoring is generosity. Real candidates cluster between 6 and 9, and the rubric says so.
The expert layer
The expertise sits in the rubric rather than in a fresh pair of eyes on every script. A CELPIP examiner is trained to apply a fixed set of descriptors consistently, and that is what is encoded here: the criteria, the band descriptors, the failure modes specific to each task, and the caps. The standard does not drift with who picked your script up, or how late in the day it was.
Our people write and maintain those rubrics, calibrate them against real attempts, watch the submission queue for failures, re-run an evaluation that went wrong, and answer you when you write in. A submission is stored before the scorer is ever called, so a failure never loses your work, and a failed evaluation is alerted on and retried rather than quietly dropped. If you think a band is wrong, email admin@celpipmocks.com with the link to your report.
How long it takes
Reading and Listening are scored the moment you submit. Writing and Speaking are evaluated while you wait, and the report link appears as soon as the evaluation returns, usually inside two minutes. The results email is sent a few minutes after that, and a short follow-up note about what to fix arrives roughly thirty minutes after the test, once you have had time to open the report.
If an evaluation fails, your submission is queued and retried rather than lost, and the report link you already have starts working once it succeeds.
The honest limits
- It is an estimate, not a result. Only Paragon Testing Enterprises issues an official CELPIP score. Nothing here can be submitted to IRCC or to any other body.
- Nobody hears your speaking. Pronunciation is inferred from recogniser confidence and from timing, not heard. That is a real gap, and it is the reason a speaking band here should be read as a range rather than a point.
- A transcript loses things. Automatic transcription makes mistakes, particularly on proper nouns and in noisy rooms. A bad recording produces a bad transcript and therefore a bad band, which is a recording problem and not a language one.
- Assessment is not perfectly repeatable. Submitting the same answer twice can produce bands that differ by one. Treat a one-band difference between attempts as noise and a two-band difference as signal.
- The rubric is our reading of the official criteria. It is not Paragon's marking engine and has no access to it. celpipmocks is an independent study site and is not affiliated with Paragon Testing Enterprises.
Evaluation FAQs
How long does it take to get my CELPIP practice results?
The report is produced while you wait and its link is shown as soon as the submission finishes, which for Writing and Speaking is usually under two minutes. The results email follows a few minutes later, and a short note about what to fix arrives roughly half an hour after the test. Reading and Listening are marked against the answer key and are scored the moment you submit.
Is my practice score the same as an official CELPIP score?
No. It is an estimate produced by this site to guide your study. Only Paragon Testing Enterprises issues an official CELPIP result, and only from a test sat at an official test centre. Treat a band here as a strong indicator of where you stand, not as a number you can submit anywhere.
Who assesses my Writing and Speaking answers?
Both are assessed against the official CELPIP criteria, using rubrics our team writes, calibrates and maintains, so the same standard applies to every attempt. If you think a band is wrong, write to admin@celpipmocks.com and it can be reviewed and re-run.
How is my Speaking assessed if nobody listens to the recording?
Your recording is transcribed automatically, and the audio is separately measured for speech rate, filler density, the longest pause and how much of the allowed time you used. The scorer works from the transcript plus those measurements. It never judges accent, tone or emotion, because it cannot hear them, and the report should not claim to.
What are the four criteria for CELPIP Writing?
Content and Coherence at 30 per cent, Readability at 25 per cent, Task Fulfillment at 25 per cent and Vocabulary at 20 per cent. Each is scored 1 to 12 and reported separately, so you can see which one is holding the overall band down.
What are the four criteria for CELPIP Speaking?
Content and Coherence, Vocabulary, Listenability and Task Fulfillment. Each is scored 1 to 12 and they are weighted equally, so the overall speaking band is the mean of the four, rounded.
Can I ask for a score to be looked at again?
Yes. Email admin@celpipmocks.com with the link to your report and say what you think is wrong. An evaluation can be re-run, and a rubric that produced a clearly wrong result gets corrected for everyone rather than patched for one report.