On the real test, Task 4 gives you back the illustration you have just spent 60 seconds describing. Same picture, no changes, and one new instruction: what do you think will most probably happen next. Thirty seconds to prepare, 60 seconds to record.
The five questions below come with pictures of their own. On the exam, one illustration serves Task 3 and Task 4 between them. In this course the two tasks have separate scenes, deliberately, so that you get ten frames of practice rather than five: here it is a self-serve car wash, a ferry terminal, an off-leash dog park, a bank branch and a neighbourhood playground. Each one is built the way a CELPIP scene is built, with several people caught mid action and at least one thing nobody in the frame has noticed yet.
Now the failure, and it deserves to be stated bluntly, because it accounts for more lost marks than every vocabulary gap on this task combined. Candidates describe the picture again. They open with "in this picture I can see a car wash with several bays and a queue of cars", spend 40 of their 60 seconds on the same material the rater heard a minute ago, and squeeze one prediction into the final ten. That answer scores close to nothing on this task, not because it is bad English, but because it answers the previous question. The rater already has your description. The marks are exclusively in the sentences that leave the frame.
Here is the test you can run on your own recording: how many verbs are in the future or the conditional? In a strong Task 4 answer, almost all of them. Is going to, will most likely, would probably, gets soaked the second she straightens up, I doubt that bag stays on the bumper. If you are hearing is and are, you are describing, and describing has already been marked.