Describing a Picture Without Listing
Task 3 hands you a busy illustration and a listener who cannot see it. Sixty seconds is not enough to name everything, and trying to is the most reliable way to score in the middle.
The task, and the constraint that defines it
Task 3 is Describing a Scene. An illustration of a busy everyday place appears on screen, you get 30 seconds of preparation and 60 seconds of recording, and the image stays visible the whole time. Nothing has to be memorised. The person you are describing it to cannot see the picture, and that single fact determines what a good answer looks like.
A typical scene is a crowded, ordinary Canadian place: a laundromat, a bakery, a school pickup zone, a bicycle repair shop, a pharmacy counter. It will contain something like seven or eight separate groups of people, most of them caught partway through an action rather than posed. The scenes are built that way deliberately, because Task 4 will hand you the same illustration and ask you what happens next.
Sixty seconds is about 130 to 150 words. Divided across eight groups, that is under twenty words each, which is one short clause per group. Divided across four, it is a sentence and a half each, with room for a verb worth hearing and a spatial anchor. The arithmetic is the whole strategy.
The difference between a list and a description
Here is a list. There is a woman at the counter. There is a man with a dog. There are two children near the window. There is a person reading a newspaper. There is a worker with a mop. Every clause is true, the grammar is clean, and it covers most of the room. It is also almost worthless, because the listener now has five nouns and no scene.
Here is a description of the same corner. The woman at the front of the queue is trying to pay while holding a toddler on one hip, and she keeps shifting him from side to side as she digs for her card. Behind her, a man in a delivery uniform has given up waiting and is scrolling on his phone with a stack of boxes wedged between his feet.
That covers two of the five, in more words, and the listener can see it. The difference is not vocabulary level. It is that the second version has real verbs, spatial relationships, and human intention. There is and has give a listener nothing to picture. Shifting, digs, wedged and given up waiting put a scene in their head.
Four moves that lift a description
Anchor everything in space. A listener who cannot see needs a layout before they need details. Start wide, then move through the frame in a direction: left to right, front to back, counter outwards. On the far side of the room, behind the till, at the window. Jumping around the picture is disorienting even when every individual sentence is good.
Choose verbs over nouns. Most weak Task 3 answers are built from is, are and has. Replacing those is the fastest available upgrade. Not there is a man with a bicycle but a man is wrestling a bicycle through the doorway.
Lay one inference over the description. Say something the picture implies rather than shows. The shop has clearly just opened, because the chairs are still stacked. Inference is abstract language sitting on concrete observation, and it is exactly what separates the middle bands from the upper ones.
Group rather than enumerate. Three separate people are waiting for the same machine handles three figures in one clause and tells the listener something a list of three descriptions would not.
Accuracy matters, and invention is penalised
Because the listener cannot see the picture, they are relying on you completely, and describing things that are not in the frame is a factual error rather than a creative flourish. Inventing content on this task caps Task Fulfilment at 8, which puts a clean ceiling on an otherwise strong answer.
This mostly happens in two ways. The first is guessing at detail you cannot actually see, such as reading an expression as anger when the face is barely drawn. The safer route is a hedge, and hedging is itself a top band marker: she looks as though she might be frustrated costs you nothing and is not a claim you can be wrong about.
The second is importing a scene you practised. If you have rehearsed a supermarket description and the picture is a bicycle shop, fragments of the supermarket will try to come out. This is why practising the structure rather than the content matters: a memorised answer bends towards the wrong scene, and every bent detail is an error.
Choose three or four groups during preparation and write one word for each, in the order you will move through the frame. Coverage is not scored and cannot be achieved in sixty seconds; a clear path through the picture is scored, under organisation of ideas.
Do not start predicting
The strong temptation, particularly for candidates who run short, is to say what is about to happen. That is Task 4, and on the real test Task 4 hands you this exact illustration and asks precisely that question.
Sliding into prediction during Task 3 costs twice. It thins the description you are actually being scored on, and it spends material you will need in sixty seconds' time, when you are looking at the same picture with nothing new to say about it.
The discipline is to describe the frozen frame and leave the tension unresolved. The child running towards the puddle is running towards the puddle. The bag balanced on the open boot is balanced on the open boot. Notice them, describe them precisely, and let them hang. You are building the inventory that your next answer will be made from.
A worked sixty seconds
Take a school pickup zone at the end of the day. Here is the shape of a strong answer, with the words a candidate might actually use.
Open wide, five seconds: This is the pavement outside a primary school at about three in the afternoon, and it is chaos. That establishes place, time and an overall impression, and the impression is an inference.
Then three vignettes, roughly fifteen seconds each, moving across the frame. Nearest to me, a woman in a hi-vis vest is holding up a stop sign, but she is looking over her shoulder at something behind her rather than at the traffic. Just past her, two boys have dropped their bags on the ground and are halfway through a conversation that clearly is not going to end quickly, while one of their mothers waits with the car door open. On the far side of the road, an older man has parked badly, half on the kerb, and is trying to reverse out with a cyclist coming up behind him.
Then close, five seconds: Nobody seems to be in any hurry except the people in the cars. That is roughly 140 words, three groups described properly, a spatial path through the frame, two inferences, and a closing observation. It does not mention the other five groups in the picture at all, and it does not need to.
Ban 'there is' and 'there are' from your practice answers for a week. Forcing yourself to find a real verb for every group is the single fastest upgrade available on this task, and it survives into Task 4 and Task 8.
Practise it
Five judgement questions on describing a scene to somebody who cannot see it.