Part 5: Discussion, the Hardest Part
Three speakers, eight questions, and positions that move while you are writing them down. This is the hardest part of the section and it carries as many questions as Part 1.
What makes Part 5 different
Part 5 is Listening to a Discussion. Three speakers rather than two, a clip of roughly a minute and a half to two minutes, and eight questions, which is the joint largest block on the test alongside Part 1. On the real test this part is presented as video rather than audio alone, and it is the only part that is.
The tone is relatively informal and it mixes three things that the earlier parts keep separate: facts, opinions and feelings. Crucially, the speakers will sometimes disagree with one another, and disagreement is the engine of most of the questions.
Like Part 4 and Part 6, the questions appear together on screen after the clip finishes and share one time limit, answerable in any order. So you are reconstructing a three-way discussion from notes, several minutes after it ended, across eight separate questions.
Difficulty in this section is designed to rise from Part 1 to Part 6, and Part 5 is where most candidates feel the step change. It is not the longest clip and it is not the most advanced vocabulary. It is the highest number of things to track at once.
Three speakers is more than one extra speaker
Going from two voices to three does not add fifty per cent to the difficulty. It changes the nature of the task, because with two speakers, identifying one identifies the other. Hear a woman and you know it is not the man. With three, every turn requires a positive identification, and if two of the three voices are close in pitch or manner, that identification is genuine work.
It also multiplies the relationships you have to hold. Two speakers can agree or disagree, which is one relationship. Three speakers produce three pairwise relationships plus the possibility of a two-against-one split, and Part 5 questions ask about exactly those configurations: which two agreed, what all three accepted, who was alone in objecting.
This is why the part is presented as video. Seeing who is talking removes the identification problem and leaves you able to track what is being said. The official guidance is explicit that you are not tested on visual details unrelated to the conversation, such as clothing or objects in the room, so the picture is there to help you attribute speech and read attitude, nothing more.
The three-column grid, built before the clip
You are told there are three speakers before the clip plays. Use that. Rule two vertical lines on your page during the introductory statement so you have three columns waiting.
Label the columns by position, left, centre and right, matching where each speaker sits on screen, and write each name at the top of their column as they are introduced. Position is a more reliable label than voice, and it is available instantly, whereas distinguishing voices takes a few turns to settle into.
Then write everything in the column of whoever said it. You are no longer recording attribution as a separate fact; the geometry of the page does that for you. That saving is substantial, because attribution is the first thing memory drops when three voices are involved and it is what a large share of the questions turn on.
Keep entries short: a stance, an objection, a number, a concession. Two or three items per column across the clip is a realistic and sufficient target. Trying to fill three columns evenly is a transcription trap in a new shape.
Rule three columns before the clip starts and label them left, centre and right to match the speakers' positions on screen. Position is available immediately and reliably, while telling three voices apart takes several turns to settle.
Positions move, and the question asks about the end state
The most reliable Part 5 trap is the speaker who changes their mind. Someone opens by objecting to an idea, hears an argument, and ends up going along with it. Or someone begins enthusiastic and is talked down by cost.
Questions are usually about where things ended, not where they started, and the opening position is the one that got written down first and most clearly. So the note that is easiest to read is frequently the one that is now wrong.
Handle this with movement notation rather than memory. When a speaker shifts, draw an arrow in their column from the old position to the new one. It takes a second, it is unambiguous afterwards, and it also directly answers the question type that asks who changed their view during the discussion.
Listen for the specific language of the turn. 'I suppose if that is the case', 'all right, but only if', 'that is fair, actually', 'I had not thought of it that way'. These are the audible hinges of a discussion, and once you are listening for them they are hard to miss.
Draw an arrow inside a speaker's column the moment they change position. It costs a second, it answers the who-changed-their-mind question directly, and it stops your clearest note being your most out-of-date one.
Answers assembled from two places
In Parts 1 to 3 most answers sit at a single point in the audio. In Part 5 they often do not. A question asking what the three speakers agreed on cannot be answered from any single line, because agreement is distributed: one proposes, one endorses, and the third concedes forty seconds later.
This is the real reason the part is hard, and it explains a common frustration. Candidates say they understood every sentence and still got the questions wrong. Understanding each sentence is not the task; holding the state of a three-way discussion and updating it is.
Watch for partial agreement especially. People in discussions agree with part of a proposal and not the rest, and an option that says a speaker supported an idea can be wrong because they supported only the first half of it. Qualified agreement is agreement with a condition attached, and the condition is very often the tested point.
Be wary too of the speaker who says a lot. Volume of speech is not the same as carrying the group's conclusion, and a talkative speaker whose position was rejected is an inviting wrong answer.
The question screen, and how to attack eight at once
Eight questions on one shared budget is the widest screen in the section, and answering in printed order is how candidates end up guessing blindly on the last two.
Sort them instead. On a fast first pass, answer every question that is about a single speaker's stated position, because those come straight out of one column and take seconds. Leave the questions about agreement, about the group as a whole, and about what someone will probably do, because those require reading across columns.
On the second pass, work the cross-column questions with your grid in front of you rather than from memory. This is the pass the notes were made for, and it is why the columns had to be built before the clip rather than during it.
Throughout, remember that the correct option will be a paraphrase. Discussions produce memorable phrases, and an option reusing one is more likely to be the trap than the key.
On the question screen, answer all the single-speaker questions first and leave the ones about agreement or the group for a second pass. Eight questions share one budget, and the cross-column questions are the ones that need your notes rather than your memory.
Practise it
Five questions on the part that carries eight marks and the most moving pieces.