Task 4 shows the picture you have just described and asks what will happen next in it. You get 30 seconds to prepare and 60 seconds to speak, and the whole task is carried by speculation language rather than description.
Walkthrough published by the official CELPIP Test channel.
Overview
Making Predictions is the second half of the picture pair. The image does not change; the question does. Instead of saying what is happening, you say what is about to happen - to the people, to the situation, in the next few minutes.
There is no correct prediction, because the picture has no real future. What matters is that each prediction is anchored in something visible and expressed in accurate future and modal forms.
Structure
Thirty seconds of preparation, separate from Task 3.
Sixty seconds of recorded speech.
The same picture as Task 3, with a new question.
One take, recorded immediately after Task 3.
What to expect
Because you have already described the scene, the material is familiar and the difficulty moves to grammar. Expect to need a range of future forms inside a single minute.
Confident predictions for whatever the picture makes obvious.
Hedged possibilities for what it only hints at.
A visible reason behind each prediction.
A short overall outcome to close on, rather than a list that simply stops.
Suggested steps and timing
Spend the first 10 seconds of preparation picking two or three people or elements you already described, and the next 20 attaching one next step and one visible reason to each - the dark sky, the packed bags, the queue that is not moving.
Open the 60 seconds with a pivot sentence that moves the answer into the future - "In a few minutes, I think…" - which takes about 8 seconds. Give each prediction roughly 15 seconds, evidence first and prediction second, and use the last 10 to say how the scene ends up overall.
Key strategies
Vary the future forms: "will" for confidence, "is going to" where the picture shows evidence, "might" and "could" for possibilities, and one conditional if you can reach it.
Anchor every prediction to something visible. "The sky is dark, so they will pack up" is a full sentence of value; "they will pack up" is half of one.
Reuse the people, not the description. Naming them again is efficient; describing them again is off-task.
Keep the predictions plausible. The scene is ordinary, so the next few minutes should be ordinary too.
Common pitfalls
Repeating the Task 3 description with the verbs changed, which reads as having nothing new to say.
Using "will" for the entire minute and showing no range of forms.
Predicting events with no visible reason behind them.
Inventing a dramatic ending the picture cannot support.
Final advice
Practise Tasks 3 and 4 as a pair on the same image, back to back, exactly as the test presents them. The transition is the skill: 60 seconds of present-tense description followed by 30 seconds of preparation and 60 more seconds in which you must not describe anything again. Rehearsing one pivot sentence you trust makes that turn almost automatic.
Frequently asked questions
Quick answers about Speaking Task 4.
There is no wrong prediction - the picture has no official future. What counts is that the prediction is grounded in something visible ("the sky is dark, so…") and expressed in accurate future and modal forms.
Task 3 is description in the present; Task 4 is speculation about the future. Repeating description in Task 4 counts as off-task. The pivot phrase worth rehearsing is simply "In a few minutes, I think…".
A range of future forms: "will" for confident predictions, "is going to" where the picture shows evidence, "might" or "could" for possibilities, plus a conditional if you can ("If the rain starts, they will…"). Using only "will" for the whole minute undersells your range.
Two or three, each with a reason from the picture. That is what 60 seconds holds at a natural pace. Five quick predictions with no evidence behind them use the same time and show much less.
Yes. Task 4 has its own 30-second preparation window, the same as Task 3, even though the picture is one you have already seen and described.
Task 4 is a separate answer and does not depend on what you managed to say before it. In practical terms a weak Task 3 can even help: the details you did not use are still on screen and available as evidence for your predictions.