Task 3 puts a picture on the screen and asks you to describe it to someone who cannot see it. You get 30 seconds to prepare and 60 seconds to speak - and the same picture returns in Task 4, so whatever you notice now pays twice.
Walkthrough published by the official CELPIP Test channel.
Overview
Describing a Scene is the first of the two picture tasks. The image shows an ordinary situation - a park, a street corner, a family kitchen, a busy shop - with several people doing several things at once.
Your listener cannot see it, so the test of a good answer is simple: could they rebuild the scene from what you said? That rewards order and detail over speed and coverage.
Structure
Thirty seconds of preparation with the picture on screen.
Sixty seconds of recorded description.
One everyday scene containing several people or actions.
The same picture returns in Task 4, which asks what happens next.
What to expect
The scenes are deliberately unremarkable and nothing in them is hard to name. The difficulty is organisation: there is always more in the picture than a minute can hold.
A public space - a park, a beach, a bus stop, a market.
A household scene - a kitchen, a garden, a family moving house.
A workplace or shop with staff and customers.
A community event - a fair, a class, a neighbourhood clean-up.
Suggested steps and timing
Spend the first 10 seconds of preparation deciding your one-sentence overview - what is this a picture of - and the next 20 splitting the image into foreground, middle and background, naming the two or three things in each you can genuinely describe well. Note what looks about to happen while you are there; Task 4 asks exactly that about this picture.
Then run the 60 seconds in that same order: about 8 seconds on the overview, 25 on the main people and what they are doing, 17 on the surroundings, and the last 10 on the detail that fixes the scene - the weather, the mood, the time of day.
Key strategies
Give the overview first. One sentence of context makes every detail afterwards easier to place.
Stay in the present: present continuous for actions, simple present for states.
Use position language - in the foreground, on the left, behind, next to - as the spine of the answer.
Describe two or three areas fully rather than naming ten objects in a rush.
Hedge where you are unsure. "It looks like" and "they seem to be" are accurate and they sound fluent.
Common pitfalls
Darting around the picture, so the listener cannot assemble it into one scene.
Slipping into the past tense, which is the most common grammar error on this task.
Predicting what happens next, which belongs to Task 4 and is off-task here.
Naming objects with no verbs, so the answer becomes a list of nouns rather than a description.
Final advice
Practise with ordinary photographs, one minute each, describing them to someone in another room until they can sketch what you said. That test is harsher than any checklist and it trains the one habit this task rewards: a fixed order, held all the way through, even when something interesting catches your eye in the corner of the frame.
Frequently asked questions
Quick answers about Speaking Task 3.
No. One overview sentence and then two or three areas in detail beats ten items named in a rush. Choose the parts you have vocabulary for and describe those fully: what the people are doing, what they are wearing, what seems to be happening.
Present continuous for actions ("a man is loading boxes") and simple present for states ("the café looks busy", "there are two cyclists"). Slipping into the past tense is the single most common grammar error candidates make here.
Position language: in the foreground, in the background, on the left, next to, behind, in front of. These phrases organise the whole answer, and they transfer directly to Task 4, which uses the same picture.
Yes. The image is on screen through the 30 seconds of preparation and stays there while you speak, so preparation is about deciding an order to describe things in, not about memorising the scene.
Move outward and then inward: the background, the weather, the time of day, then what the people might be feeling or how they seem to be related. Careful guesses framed with "it looks like" count as description and fill the last stretch naturally.
They share one picture. Task 3 describes it in the present; Task 4 predicts what happens next in the same scene. Each has its own 30 seconds of preparation and its own 60 seconds of speaking, so noticing the details in Task 3 gives you material for both.