Speak About the Photo
An image, twenty seconds to look at it, then up to ninety seconds describing it aloud. The content is handed to you — which is exactly what makes it hard.
- 20s prep on the clock, 1+ per full test
- Feeds your Speaking subscore
- Built for the current July 2025 DET format, including the new Interactive Speaking task
Speak About the Photo shows you an image, gives you about twenty seconds, and records you describing it for up to ninety seconds.
Candidates assume this is the easy speaking task because there is no question to misunderstand and no opinion to construct. In practice it produces more dead air than any other, for a simple reason: most people exhaust the literal contents of a photograph in about fifteen seconds and then have nothing left.
The task is not really testing description. It is testing whether you can keep producing organised, varied English about a fixed subject — which is a different and more demanding thing.
What you actually see
A photograph, a preparation timer, then a recording indicator. Everyday scenes: people doing ordinary things in ordinary places.
- A single photograph of a recognisable everyday scene — a street, a kitchen, an office, a park.
- No specialist or cultural knowledge is ever needed to describe it.
- About twenty seconds of preparation while you look at the image.
- Up to ninety seconds of recording.
- At least one appearance in a full test.
How to have ninety seconds of things to say
Imagine a photo of a woman in an orange jacket walking a dog along a wet city street.
The fifteen-second answer that most people give: "There is a woman. She is wearing an orange jacket. She is walking a dog. It looks like it is raining." Then silence.
A structure that fills the time — four layers, each one opening up the next:
- Overview (10s). What kind of scene is this, in one sentence. "This looks like an early morning on a city street just after rain."
- Main subject in detail (25s). Who or what dominates, described properly — clothing, posture, action, direction. Specificity is where the vocabulary marks are.
- Background and setting (25s). What else is in the frame. Parked cars, shop fronts, reflections, other people, the weather, the light.
- Inference (25s). What you can reasonably guess — where she might be going, what time it is, what season, what happened just before. This is the layer that never runs out, and it lets you use conditionals and modals: she might be, it looks as though, she may have just.
The fourth layer is the one people miss, and it is worth the most. Inference is where you naturally produce the more complex grammar the scorer is looking for, and unlike the literal contents of the photo, it is unlimited.
How Speak About the Photo is scored
This task feeds Speaking. What is assessed is the language you produce, not whether your description is factually complete — nobody is checking that you mentioned the bicycle.
That is worth internalising, because it changes strategy. A candidate listing eleven objects in short simple sentences scores worse than one describing four things with subordinate clauses, precise vocabulary and varied structure. Coverage is not the objective; range is.
Relevance still applies — the language does have to be about the photograph. A memorised paragraph deployed regardless of the image is detectable and scores as off-task, in our grading engine and, by all indications, in the official one.
| What the scorer looks at | Why it matters |
|---|---|
| Fluency and continuity | Sustained speech without long silences. This is the criterion the task most directly stresses. |
| Range of grammar and vocabulary | Varied structures and precise words. The inference layer is what produces these naturally. |
| Relevance to the image | The speech must be about this photograph. Prepared paragraphs that ignore the image score as off-task. |
| Intelligibility | Whether you can be followed easily. |
| Organisation | A described scene with a shape — overview, detail, inference — reads as more competent than a list. |
The mistakes that cost the most
One failure dominates this task, and everything else is a distant second.
Do this
- Open with a one-sentence overview of the whole scene before any detail.
- Move deliberately through layers — subject, then background, then inference — so you always know what comes next.
- Use the inference layer generously. It looks as though, she might have just, presumably — this is where the complex grammar lives.
- Be specific. A woman in a bright orange waterproof demonstrates more than a woman in a coat.
- Keep looking at the photo while you speak. It is still there, and it will give you the next detail.
Not this
- Do not list objects. A rapid inventory of the frame is the most common answer and one of the weakest.
- Do not stop when you run out of literal content. That is the moment to switch to inference, not to fall silent.
- Do not deploy a memorised description template that ignores what is actually in the image.
- Do not narrate your own uncertainty — "I don't know how to say this" spends time and demonstrates nothing.
- Do not repeat the same structure for every sentence. There is… there is… there is… caps your grammar range immediately.
How to actually get better at Speak About the Photo
The fix here is a drill rather than a course of study, and it works quickly.
Set a timer for ninety seconds and do not stop early
Any photo will do. The discomfort of the last thirty seconds is the entire exercise — that is the gap you are training out.
Practise the four layers explicitly
Say them out loud in order until the sequence is automatic. Overview, subject, background, inference. Structure is what prevents dead air.
Drill inference language on its own
Take one photo and produce ten speculative sentences about it — what happened before, what happens next, why. This is the highest-value language on the task.
Build precision vocabulary for common scenes
Streets, kitchens, offices, parks, transport, weather. Knowing the specific word for the thing is what separates a described scene from a listed one.
Record and count your silences
As with the other speaking tasks, silences over two seconds are the measurable thing. Track that number across a week rather than judging your fluency by feel.
How much this one task moves your score
Speak About the Photo feeds Speaking, and so Conversation and Production. It appears at least once in a full test.
Its usefulness is out of proportion to its frequency, because the skill it isolates — sustaining organised speech without external prompting — is the same skill that carries the much longer Speaking Sample. Practise this and the harder task gets easier.
It is also the safest place to practise speaking at length, because there is no question to misunderstand. If Listen, Then Speak is going badly and you cannot tell whether the problem is listening or speaking, your performance here answers that question.
Speak About the Photo: questions people ask
Do I need to describe everything in the photo?
No. The task assesses the language you produce, not the completeness of your inventory. Four things described with precise vocabulary and varied grammar score better than eleven things listed in short simple sentences.
What do I do when I run out of things to describe?
Switch to inference. Speculate about where the people are going, what time it is, what happened just before, what the weather has been. This layer is unlimited, and it naturally produces the modal and conditional structures that demonstrate range.
Can I prepare a description template in advance?
A structure, yes — overview, subject, background, inference is a reliable shape. A memorised paragraph, no. Prepared text that does not fit the image reads as off-task and is scored accordingly.
Does the photo stay on screen while I am recording?
Yes. You are not describing from memory, so there is no need to rush through the details. Keep looking at it while you speak — it will supply your next sentence.
How long should my answer be?
Use most of the time available. Very short answers give the scorer little language to assess, and thin responses tend to score low for that reason alone. The common failure is stopping at fifteen seconds, not overrunning.
Does my accent affect the score?
No. Accent is not assessed and never has been; what the scoring listens for is whether your speech is intelligible and whether it keeps flowing. A strong regional or non-native accent costs nothing as long as the words come through clearly.
Can I re-record if I stumble at the start?
No — you get one take and the recording runs to the end. A shaky opening is recoverable, though, because the score reflects the whole sample: keep going rather than trailing off, and the rest of your ninety seconds outweighs a rough first line.
Do grammar mistakes while speaking count against me?
Minor slips are expected and cost little; freezing or self-correcting every one costs far more, because the hesitation it creates damages fluency directly. Accuracy and range are both assessed, so keep producing varied sentences and let the occasional error stand.
How long do I get to speak about the photo?
About twenty seconds to look at it and up to ninety seconds to describe it. The photograph stays on screen throughout, so nothing depends on memory.
What if I do not know the English word for something in the image?
Describe it instead of stopping — "the thing you carry water in" demonstrates more usable English than silence does. Circumlocution is a real skill and it reads as competence, not as a gap.
Do I need to mention everything in the picture?
No, and trying to is counterproductive. A rapid inventory of eleven objects in short simple sentences scores worse than four things described precisely with varied grammar. Range is what is assessed, not coverage.
What kinds of photos appear?
Everyday scenes — streets, kitchens, offices, parks, people at work or travelling. Nothing requires cultural knowledge, and nothing is ambiguous enough that you could reasonably misidentify what you are looking at.
Is it acceptable to guess what is happening in the photo?
More than acceptable — it is the highest-value thing you can do. Speculation naturally produces modal and conditional structures (she might be, it looks as though) that demonstrate exactly the grammatical range the scoring rewards, and unlike the literal contents it never runs out.
What happens if I finish early?
Nothing is deducted, but a thirty-second answer to a ninety-second task gives very little evidence to score on and tends to land below your actual level. Running dry early is the characteristic failure here, and a four-layer structure prevents it.
How is this different from the written photo task?
Same stimulus, opposite pressure. Speaking rewards sustaining organised language for ninety seconds; writing gives you sixty seconds and rewards finishing accurately. One feeds Speaking, the other Writing.
Where Speak About the Photo sits in a full test
No task on the DET is scored in isolation. The adaptive engine is building one picture of your English out of every response, so it helps to know how much of the hour this particular task accounts for and which of the others are drawing on the same ability.
The table below is the whole test. Your current page is highlighted.
| Question type | Time | How often | Subscores it feeds |
|---|---|---|---|
| Read & Complete | 3 minutes | 3–6 times | Reading, Writing |
| Read & Select | 5 seconds | 15–18 times | Reading |
| Fill in the Blanks | 20 seconds | 6–9 times | Reading, Writing |
| Listen & Type | 1 minute | 6–9 times | Listening, Writing |
| Speak About the Photo — you are here | 90 seconds | at least once | Speaking |
| Interactive Speaking | 35 seconds each | 6–7 questions | Speaking, Listening |
| Write About the Photo | 1 minute | at least twice | Writing |
| Interactive Reading | 7 minutes | 2 sets | Reading |
| Interactive Listening | 6 minutes | 2 sets | Listening, Speaking |
| Speaking Sample | 3 minutes | once | Speaking |
| Writing Sample | 5 minutes | once | Writing |
Two things are worth reading off that table. The first is that the tasks feeding Speaking are not only this one — improvement transfers, so time spent here shows up elsewhere. The second is that the tasks taking the largest share of the hour are not the ones most people practise most.
Question types that use the same muscles
Speaking Sample
An extended spoken response to an open prompt. It is sent to institutions alongside your score.
Read the guide → Speaking + ListeningInteractive Speaking
A short spoken conversation with an animated character. You hear one question at a time and record a reply to each, and every question builds on what you just said.
Read the guide → WritingWrite About the Photo
An image appears and you write one or more sentences describing it.
Read the guide →Start a practice test
An hour under real conditions, all eight subscores, and feedback on every written and spoken answer.
Take a test