Back to blog
Speaking · Task 33 min read

Task 3. Describing a Scene

You see an image and describe it in detail for someone who can't see it.

Prep
30 s
Response
60 s
Format
Voice recording

Task objective

You look at an image and describe it in detail, as if the listener can't see it. The focus is on precision and spatial organization.

How to complete the task

  1. 1

    In prep, scan the image: foreground, background, left and right.

  2. 2

    Start with an overview: what it is and where it happens.

  3. 3

    Describe by zones (left to right or front to back), not at random.

  4. 4

    Include people, actions, objects, colors and spatial relationships.

Tips & tricks

Use prepositions of place: 'in the foreground', 'on the left', 'next to'.

Describe actions in present continuous: 'a man is carrying…'.

Don't invent what you can't see; describe the observable precisely.

Follow a fixed order so you don't repeat or skip zones.

How it's scored

Your recording is assessed on four equally weighted dimensions, and the result is reported on the CLB 1–12 scale. Isolated mistakes aren't deducted: what counts is the overall impression of your full response.

Content & Coherence

They look at how many ideas you give, how good they are, how you organize them and whether you back them up with examples or details. Two well-developed, connected ideas score higher than five loose ideas rattled off in a hurry. A clear order (opening → development → close) makes your answer feel complete.

Vocabulary

They assess the range of words and phrases, whether you use them naturally, and whether they're precise for the context. Repeating the same word or falling back on vague terms ('thing', 'good', 'nice') lowers your score; synonyms, natural collocations and topic-specific vocabulary raise it. It's not about rare words — it's about the exact word.

Listenability

This measures how much effort it takes the listener to understand you: rhythm, pronunciation and intonation; pauses, fillers and self-corrections; plus grammar and variety of sentence structures. You don't need a native accent — you need clarity and a steady flow. Mixing short and long sentences sounds far better than a monotone rhythm.

Task Fulfillment

They check four things: relevance (you answer what's asked), completeness (you cover everything requested), tone (the register fits the person you're addressing) and length (you use the time without trailing off or getting cut short). Here that means describing precisely what you see, without inventing or over-interpreting.

Most lost points don't come from your English — they come from not fulfilling the task: answering something else, falling short, or using a tone that doesn't fit. Before you record, repeat to yourself exactly what's being asked.

Ready to try it?

Practice this task with a real question.

Practice now